Transcript
Speaker: You are listening to Kubernetes Bytes, a podcast bringing you the latest from the world of cloud native data management. My name is Ryan Wallner and I'm joined by Babin Shah coming to you from Boston, Massachusetts.
Speaker: We'll be sharing our thoughts on recent cloud native news and talking to industry experts about their experiences and challenges managing the wealth of data in today's cloud native ecosystem.
Speaker: Good morning, good afternoon, and good evening, wherever you are. We are coming to you from Boston, Massachusetts. Today is August 14th, 2026. Hope everyone is doing well and staying safe.
Speaker: um My summer has been great. My parents are visiting us from from India after a ah decade. So it has been fun showing them around Boston and also taking them to some of our favorite national parks.
Speaker: right um I'm hoping that all of you are also getting some much needed R&R this summer um and are enjoying some family time. ah Today, we have another great episode lined up for you, an interview with Phil Andrews, who's the Global Field CTO or Chief Technology Officer at Cast AI, ah focusing on how ah customers can optimize their AI spend.
Speaker: So it be a fun and great should be a fun discussion, um very relevant. So without further delay, let's get Phil on the board. Hey, Phil. Welcome to the Kubernetes Bytes podcast. ah Thank you for joining us. Can you please take a moment and introduce yourself and and what you do at Cast AI?
Speaker: Sure. um Phil Andrews, I'm Global Field CTO here at Cast AI. So I've spent the past kind of five years of my life working with customers on optimizing their Kubernetes. And now we've been kind of tuning towards more of optimizing AI usage, actually.
Speaker: Okay. Okay. Nice. So like right before we hit recording, right? We we were having a conversation about different models. I know were talking about healthcare, but if we were talking about how AI can help improve it. um And one of the, I think comments, maybe not today, but in our first conversation, you said like, instead of using Opus,
Speaker: At that point, it was 4.8 for everything. we and now it opens 5 for everything. People should really be choosing the right model for the work that they want or the expect outcome that they want. um Can you talk about how a user should be making that decision? right Let's keep the the technology and the Kubernetes integration out for a minute. But like if I'm just a user, how do I choose the right model?
Speaker: Sure. And that's that's probably the biggest challenge with ai efficiency today. You can either use a default and we'll we'll talk Anthropic because everybody everybody's you know familiar with Anthropic pretty well on on the coding side of things. But so you choose Opus, you know, whether it's or or Fable and you set it up because you have a hard coding task that you want to give it.
Speaker: And then you forget to change your default. And then you ask it some silly question or some small rudimentary question. Now you've spent $100 having Fable do a simple bug fix that's a 10 line code change because it thought about it for 20 minutes.
Speaker: That's it's it's extraordinarily inefficient. But to have to constantly be manually thinking, is this a sonnet task? Is this a haiku task? Is this an opus task? Is this a fable task?
Speaker: It puts a load on the developers that is very hard to quantify. And it's also very hard for developers to constantly be context switching every time they want to do a new task.
Speaker: Am I using the right prompt or the right model for this task that I'm asking it? And as a result, people just pin to one that's reasonably good and getting them good results, and they just leave it there and just work everything through that model. So whether it's a simple bug fix or whether it's a really elaborate you know feature change, it all goes through whatever the default model that they select is because they don't want to have to be tasked with constantly moving things around.
Speaker: And that just that leads to a tremendous amount of of spend and inefficiencies with AI usage in general. No, 100%, right? Like if I'm choosing Overs for everything, i'm i'm using up all my tokens or adding to the bill. But if I'm choosing, i don't know, Haiku for the complex tasks, then I'm not using all of the, like, I don't know, capabilities that an AI model in 2026 has, right? I'm just restricting myself.
Speaker: and And this is just weird, right? Because at some point we thought the latest Haiku model or the latest Sonnet model was best in class and until the next one and the next next one comes along and like, no, this this is what I need, even though we can solve the the same problem through an older model. um But then like as you ah like Being ah a a CTO, right like you talk to customers, I'm assuming, multiple times a day. ah where do Where do people stand or how do you guide them in consuming models as a service through things like Bedrock or Vertex AI or like... self-hosting models, open source models, and especially right now, right over the last month, we have seen the interest spiking up against ah ki community models and then open source models as well. Like, how do you have those discussions? How do you guide those customers?
Speaker: Yeah. So it's your point in the past. four to eight weeks, we've seen a lot of negative press around, um you know, either AI data leaks, you know, entire code bases being uploaded to places they weren't supposed to be uploaded or ai tooling breaking out of sandboxes.
Speaker: It's dramatically increasing the awareness that just grabbing the latest frontier model might not be the greatest business solution. And so we're seeing more interest in, okay, how how can I bring these models into my environment?
Speaker: How can i maintain data sovereignty? um You know, what what can i use a a public model for versus what should i use an internal model for? you know, to your point on on the open weight, open source models that are out there.
Speaker: They're getting very, very powerful. ah You know, Kimmy K3 is launching. It's in various stages of launch already with different providers. ah It's incredibly powerful. It's it's benchmarking up with, you know, i think better than Opus 4.8 right now. um very, very good model.
Speaker: Do you bring that and host it internally? you know But then you've got to get GPUs. you know So then GPU access becomes a new a new challenge. So there there's a lot of trade-offs in that space right now. And I'm seeing more of the CTOs and the higher level executives that we're talking about starting to prioritize governance and data sovereignty over ease of use and, you know, swiping a credit card.
Speaker: You know, you've got the big providers with the frontier models that are easy. And that's what people started ramping on in the past year. And they went from $1,000 month to $100,000 a month in months.
Speaker: And they're going, this is this is unsustainable, right? you this we We can't keep growing like this. And that's where people are really starting to figure out these these trade-offs. And so this either self-hosted or self-governed model, I think, is really going to start taking off, particularly in any kind of regulated industry.
Speaker: Any industry that's dealing with high levels of PII, financials, yeah you know even just... ah Even marketplaces, right? You've got e-commerce.
Speaker: yeah There's a lot of PII in e-commerce between credit card information and user information, emails, addresses. You can't be sending that to a model that's going to use that data to train on.
Speaker: And if you're using an outside model, somebody else's hosted model. You have no idea, unless they've guaranteed to you that they will not use that data for training, you don't know.
Speaker: um And that's, that's i think, a lot of the big concerns that are coming out now is is how to handle these different trade-offs. But then, Phil, do you categorize like buying codecs directly from OpenAI or buying cloud subscription directly from Enthropic in the same bucket as like a Vertex or Bedrock? Because like at at this point, even large financial institutions have made peace with are consuming public cloud, right, as a service. So they don't have to buy the GPUs, quote-unquote. The hyperscalers worry about procuring them. These guys can just rent it and and and maybe do a three-year contract, right, to get the latest and greatest.
Speaker: Like, do you categorize both of them in the same bucket versus self-hosted? I'm in my own data center or a co-location facility where I physically know the hardware that it's running on. Or there is are are customers more comfortable with trusting AWS and Google with with your ah applications and and and website, like e-commerce applications as well?
Speaker: Yeah. I think there's sort of the the three tiers, right? The straight externally hosted, you know, cool we'll call it kind of SaaS versions of of ai you know, the the anthropic and the open AIs.
Speaker: those are getting viewed as the most scary. And then you got the bedrocks and the AI foundries out there as, okay, this is better. This is a better compromise. you know they're They're not using my data to train on, they're hosting it inside of their environment.
Speaker: um It still might be multi-tenant, yeah know depending on which options you choose. but it's at least a little bit better and a little bit more secure from that perspective. And then you've got the, I want to own it end to end, you know, the really, it depends on kind of what the security model is, how comfortable they are with both public cloud, what the experiences are, you know, you've got,
Speaker: cloud first financial companies that maybe they're completely comfortable with, hey, you know, AI Foundry or Bedrock are are solid. i I trust that they are going to be able to protect my data. I paid for the provision capacity.
Speaker: It's mine. ah But then there's some other ones out there that are like, that's still not good enough. but You know, there's there's a lot of paranoia running pretty deep, particularly with every news cycle when there's a new ai data leak that's coming out or or ah an agent that went rogue and started sending data somewhere because it thought it needed to. And, you know, it's just um it's it's a new world out there and it's evolving so incredibly fast.
Speaker: most teams can't keep up. yeah I mean, week to week, the industry is changing. The tooling is changing. What you adopted six months ago might be trash today compared to what the current available best in classes.
Speaker: And I feel like people are getting whiplash from that because it's, you feel like every time you make a decision, six months later, you made the wrong decision and and people are almost getting like, um you know, ah or analysis paralysis. Mm-hmm.
Speaker: Because every decision seems like it's the wrong decision when you look at it in hindsight six months later. So people are terrified about making the next wrong decision with with how fast things are moving. Yeah, like all the all the Amazon principles about like one-way doors and two-way doors and everything. Yeah, yes, these are two-way doors. But then if you have to go through them or make a decision every six months or every quarter to change strategies, that doesn't look great for somebody in a leadership position, right? Like yup you pivoted an entire organization, asked them to use XYZ model and use these tools. And now a quarter or two later, you're like, nope, let's... uproot all of that and then go over to the new thing, new way of doing things. ah so so then Sorry, go ahead.
Speaker: Oh, and that's exactly what we're seeing with like dev tooling. Yeah, right. Everybody we talked to, oh, we just cut over to cloud code, we just cut over to cursor, we just cut over to open code.
Speaker: just recently, you know, so like, can you work with this? Can you work with this? Because everybody, is nobody wants to go and disrupt their developers again. i mean, developers are set in their ways as it is.
Speaker: They don't like change as it is. yeah just Just adopting AI and agentic coding and everything else has been a huge hurdle, especially for anybody who's been 10 to 15 years in the industry. This is a massive change in mindset.
Speaker: And now to be swapping out the IDE that you've worked with for 10 years for some new tool. And then six months later, everybody says, oh, why are you using that tool? That's way outdated. You need a new one. Everybody's having this this yeah these issues with how do you keep up with the industry yeah without causing you know your yeah mutiny amongst your amongst your team.
Speaker: No, 100%, right? Like even like i'm i'm I'm a PM for my day job. I have cursor up and running. I have skills. I have neat project structures where everything has the right context.
Speaker: And now we are evaluating cloud code. I was like, yeah, cloud code great, right? Like everybody on the internet talks about how great it is, even for PMs. But I don't know. I'm happy with like getting access to the same cloud models in Cursor. like Yeah, I wouldn't even think about how my devs would handle that like if we asked them to to switch IDEs or or change change how they're doing work again. um But Phil, I wanted to pivot now into, OK, Perfect. we We spoke about the challenges that exist today, but how is Cast.ai helping them? I know in our discussion, you brought up something like something called kimchi. Perfect naming. I don't know i don't know what what it's supposed to mean or if it's an acronym for something, but it's a great name just to share amongst people. Can you talk about ah what it is and how it's useful to organizations, leadership, and then developers, as you said?
Speaker: Sure. ah So the funny thing about the name was it was literally come up by the development team. so they were They were having dinner, you know, a couple of beers on a Friday afternoon. the original name, I think, was AI Optimizer or AI Enabler. You know, we had, it was, it was some very generic name. And they said, we, you know, we hate the name. We're like, we need a better name for this. And, you know, if they were out at, um they were getting, they were eating Korean food. They were out at, you know, eating eating, you know, Asian food and
Speaker: but they had kimchi and we're like, why don't we just call it kimchi? Everybody's like, yeah, yeah, that's good. Let's call it kimchi. And so like three or four developers mutinied on a Friday afternoon and just decided to rebrand the product to kimchi and it stuck. And now we've got a whole branding around kimchi. And people like you said, people love it. They're like, oh, that's a cool name. That's a cool product name. We've got the the pepper around it. like It's all branded to it. So it was developer-led naming.
Speaker: That's awesome. Marketing rhythm doesn't get involved. They're not invited to those happy hours. Yeah, no, people love it now. But the goal behind Kimchi is to to solve a lot of what these challenges are. how How do you efficiently use AI and make sure that you're getting the most out of it without spending a tremendous amount of money?
Speaker: So Cast AI was founded on the principles of automation first optimization. So we don't want people having to do everything by hand. Manual optimization is tedious. It's toilsome.
Speaker: Like we talked about earlier, manually switching models between every task that you do is a horrible user experience. Kimchi is designed from the basis to stop having to do that manually.
Speaker: So in our harness, we have a multimodal interface. You can select different models for different agent tasks. The agents will automatically steer between each other and break down chunks of work so that You might build out a plan and that plan is using one of your highest level, most powerful reasoning models.
Speaker: ah For instance, personally, I use GLM five to fantastic reasoning model works very, very well. I'll use that to build out all my plans. And I would specifically say build out a plan that can be implemented, coding implemented by agents like Kimmy or Minimax.
Speaker: You know, using smaller models, make sure the chunks are small and consumable, be able to to work within the the context and the reasoning abilities of these smaller models. And then I'll submit that plan to Opus 4.8. And I'll say, Opus, I want you to judge this plan and grade it against the original asks of the prompt.
Speaker: It'll come back and grade it. And it'll iterate on that until I have, you know, an A grade on my plan. And then it'll delegate it automatically. So we've got a, it's kimchi. So we use, you know, fun language.
Speaker: We have ferment. Ferment is kicking off a long running task. It's an automated long running task where the orchestrator will keep iterating on it until all of the work is done. So I'll run a ferment on the plan that I just created.
Speaker: And the orchestrator will delegate all the tasks to the right model, the right agent. It'll use GIMI, it'll use Minimax, it'll orchestrate however it needs to, and it'll keep going until the end work grades whatever I tell it to. So I'll say I want A minus or above compared to the original plan.
Speaker: And again, I'll have Opus do the grading. So I'm still using those frontier models, but I'm using them in a small amount for checking all of the work of the more simplistic models and the more the more basic models. So the coding work can be done just fine, but they're not as complex at the reasoning and the thinking level work as something like Opus. So I want Opus to check the work. It's like having a senior architect Check the work of your of your you know junior engineers. New college grads. Exactly. And and that's that's the concept that people, that that's how we need to liken model choices.
Speaker: Yeah. It's your staffing, every every engineer now, every software developer is staffing their own team. Every one of them becomes their own manager of their own team.
Speaker: You have your architect, your that's doing your high level reasoning, your thinking, your planning, your creating your plan in your project architecture. And then you have your senior engineers that are doing all of your actual implementation of major feature chunks.
Speaker: And then you have your junior engineers that are doing bug fixes and simple stuff. And then the architect signs off on the work at the end to make sure that it's all you know done correctly. It's the same mindset. That's how you break down your models into that same working model that we use humans for today. the know That's why we have different degrees of humans in an engineering team.
Speaker: You don't build an engineering team out of 10 architects and then have them go do bug fixes. It will cost you a lot of money. Exactly. It's just the same as using Fable 5 for literally everything. so So we need those efficiencies to be translated into coding and that's what Kimchi does. um you know and It uses a mixture of hosted frontier models as people choose. They could use OpenAI, they could use Anthropic, whatever they want, as well as a mixture of serverless, either serverless models that CastAI hosts for like
Speaker: me, and Minimax, GLM, or they can actually self-host those models inside of their environment. OK. And I just provide those API keys or LLM proxy keys. um OK. OK. And we can actually spin them up dynamically inside of their Kubernetes clusters because we do Kubernetes automation as well.
Speaker: They can say, hey, I want to bring Kimi into my environment to get the best value out of it and for data residency reasons. Cool. Great. We'll spin it up and launch it for you. We'll provision a GPU. we'll The proxy will automatically routed into there.
Speaker: The proxy is tracking all of this usage. So you have cost breakdowns on a per project basis, per user basis, per model basis. You get to see where your total spend is. So, for instance, in your case, maybe you've got open AI, Anthropic, and some front and some ah open source models. You can see the full cost breakdown across all of them in a single place.
Speaker: So now you've got ai cost governance, budgets, crackability on everything that you're spending, as well as giving your developers the ability to use the right tool for the right job and get the best out of it.
Speaker: I think I really like it, right? Because um obviously we are having this episode because we we like like the discussion earlier, but yeah I'm thinking like today, if I, if I like, obviously at work, I use models, um but I only get like a a round number at the end of the day, right? Hey, this is how much my spend was, but it doesn't show me the breakdown, right? Like,
Speaker: This is how much you use from Opus, this is how much you use from Sonic, right? Like I use mostly anthropic models. But if at the end of the day, end of the week, end of the sprint, if a dev sees that, yeah, their Opus consumption was too high, even though it shouldn't be as high, like it was 60% when for all the tasks that you described, maybe it should just be averaging out at around 30%, for example, right?
Speaker: Then maybe I can also update my workflows that, hey, yes, maybe I can bring in more efficiencies, compare across teams or across developers that how are you um orchestrating some of these things? And then my question is like the the workflow that we described, right? Like building a plan through GLM, having Opus review it and then having Kimi ah ah run with it.
Speaker: is this a decision that me as a developer has to take? Or is this something that Kimmy can take on my behalf as well? Like if I give it the high level threshold once for each task, does it automatically use these defaults? Or do I need to do it for each task and be mindful of it?
Speaker: Yeah, so the the orchestrator, so that kind of there's an orchestration layer in there, that will automatically delegate to the right agent based on the task and based on what it's actually consuming for for work. So when you when you for instance, when you kick off the ferment, the orchestrator is responsible for getting the whole ferment done.
Speaker: So it'll look at what the chunk of work that's being done is, how big it is, what the complexity is, and it will assign it to research agent or a builder agent, and all of those agents will have the right model assigned to them. So In your workflow, you know, it's kind of a set it and forget it thing. You you you set all your model mappings in the beginning, and then you just run. And it's amazing how good of a quality you can get for results, you know, at the end of it, as long as you're you're thinking in that mindset of,
Speaker: I'm going to have a junior engineer implement the chunks. So I need to make sure that when I build my plan, it's junior engineer readable. Okay. there's The kind of mindset that you get in working in it, but it allows you to run so many parallel tasks.
Speaker: A ferment could take two hours. Cool. yeah You kick it off, let it start running in the background. It'll notify you when it's done. And then you go start on a different task somewhere else. And now you've got two or three or four simultaneous parallel running tasks.
Speaker: Yes, it drives up your AI spend, but you're getting five times as much volume of effort done in the same hours. And you're using less AI for each of those individual tasks because you're using the right models for it for the right chunks of work. It's just within cast, just just in our company, we've switched over to running kimchi as our primary for every developer.
Speaker: Yeah. We have cut our anthropic bill by two thirds. Wow. We've only added about 25% of costs with the kimchi cost. So all in all, we're doing about $400,000 a month worth of token work. If it were say entirely opus 4.8, and we're getting that for around $60,000 month.
Speaker: that That's awesome, right? Like that's less money that you have to raise, um higher margin that you can operate at while getting to the same outcome. So like, that's awesome. Okay. But how do I install kimchi? Right? Like, I also want to talk about like how you can orchestrate on Kubernetes, but is it like an IDE extension? What do, how do I tell kimchi that this is what I want to do and please orchestrate the models for me?
Speaker: Yeah, so we've got, so it's kind of the three pieces. There's the backend serverless models, right? There's the actual open source models that we host. And then there's proxy layer, which does all the reporting and analytics that allow you to get the visibility, setting budgets. You can set model level budgets. So you can say, hey, you can use Opus, but you only get a $200 budget for that.
Speaker: But you've got unlimited budget for all the open source models. Yeah. you know And that encourages developers to make sure they're using the right mixtures. And then you have the the harness layer. So we have kimchi CLI, which is very comparable to cloud code. OK. We're actually benchmarking slightly better than cloud code. If you're doing a straight terminal bench, we're about seven points higher than cloud code and terminal bench.
Speaker: And so that's an option. We also have a desktop engine, which is going to be comparable to a full IDE. yeah You know, can do a full coding interface, agentic coding interface, similar to a cursor.
Speaker: but we also act as a plugin. So if you're a VS code fan boy, you can plug kimchi into VS code, still use your same environment, still use all your existing plugins. And it's just lives within VS code. Same thing with cursor.
Speaker: You can be a plugin within a cursor, same thing within co-pilot within co-pilot. So we're trying to meet developers in the tooling. They already are running in. I think we've got a Zed plugin. That's, uh,
Speaker: in beta stage right now. We're trialing it with some some customers. So our goal is to bring that efficient, multimodal, agentic workflow into the existing developer experience. We don't want to force developers to switch to change to change hardware again. ah you know We all know it's it's challenging, so we really don't want to have to force them to to change tooling yet again.
Speaker: No, no, 100% agreed. um You described three things. if i'm what's Are all three of these paid? Can I get so started with something for free? just to kick not Not even just to kick the tires. Is there a freemium thing that I can just use to optimize my workflow and then obviously be the, I don't know, um leader to to push it inside my organization? like What does that look like?
Speaker: Yeah. So the, we have a community plan, you know, it's a free community plan. It's, There's no specific cap on usage. It's just rate limited. um and And you can use the kimchi CLI with the community plan. You get all the visibility there. You get all the capabilities that the kimchi CLI brings to the table. um And that's that's part that's our freemium, right? It allows uses of our backend serverless models with rate limits on it. yeah um And then it allows all the reporting. And then it also allows you know being able to use the the harness and get the full capabilities out of out of the CLI.
Speaker: we've got, I forget how many it was, it was over 3000 daily active users um between our our freemium plan and our um kind of B2C, you know, consumer coding plans, ah which for launching little over a month ago is is pretty, pretty amazing. So we've been $20 per month plan or for the B2C?
Speaker: Yep. and we are our month plan Yep. um And People would be getting fantastic feedback with it. yeah yeah We're also sponsoring hackathons at companies that we're working with.
Speaker: you know So we've told them, hey, look, run your hackathon using kimchi and we'll give you credits to be able to run your hackathon. We ran our own internal hackathon. We did 10 billion tokens in two days for $4,400.
Speaker: worth of worth of price expenses. yeah If that was a mix of ah sonnet five and fable it was gonna be like fifty eight thousand dollars Yup. Yup. No kidding. That's why when you said 44, I was like wow, holy shit, like 10 billion tokens. Yeah, no, it was huge. i mean, we came out with some amazing hackathon projects. We did it at our company summit.
Speaker: It was an amazing event. Everybody was using kimchi as the, as the core platform. And yeah, we were really psyched about, about the result. This is turning out to be like a, I don't know, job board for, for Cassie. Like yeah, decisions get made at at dinner tables during happy hours. Yeah. This is awesome. I might look at the the individual plan as well, because I'm an Anthropic subscriber even for personal use. But then I was like, if I'm using it for work, I don't know, maybe I do want to try something else to make sure that I'm keeping up with other parts of the the ecosystem as well. So i'll maybe i'll I'll be one of your customers as well and a daily active user. well you can And you can register your Anthropic account.
Speaker: with us and then you can actually have it measure your usage between the open source models and your anthropic usage oh that's true yeah 100 maybe i can switch off the the paid subscription still get basic access to anthropic and then just have kimchi orchestrate um pull it in when when it's needed okay that that's a good plan um but i want to go back to kimchi and how it makes those decisions right like is it a basic router, like we we we have seen a few of these options being available now in the ecosystem or um is it does it apply some of the intelligence to dynamically figure out that yeah even though the user wants us to do Opus 5 for reviewing the plan, maybe Opus 4.8 is good enough because it's like...
Speaker: I don't know, 50% cheaper. I'm just throwing that number out there. I don't know if it's actually 50% cheaper or not, but um is there intelligence in the ah um ah in the routing capabilities to help optimize the token spend in in kimchi as well?
Speaker: Yeah, absolutely. that's, and we're continuing to build up that muscle, right? I mean, it's, it's a, it's a very complex task and we're seeing that with some of the other routers that are coming out out there. um You know, there's, we have quite a bit of intelligence in there, especially with, you know, cascading fallbacks. If you run out of budget, you can automatically default to the the next best option that you have budget for. So there's some kind of interesting routing there around budget maintenance and budget management, as well as to your point,
Speaker: Does this task really need what you what you're what you're seeing it needs? And that's part of where the harness makes some of those decisions by chucking things up and assigning it to the the correct area. And then the router is doing some of those decisions as well. When it reads in and says, hey, you're're you're trying to go to Fable 5 because you want to know what the the capital of France is. you know that's That's not a good idea.
Speaker: question to answer with fable five you know so some of those some of that kind of logic is is included in there and we're constantly iterating on that rule set and improving it over time but then so again this is this was ah again doing some research on cloud compliance api for for work right um apparently you can look at like an admin a security admin for the enterprise subscription can look at the prompts uh user prompts so is ah Is kimchi actually looking at the text, right? Like, Hey, this is what the actual text of the prompt is and then making decisions. So are our customers okay with, with vendors like reading what their developers are telling their AI models to do and, and being not mad and man in the middle, but like, it's still like you're reading what yeah somebody, some employee wants to do in their own, on their own.
Speaker: So it's optional, right? You know, it's, it's a, something you can turn on or off. So if, If you're a ah financial firm that needs auditability on every single prompt that goes through in order for for compliance reasons, yeah, yeah you can you can enable that.
Speaker: If you're a firm that doesn't need that and you want to trust users to to do the best they can and and you don't want the the intelligent routing, you don't want the man in the middle kind of kind of reading and monitoring the the traffic, you can pass it through directly and and there's not going to be any mutations or anything like that that happens. So it's definitely configurable. We understand that there's a place in the market for both Yeah.
Speaker: sides of it compliance and governance around ai is is evolving extremely quickly yeah know we're seeing customers that say hey we want an audit log of every single prompt that comes through because we're gonna use this for our chat bot and we need to be able to audit all of the interactions from customers we wanted to see know what all of our engineers Yeah, it's a bit of big brother. But when you start talking about, you know, Fortune 500s, every piece of software you run on your laptop in a Fortune 500 has some level big brother written into it. Yeah, i know. It's the way the world, right? yeah I agree, right? But once you find out that yeah whatever you put in a prompt, which you thought were private, are open to specific people in the organization, like, oh, okay, I don't know how I feel i' about that now.
Speaker: Right, exactly. that's, I mean, there's there's the ethics concerns there. And we're trying to focus as best we can on protecting developer experience yeah to all extents possible.
Speaker: um You know, our goal is we haven't implemented the full audit logging for those governance purposes yet. It has not been a ah ah mandatory requirement for us yet. So right now we are doing nothing with data retention on any kind of prompts.
Speaker: Everything is pure, you know, pass through. If it's being read or adjusted, the model is being read or adjusted. There's no data retention there. Okay. Today. Now, obviously, if we have business requirements to to change that, to do 100% audit logging, it would be a specific use case it would be specifically enabled it would need to be, we would probably even just put a banner in the actual, you know, interface, the harness saying, you know,
Speaker: These problems are being audited, whatever. Managed by your enterprise. That's it. like Exactly. Exactly. Because you have to, right? and And I mean, this product is being built by our own team. They have to use it. yeah They don't want to put police state stuff in there. They don't want to put big brother stuff in there. And and we're trying to build it for developer experience.
Speaker: But we also have to be able to respond within the market. So it's it's trying to find that good middle ground of ethical governance while still being able to deliver a good user experience.
Speaker: we're leaning more towards going to a, like a prompt grading system. So how good of a use of AI was this prompt? Oh, damn. I like that. Yeah. So then it's not data retention. It's not pure. I'm reading the prompts that my engineers are typing. It's,
Speaker: OK, what's the capital of France went to Fable 5? OK, that's a one, right? That's a terrible use of AI. Hey, build me out a function that does this extravagant feature and, you know, handles all this. And here's, you know, a bunch of context and here's all the background behind it and all of that.
Speaker: 10, great use of AI, right? So you know being able to kind of grade that scoring and show which users are excelling at using AI efficiently and properly versus which are maybe not using it as effectively, it also enables training.
Speaker: right You can go to those users who are not understanding how to get the most value out of AI and do some learning sessions. Have your ACE developers train you know the developers that are struggling and not not using it as efficiently as possible.
Speaker: We'd rather build tooling around that that's enablement tooling than Big Brother tooling. and So that's the that's the idea set that we're trying to work towards as we build these enterprise features. And I really like that idea because I'm thinking like not all my problems are great. Right. Sometimes I do ask.
Speaker: the AI model to generate a prompt for me for a specific use case to make sure it's good. um But if I can, at the end of the day, just do a slash something like slash judge my prompt.
Speaker: And then it looks at all the prompts and tell me, yep, these are all the ways you can improve it. And then I get to iterate over how I am writing prompts every day. I think that would be super helpful, right? Like to any any developer or any user for that matter, any persona, like product managers as well.
Speaker: Like, hey, you didn't provide this context. If you did this, it would have reduced your token spend. But if you did this, you would have gotten a ah that the you have you would have crossed that A- minorus or A-plus benchmark sooner ah that that you wanted to hit. So I i really like that, man. Damn, so many good ideas.
Speaker: we're We're kind of excited about it. um you know we're We're actively kind of building and tuning that and testing that capability. And at the end of the day, these companies that that think they want 100% auditability of every prompt that goes through, they're going to absolutely drown in data where it's no longer actually useful.
Speaker: yeah right the sampling on it is is useless at that point. When you're talking, you've got 50,000 developers and they're doing 5,000 prompts a day, you know especially when agents are iterating on prompts and you've got just authentic iterations on prompts.
Speaker: you know agents generating prompts for other agents, you could end up with, you know, millions of prompts a day. You're not going to be able to audit that. But but if you can distill that down into metrics, right? If you can distill that down into, is this a good or is this bad? know, is this complex? Is it simple?
Speaker: Now you can distill that down into a more grading rubric. And maybe for the individual users, like you said, on their specific CLI, they get to see a great prompt.
Speaker: But in the aggregate, you see a user and then scores of how important their usage was. That's the tactic we'd much rather take. That way, individual users can get good feedback on a prompt-by-prompt basis, whereas the enterprise can get feedback on overall content.
Speaker: ai efficiency and AI efficiency. Yes, 100%. No, this is great, right? Like, um I don't know. But I do want to go back to the Kubernetes thing that you brought up earlier.
Speaker: how How is Kim Chi helping orchestrate workloads on on Kubernetes? And are these only like, are EKS, AKS, GKE like things in the cloud? Or can you also orchestrate things on-prem and help me deploy these models on demand and and basically...
Speaker: make make use of this this infrastructure or or Kubernetes ah stack that I have? Sure. So the way we've built everything is it's compartmentalized. So for instance, the proxy itself can be run within the um within the customer's Kubernetes environment.
Speaker: So okay when we start talking about sovereignty and data residency, you can host the proxy in your Kubernetes environment because we're as a company, we manage Kubernetes for a living. So we can launch the proxy into your environment.
Speaker: That keeps your developers talking to the proxy inside of your environment. Now, where you go from there is the next question. Do you go external to an Anthropic?
Speaker: Or do you have models internal to maintain that data residency or bedrock or something like that? yeah You can also self-host any of the open source models. And we can spin those up inside of, again, your Kubernetes cluster. We can launch those models.
Speaker: The GPU challenges are the next area there. Like getting a B300 in AWS right now in US East 1, good luck. Like that's like winning the lottery, right? You're not going to get an on-demand B300 in AWS in US East 1. It's not going to happen.
Speaker: Using Omni, we can actually source GPUs from other NeoClouds and be able to attach them to your Kubernetes cluster. So we can get a B300 from, you know, one of the big NeoClouds out there.
Speaker: Get it, maybe on demand, maybe a one year commit depends on kind of what the what the terms are. And then we can load that into your Kubernetes cluster and load Kimmy K3 on it. Sweet. Now your proxy is in your environment.
Speaker: Your Kimmy K3 is in your environment. The only thing at that point that would egress to Cast ai is your ones and zeros, your metrics. So all of your prompts, all of your data, all of your important stuff is staying entirely in your environment.
Speaker: And the next step is we'll actually be able to launch the metrics ingestion inside of your environment as well. So we're taking that whole UI console, all of that will be the next piece that we can launch inside of a self-hosted environment. Our goal is that end-to-end, all data, all metrics is entirely governed by your environment. And it's all running within your Kubernetes cluster without anything coming back to CastDI. That's where we're headed.
Speaker: ah copies Sovereignty is... is rapidly becoming one of the biggest concerns for the CTO level and CSO level that's out there. Okay. So then you you just slid in Omni right there. Let's talk about what Omni is. Let's take a step back and and talk about what what that is. um How does it compare to kimchi? And and yeah, I know we already referred to a use case, but if you can expand more on that.
Speaker: Sure. um Omni is solving the problem of GPU availability in the public clouds. um What we're finding is the only way to get GPUs in most cases in any of the big public hyperscalers is three year commit with a large amount of spend.
Speaker: yeah You have to say, i want 100 B300s and I'm going to commit to three years and I'm going to pay you 50% upfront and then the rest monthly over the next three years. You're now locked into a capacity where If you're ah Fortune 100, that's probably not an issue, right? You write the check, you get the capacity.
Speaker: You might be burning a little bit of cash for the first six months while you're figuring out how you're actually going to use them, but you're willing willing to tolerate that because you know having the capacity is more important than not. If you're further down in the Fortune 100 list or Fortune 500 list, or if you're a Series D startup that's trying to run ai in in your platform and in your tooling,
Speaker: you don't have $20 million dollars or $30 million dollars to commit to a three-year deal with AWS to be able to get that capacity. You don't even know what you need. If you're talking inference work, like you have no idea what your inference load is going to look like a year from now. So if you buy a hundred GPUs and it's too little, okay, that's a good problem to have. You buy a hundred GPUs and you need five of them.
Speaker: You've just burned a tremendous amount of cash for no reason. And you got to figure out how to offset that. Like that's, that's a make it or break it for some of these, you know mid-level startups. That's where Omni comes in. You need,
Speaker: Five GPUs on a one-year commit while you're starting to develop your project, while you're starting to figure out what you're going to do. You want to run maybe Kimi K3 or GLM 5.2 and stuff inside of your environment.
Speaker: Cool. We'll go get those from one of the one of the other providers. We've got partnerships with a bunch of the big GPU NeoClouds out there. okay We'll use those. They wire into your Kubernetes cluster.
Speaker: They look like native nodes in your Kubernetes clusters. So when you spawn a pod that's requesting a B300 and pulls down a model, it might get launched in Pennsylvania in a data center of some, you know, NeoCloud.
Speaker: But that's okay. The the latency is negligible. You're talking most inference calls are 500 milliseconds at best, possibly one to two seconds, you know, at worst.
Speaker: you're adding 20 milliseconds of latency at worst. It's negligible in the grand scheme of inference, but it's allowing you to avoid those huge three-year commits. It's getting you better capacity with a lower cash outlay at really good prices. And that's where Omni is really useful. And honestly, it's probably our...
Speaker: it might be our second largest, or it might be our first fastest growing product right now. Like it's the, the adoption of Omni has been phenomenal because a lot of these midsize companies just can't get the GPU capacity they need, without bankrupting their company. And so we've been enabling that and it's, it's, it's been amazing how, how much adoption we've been getting.
Speaker: So is, is Omni, um, I'm thinking about two use cases, right? One is, is it like a spot instance, right? But does it doesn't sound like it if it's just like if if the startup, seriously startup still has to commit for a year, right? So it's not like spot right now, but it can be like a spot where if, ah i don't know, older generation of GPUs are just sitting around and in new clouds, it can go and pull them and and make that connection. Right now, it it more sounds like it's an agent-based GPU finder where you just put your needs in and then it goes in inventories or the infrastructure that's available at different NeoCloud providers. And maybe you're partnering with multiple one of those multiple of those. It has the intelligence to figure out, hey, my source workload is in Virginia and US East 1. So I need to pick a GPU that's closer or at least on the East Coast. And that's the level of intelligence that it brings in. It's not really a spot where it's on demand. It can come and go away. GPUs can come and go away. Is that right?
Speaker: So it's it's going to be a bit of both. Right now, it's more of the, you know, we can get some on demand. You know, they're not spot spots, but they're you on demand. We have a certain level of that today.
Speaker: a lot of it is going to be short commits, six months to one year commits, you know, so that they're much less painful on on the commit side of things. And to your point about older capacity, we are seeing quite a bit of interest in that because there's a lot of companies that they are happy with an A100. Right.
Speaker: Or they're happy with an H200 still. H200s are rapidly becoming old capacity, you know which is crazy. But there's companies that need those, right? And there's still a big demand for those.
Speaker: Fantastic. Great. you know Being able to pull and get more value out of those, regardless of where they're located, um Our goal is to move to more of a spot capacity. And we have a couple of partnerships we're working on to develop that that kind of framework out of.
Speaker: We ask them what inventory is available right now. We grab a machine, we lock it in, we load it and maybe we have seven days guaranteed usage without any an interruption.
Speaker: or Or they give us, we use it as long as we want, but they give us seven days of notice before they interrupt it. So that way we can tell the user, hey, this machine is going to be going to be turned off in seven days. make sure you migrate off of it. So we're working on some of those relationships right now because that's what's really needed for this experimentation phase.
Speaker: you know If a company is just trying to figure out how AI is actually fitting in and and what the usage looks like and what the patterns look like, great. I need it for two weeks to run some experiments.
Speaker: Perfect. you know that's That's where this kind of trade-off in this marketplace goes. But if somebody else comes along and says, hey, I'll do a one-year commit or a three-year commit, they say, we need that box back. We're going to go give it to this person who's committing to it.
Speaker: Yeah, that's that's completely acceptable to most companies that are out there while they're in this figuring it out stage. OK, so that's kind of where we're headed with a lot of these relationships and and we're getting really, really positive feedback.
Speaker: Yeah, no, 100 percent. And it's a real problem, right? Not getting GPUs when you have you you need them is, is is ah I guess, the main challenge of 2025, 2026. Well, it's impacting funding, right? Like if you're a you're a late series startup, if you're a series B, C, D startup, and you're getting told by investors, your product needs to be ai based, you know, or you're not going to get your next round of funding.
Speaker: and they say, well, we, we're either going to go bankrupt going through a a hosted provider, or we got to commit to this massive three year spend on something that's really a gamble. Like that's terrifying, you know, to that the the life of your company is in the balance here. Yeah.
Speaker: and And so that's where they're that's where a lot of these founders of these, you know, 2020, 2021, 2022 startups are, is how do we become an AI native company without completely going bust and breaking all of our cash that we've got in the bank so that we can get to that next round of funding?
Speaker: Because you're not going to get another round of funding if you don't have the AI problem figured out in your workflow. Yeah. Yeah, no, 100%. But then, so like, OK, I found capacity in in Pennsylvania, as you said, right? And I need to pipe that in.
Speaker: Is Omni establishing like an IPsec VPN tunnel or something for that communication? Or Omni is just getting me the GPU and now it's on the user to figure out how to do this cross hyperscaler, cross data center connection? It handles the full end-to-end connection. So it'll wire it into your Kubernetes cluster. It'll register it with the control plane. It will show up in your list of nodes. If you go in and do a kget nodes, it'll show up as a node in your in your cluster as and as if it were a native node. And if you've got a pending pod that's asking for that GPU, it'll get scheduled on that pod and it becomes a seamless piece of your cluster. Other pods in the cluster can talk to it. You now have your your API endpoint. it's
Speaker: It's completely seamless from that perspective. So many questions there, right? Like if it's EKS, ah they have ah like the concept of worker groups or node groups or something. um Specific AMIs being launched on those worker nodes before they get added to your cluster.
Speaker: Are you faking any of these things on the NeoCloud GPU node and ah faking that relationship to EKS? Or there is a partnership and EKS has, like, i don't know, given you some templates to follow and then you can plug those in. Like, how does that work?
Speaker: Yeah, um no shenanigans, no partnerships either, really. It's just, you know, and and i unfortunately, I can't go into the deep technical details because I'll be honest, I have no idea how it actually works to that level of depth. But um no, it it'll connect it into the yeah EKS control plane. We register it with yeah EKS as a native node. We don't need to use their AMI. um We install Kubelet and set it all up. that We configure the remote machine po to map to EKS and be configured to talk natively to the EKS control plane.
Speaker: So we've got you know a Kubelet config and everything that wires in and and looks like a native EKS node. And As far as EKS concerns, it might as well be in the same data center. Yeah, so if I'm upgrading the version of Kubernetes, because of Kubernetes, it it does push out the upgrades to even the hyperscale, sorry, the NeoCloud node and upgrades the version as as needed.
Speaker: OK. Damn. that Again, that's, again, innovative. So this is this is awesome. Yeah. Like on Cast's interface, on Cast's interface, if you look at it, you'll see 10 EKS nodes. And then you'll see two nodes that have a slightly different naming convention. And it'll show a region of you know whatever the remote region is. yeah And they look like they're just part of the cluster.
Speaker: And your applications will spawn on them and pods load on them. Your daemon sets will get loaded on them just fine. Yep. OK. And like like, I just want to use an example, right? Coreweave, for example, like NeoCloud provider. Me as a customer, do I need to have an account with Coreweave or I'm just ah a CastEI Omni customer and i maybe I selected from a list of NeoClouds that I'm comfortable working with and then on demand, it's going and bringing that up. So who is signing up for that capacity from Coreweave or the NeoCloud provider? Is it Cast or is it the customer? Sorry, I know it's a lot of details. no, no, it's fine. um We've got customers that work both ways.
Speaker: okay So you can bring your own account. If you already have a relationship with CoreWeave and you want us to just leverage that relationship to do the wiring, great. If you're just like, hey, Cassia, can you just find me GPUs? We'll...
Speaker: find the GPUs, we pay for them and we just add the cost onto the, you know the Cast AI bill at the end of the month and it ends up being a pass-through. So they only have one vendor to work with. They don't have to have a relationship with six different Neo clouds with accounts set up to make it work.
Speaker: They just pay us and we figure everything out on the back end. We've got customers working both ways. It depends on kind of, usually it's the size of the customer, you know, customers that are big enough to have a commercial relationship with, you know, a core weave or a Crusoe.
Speaker: They've got some pretty big infrastructure. If you're one of those kind of mid-series startups, you probably don't. um it's them They just assume have the single vendor. They say, go get it. Here's the price I can pay for it. Here's what I can afford.
Speaker: I'd like to do a one-year term, see what you can do, and we'll go figure it out. And we just, you know, add it into the contract. Okay. Okay. No, that's super cool. um I know we are running close to time.
Speaker: I want to wrap this up by asking, like, there are so many exciting things coming ah going on at Cast.ai, but what's next? Like, what what's something that users should keep an eye out for? Are they just new releases of Omni and Kimchi or other exciting products that maybe you can give us a sneak peek into?
Speaker: ah Honestly, our our desktop kimchi is is really exciting. It has chat. It has a full design interface. So you can do UX UI design in there and then it has a full coding IDE built into it.
Speaker: It's also going to be multi user. So you'll actually be able to coordinate agents amongst multiple users in your team. So if you've got a team of 10 engineers, you'll be able to see what tasks are running on other people's work. And if they've got a dependency that they're working on that unlocks your next step on your project,
Speaker: the agents will actually be able to communicate together. And when their work finishes, it can kick off the agents on your side to start your work on your side. So that kimchi studio is going to be really powerful for team collaboration. you think Think Kanban for agents, right? yeah yeah I'm blocked by this. Okay, it's complete. Now I'm unblocked.
Speaker: Rather than waiting for the next sprint, it can be minute to minute. yeahp yeah you know so You don't have to wait for the next sprint for somebody to close their ticket before you can pick up your task. It's just going to happen automatically.
Speaker: And that's going to be extraordinarily powerful. We're really excited about that. um you know and And it all is going to be done in an efficient manner. So we're really excited about the level of ah kind of velocity that all of this is going to unlock in development teams.
Speaker: No, a hundred percent. Okay. Last question. How, where can people learn more about Cast.ai or if they want to listen more of what you have to say, like are there blogs, podcasts? don't know. Things that you can point us to.
Speaker: Yeah. So kimchi.dev has ah blog posts on it that go over a lot of you know techniques around how to use AI effectively, you know what governance looks like in the next generation. We're getting a lot of interest there. I've been on a few different podcasts. you know People can look me up on on LinkedIn pretty easily. you know Philip Andrews on on LinkedIn at CastAI. I'm happy to always talk to people. like Yeah, optimization stuff is super interesting. like That's been my focus for the past six, eight months. you know and I think that's that's where everybody is kind of heading at this point is is how do we use AI, but use it efficiently. and And I'm really excited about that space because we've been working in efficiency for five years now. And I'm really excited for this new generation of of things to optimize.
Speaker: Okay, so that's awesome. I thought I was special hosting you on the podcast, but looks like you you have been on another podcast as well. This is the first AI kind of discussion that I've had on a podcast. So this is this was kind of cool to be able to actually talk about you know some of the AI stuff. Most of the past has been strictly Kubernetes. Kubernetes related. Okay. I don't know. Phil, I do appreciate our time. This has been a super fun episode to record. Thank you so much for your time. And you are in welcome back any anytime you want.
Speaker: Awesome. Yeah, appreciate it We'll definitely take you up on that. Thank you. Okay. That was a great interview. Hope you guys liked it as well. Kimchi definitely feels something to try out. i know I'm personally going to look into it, even if the free version is is usable enough um and then something worth trying out um to to see if I can optimize some of my workflows, right? We do have similar episodes lined up ah from practitioners from our ever-evolving ecosystem.
Speaker: So I would really appreciate if you can share a link of this episode on one on that one Slack channel at work where people post random or interesting things or share it on that one text chain that you have with with your ex-colleagues or people that you used to work with but are still keeping touch with.
Speaker: And with that pitch, it brings us to the end of another episode. I'm Bhavan and thank you for listening to the Kubernetes Bytes podcast.
Speaker: Thank you for listening to the Kubernetes Bytes podcast.


