Zencastr
00:00:00
00:00:01
Speed1x
Format
Share
Embed
Report

Your CI Tool is NOT a CD Tool!

Kubernetes Bytes
Kubernetes Bytes

0 plays · Aug 28, 2026

In this episode of the Kubernetes Bytes podcast, Bhavin talks to Pushkar, Staff Software Engineer at Snap Inc. The discussion starts by talking about how platform teams support developers by building internal developer platforms, what are the capabilities developers expect in an IDP. They then talk about how reusing a Continuous Integration tool for Continuous Deployment led to issues and pagerduty alerts. Finally the discussion focusses on Argo Rollouts and the benefits it offers for rolling out updates to your production environment. This is a real world scenario where Pushkar shares lessons learnt, best practices, etc. from his experience onboarding Argo inside their organization.     Check out our website at https://kubernetesbytes.com/   Show Notes: * https://argoproj.github.io/rollouts/  * https://argo-cd.readthedocs.io/en/stable/  * https://buildkite.com/home/

Transcript

Speaker: You are listening to Kubernetes Bytes, a podcast bringing you the latest from the world of cloud native data management. My name is Ryan Wallner and I'm joined by Babin Shah coming to you from Boston, Massachusetts.

Speaker: We'll be sharing our thoughts on recent cloud native news and talking to industry experts about their experiences and challenges managing the wealth of data in today's cloud native ecosystem.

Speaker: Good morning, good afternoon, and good evening, wherever you are. We are coming to you from Boston, Massachusetts. Today is August 28th, 2026. twenty twenty six Hope everyone is doing well and staying safe.

Speaker: um Today, we have another great episode lined up for you guys, an interview with Pushkar, who's a staff software engineer at Snap. ah But we are going to talk about his experience transitioning the continuous delivery tool and adopting Argo CD and Argo rollouts while he was a staff senior engineer or staff software engineer at at Cruise. ah Should be a great conversation. So without further delay, let's get Pushkar on the pod. Hey Pushkar, welcome to the Kubernetes kubernetes Bytes podcast. If I can just say that name, I know it's it's Friday right now. Thank you. Thank you for having me, Bhavan.

Speaker: Yeah, why don't you introduce yourself to our guest to our listeners, please? Yeah, sure. Hi, um I'm Pushkar. I'm a staff software engineer at Snap. um So prior to this, I worked at a company called Cruise. I was a staff software engineer at Cruise in the infrastructure space. And prior to that, I worked at Amazon for quite some time. um So I have a little more than 15 years of software development experience, primarily on the backend side.

Speaker: And for the last couple of years, I've been mainly focusing on in the infrastructure side of ah you software development. um yeah that's that's quick introduction yeah no that's awesome thanks uh fun right like i think you were you might have been like if you said 15 years it looks like you were there when aws was on that crazy great growth trajectory i'm not saying they're not anymore but at least yeah yeah you were number one and yep uh every reinvent was crazy Yeah, yeah, yeah. There's just like so many things that the company was doing with AWS, like so many different services that were coming up. And yeah, it was crazy. Like I saw the whole growth phase and um there was a point in time at Amazon where ah we weren't really using much of AWS.

Speaker: ah then we kind of started dogfooding it ourselves and so there's a push to move towards AWS and yeah so we were trying to get all our services onto you know AWS and use it as much as possible.

Speaker: Gotcha. Okay so then let's let's focus this discussion on your time at at Cruise right as you mentioned ah you did spend some time around infrastructure was it also cloud infrastructure given your AWS experience?

Speaker: Yeah. Can you talk about what part team you were part of, what did you guys own, and then what what um yeah responsibilities look like? um Yeah, so at Cruise, so I was a part of the cloud infrastructure team. um So our team was responsible for most of the cloud infrastructure rob that the service typically ah you know needed at Cruise.

Speaker: So um like a service, like a running service in production needs a lot of different infrastructure components. um First off, it needs like a compute infrastructure. um ah where it has to run the business logic. ah It can be like a Kubernetes cluster, maybe in GKE, EKS, or in Azure, or it can be like a VM, or it can be like a um a serverless sort of a thing like Lambda. so um ah It needs like some compute to run the business project.

Speaker: and In addition to compute, it might need like a storage infrastructure, like a SQL database, NoSQL database, or a blob store. um so It needs like storage infrastructure. It needs like different other den know sorts of infrastructure, like a CI, CD infrastructure. um so You'd want like you know pipelines, CI, CD, pipeline to be able to you know build your code, generate the image, run the tests, and then deploy to the different environments that you might have. And it like a service would need like ah observability infrastructure as well, like you know metrics, logging, monitoring, all of that.

Speaker: so there are a lot of different infrastructure components that are needed, and some of it is like visible, some of it might be invisible. yeah ah So what we did at Cruise was um Essentially, we were responsible for ah creating all the different infrastructure components.

Speaker: ah We had like an internal developer platform called Juno. and um so that was responsible for um us when When a service developer wanted a service, ah it was responsible for ah you know creating all these different infrastructure components and also linking them together.

Speaker: of That was actually the important part because, you know, um anyone can just create like a cluster and, you know, have pipeline and create these isolated components. But stitching it all together was the hard part because, you know, they had to kind of be integrated and work together as one cohesive piece. And only then, you know, it's like useful to the developers who want that. right like For instance, ah know if you have a Kubernetes cluster, you should make sure that the CICD pipeline has the right permissions to be able to deploy to the cluster.

Speaker: Or ra like... um need to make sure that the cluster is on the right you know networking stack that you would want. Or ah if you have like a database or pubsub queue, whatever, you need to make sure that it has permissions to read and write to these you know infrastructure components.

Speaker: So yeah, that that was essentially what we did at Cruise. No, no, no, for sure. that Again, you summarized it in like a minute, but obviously it's not as easy as that. Was it something that was built in-house or um yeah something like backstage, right? It it was popular. ah Yeah, yeah. We did we did briefly ah play around with pla backstage as well, but ah no, we had our own custom platform.

Speaker: ah platform that was you know doing this. um But we used a lot of Kubernetes controllers behind the scene to you know accomplish this. ah so we used stops We had our own Kubernetes operators to create a cluster, create namespaces on it, integrate it with you know the rest of the ecosystem. We used something called cross-plane for infrastructure as code.

Speaker: ah that that was really nice as well. so We were able to use it to create like opinionated infrastructure components. That was another you know great value out of this. um So we would create like ah really opinionated you know infrastructure components. like If you think about it, if you create a database, so it's not just about having the database. rate You need to make sure that the database has like the right auto-scaling setup or ah it has a point-in-time recovery setup, it has automated backups, it has the right permissions. and So there's a lot of um

Speaker: best practices that like a lot of infrastructure components need. and We packaged it all into opinionated infrastructure components and integrated it with the internal developer platform.

Speaker: so Essentially, like you know anyone any service developer at our company was able to use like infrastructure with all the best practices already incorporated into it.

Speaker: Interesting. So you brought up like databases and having backups and point in time recovery. So were you also orchestrating production deployments using the IDP or like because you think about test dev or development life cycles, right? Like you don't have really have to worry about it unless you are like cloning a production database to make sure that you are testing against the right set of things. Can you talk more about that?

Speaker: so ah so There are two sorts of deployments. right There is the application deployment to production and there is the infrastructure deployment. and so We had like ah you know two different CI-CD pipelines for ah you know both of them. There was definitely some infrastructure that was kind of managed through a UI and it didn't really have like a deployment per se. But there was you know a lot of opinionated infrastructure that we could deploy through a CI-CD pipeline.

Speaker: and so That was where we had ah you know a promotion process where, ah let's say, we set up auto-scaling for our Spanner database and it obviously wasn't there. The change would get applied to the development you know database and where you know we could run some automated tests. and If it was working fine, it would progress the staging and then to prod.

Speaker: yeah so We had two different pipelines. No, thank you for going into that level of detail. Right. So I know we we are a few minutes into the recording. We never introduced the topic of the episode. And on our first call, at least you helped me ah realize that, hey, we'll talk about Argo, but we sure that you understand that Argo CD is different than Argo Rollout. So can you talk about the Argo ecosystem first and then we'll talk about how how ah that that journey happened at Cruise, but how how is Argo CD different from Argo Rollout to anything else that Argo ecosystem might have?

Speaker: Yeah, yeah, yeah. So um Argo CD is um a continuous delivery system. um so um Typically, you know you'd want like some CI, CD system ah to deploy to, you know first off, like build your code, test your code, run you know different sorts of integration testing, competent testing, load testing, and then ah ah then be able to you know deploy your code to different environments that you would have in in a typical service.

Speaker: so There are two distinct aspects of this. One is the CI aspect, the continuous integration part of it, and then you have the continuous delivery part of it. um Argo CD helps with the continuous delivery part of it. So Argo CD essentially helps um ah service developers like you know ah implement GitOps where you know you can define your your state of the system in in a Git repository and Argo CD updates it, like make sure that the state of the code on the cluster matches the state of you know your code and your configurations in the Git repository. um so Whenever a deployment is needed, um the CIA part of it essentially you know runs your code, builds it, and uploads the images to ah some sort of a container registry, and then updates like um

Speaker: a Git repository or like you know there are different sorts of you know riing like trigger operations for ah deployment. But um what I'm describing here is like more of a GitOps model where um you know the Git repository is updated and then Argo CD is responsible for ah ah like deploying the code to the cluster.

Speaker: Yes. Essentially, it starts the port, it gets the OCI image from the container registry, um and then it starts the deployment. so that That is the Argo CD part of it. and Argo rollout is um it's but It's specific to progressive delivery. And um so with the typical ah deployment that you get with Kubernetes, ah you know you get click you know simple, like straightforward like deployment strategies. And there are like you know a lot of downsides to the basic deployment strategy. So Argo rollout helps with the progressive delivery aspect of it.

Speaker: Right. and um so The deal here is like with progressive delivery, you can essentially um roll out to your code in a controlled manner where you can know start off very small and maybe evaluate the state of your service as the rollout progresses. and If in case ah you know you you have some metrics that determine that the rollout is not healthy, you can roll it back.

Speaker: um So that is progressive rollout, like the concept of like you know introducing your code, your new version of the code to like a small subset of customers first, and then you know trying it out, looking at metrics, and then taking a decision as to you know whether to ah continue the pro rollout or to roll it back.

Speaker: Okay. no the it seems super helpful, right? Like, and and needed, especially now with AI-generated code, you might not want to release it to all of your customers at the same time. And we have always spoken about, like, as an ecosystem, right? Like A-B testing and then blue-green testing, blue-green testing. Yeah.

Speaker: This just makes it efficient with those gates in place. That's what it sounds like, right? You pre-define those success criteria or metrics and only it meets those criteria, a human doesn't have to intervene, right? Exactly. Rollouts can can take it over. Okay. no yeah Thank you so much for that explanation and and description. i I'm sure it will help a lot of our listeners as well.

Speaker: um But let's go back to your time at Cruise. right Can you talk about ah the entire SDLC, CI, CD and the progressive delivery components? um What was the tool of choice? I know like Argo wasn't used at that point. What was the tool and then how did we can focus more on like how do to migrate? Yeah, yeah for sure. for sure so we um So before we introduced Argo, we had this tool called BuildKite. So BuildKite is a great CI tool. We were using it for CI as well as CD. um It can be used as a CD tool as well, and it works out great in certain scenarios. ah But it's just that the way we had it set up at and the company at the time, ah it wasn't, you know, ah we we ran into like a lot of different issues with it.

Speaker: um So to explain the setup, um so... um each of those services would get their own build guide pipelines. and that Within the pipeline, like developers would define like those steps that you know have to happen after a code is committed, like merge to the main branch. um know there is There's a series of steps that need to happen. right like The code needs to be built, the container image needs to be built and uploaded to you know some registry and like some testing to happen. and Once you're satisfied with the quality of all that, like ah the pipeline needs to you know deploy to the dev environment environment and then maybe a staging and then to you know production.

Speaker: um and so For this whole lifecycle, we used um Buildkite. and um One of the ah main issues that that we were facing was that, ah so with BuildKite, we use a push model where, um so the BuildKite is the CD tool and once you know an update is available, it pushes the new version of the code to the compute instance, which is the Kubernetes cluster here.

Speaker: um So it was using this push model and um so with the push model essentially continues delivery tool which was built right at the time had to have credentials to the Kubernetes cluster ah because without credentials it can't do the deployment.

Speaker: and So this is where it kind of, you know um so the way we handled this authentication was ah Maybe not the best way to do it, but what what would happen was um when a new service was created and when a new Kubernetes cluster or namespace was created, we'd first have have to have ah some sort of a kubeconfig which had like the service account details ah that would allow the external system to you know start the Helm upgrade command.

Speaker: um So create this kube.config file and this was like this would have long lived credentials. And so these weren't short term credentials, but these were permanent credentials that were then stored in Vault. And yeah, I mean, although we didn't store it very securely in Vault and we didn't give access and all of that, but that's generally not a good idea to have longer credentials. So that was one piece of it.

Speaker: The other piece of it was, ah ah now we have like a lot of different authorizations or permissions at play here. so First off, ah the BuildKite pipeline that itself should be able to read from Vault.

Speaker: and um Also, we had like ah different Vault paths. we would store the long lived credentials that was specific for that service and as well as like ah Kubernetes namespace. so Each service could be deployed across like multiple server Kubernetes clusters or even namespaces.

Speaker: um so but We had a combination of vault paths that were specific to you know the different servers, namespace, and repository. and so We had to store it in the right structure. and We then had to give permissions to the BuildKite pipeline to you know read from this. and and um so Within the BuildKite pipeline itself, like developers had to you know configure all this code to first of read from vault, get the kubeconfig file, and then you know log into GCP, then log into the cluster, and then start off start the deployment.

Speaker: ah So there was a sequence of steps that we expected developers to get right and a lot of times you know they wouldn't. And um so they would you know see the random permission issues like either the BuildCut type 10 could not read from Vault or they weren't reading the right path or Even if they're at the right path, they weren't trying to you know deploy to the right cluster or namespace or they would get the GCP permissions wrong. It was a source of like a lot of um you know secondary tickets that they got ah ah because it was so confusing to you know set up a new pipeline or set up a new stage in in in the cluster, like a deployment ah stage in their servers. It would take a lot of time.

Speaker: And it was super frustrating for the developers who used the tool as well as the platform team as well, because the secondary on call of the platform team now had to go, you know, debug like lot of different deployment failures and figure out, you know, what's going on. So it was a source of like a lot of OE issues for the team in general.

Speaker: yeah so that was I would guess so, right? like i think And it's not just for one specific team. If you were doing this across different teams, oh because of all of these issues, the platform team, basically your team would have had to build up expertise into their code, not code basis, but like in their modes of operations, figure out how to do it usually. And yeah these are the common pitfalls that that they run into any time. So yeah, you but you your team would have had to build that ah knowledge base ah across the field instead of using a smarter tool. Yeah, yeah, exactly. Exactly.

Speaker: um Probably in the current landscape with, you know, AI and cloud code and all that, maybe some of things would have definitely been simplified. But yeah, back then, yeah, like ah the on-call engineer would have to go look at, you know, a lot of different pipelines, understand what's going on, which was, yeah, ah super frustrating. Unless you are getting measured on the number of tickets you resolve, then this is like the dream scenario, right? Yeah. I sold like 100 tickets this week because the tool sucks but no that's not about it that that was another annoying thing ah the other piece of it was um

Speaker: so ah Buildkite essentially would start off the deployment. It would run you know Helm upgrade install command. and um Once it started the command, it didn't really have visibility into you know what what was going on on the cluster.

Speaker: So um you know like sometimes pods might not come up, there could be issues with you know the new pods, like it might not read be able to read the container image or like pods could crash loop. There are like so many different errors. right and um So ah we didn't really have the possibility on the BuildKite pipeline. um So one of the reasons was that we would only give permissions to you know run Helm upgrade, but not any of the other kubectl permissions.

Speaker: um and But even if you know we gave it, like ah service developers would have to manually run QCTL commands from the builtite pipeline to figure out what was going on, it was super tedious. and um so there is no visibility into like and There's no good UI where you can look at, hey, I want likes say 10 pods to come up, but maybe you know five of them have come up and the rest of them are running into some some sort of an issue. um so That visibility wasn't there. um so that That was another you know frustrating piece of um like an experience for not not just the service developers, but the infrastructure team as well. Because you know again, we would have to go look at what's going on in each of these pipelines. and After having dealt with all the permission issues at the initial stage, like these sorts of issues would you know kind of keep happening. so um

Speaker: ah so We spent a lot of time you know debugging these sorts of issues as well. And ah so one thing that we realized was that, yeah, Bitkite, didn't really give like good visibility and that that was the source of you know a lot of ops ah you know burden for the team.

Speaker: um Yeah, so that that was when we kind of realized that we probably need a different sort of different tool that you know maybe handles CI separately and CD separately. And yeah, that was when we started looking into alternative tools to you know handle our CI, CD yeah lifecycle.

Speaker: So again, just to base this discussion in some sort of timeline, right? Like when was this? Like, was it early 2018? This was like 2022-ish, 2022. Yeah, 2022. 2022. Okay. That's when when you were at Cruise, you you guys like, okay, enough of this.

Speaker: ah This tool, which was meant for CI, is not really working across the board for everything that we would expect it to do so let's go and evaluate other tools in the ecosystem so um yeah can you walk us through that bake-off that that you guys conducted ah what were the tools that you evaluated and what were some of the criteria like I'm sure you guys would have like a success criteria before starting the investigation Yeah, yeah, yeah. So we ah kind of evaluated like a lot of different tools, ah but the top three kind contenders that we had were like Argo CD, Flux CD, and Spinnaker. Okay. And... um

Speaker: These were like you know ah these were like you know really great tools and we really liked them. and We did like you know have a thorough design doc where we you know compared to you know tools. with you know We had a list of you know criteria in our mind and we did evaluate these tools against ah the criteria that we had.

Speaker: um The first criteria was ease of use. um you know We wanted the tool to be as easy as possible to use, not not just for the service developers, but for the platform team as well to be able to you know set it up and um you know just get things going.

Speaker: In that sense, Argo CD and Flux CD were really nice. like It was super easy to you know set up, super easy to learn. like The learning curve was like you know very short. um ah With Spinnaker, we kind of felt that um it was probably not that easy to set up and the learning curve was definitely higher. There were a lot of different you know microservices that had to be set up and so this whole you know suite of services that had to be set up and maintained and lot of concepts that had to be you know learned. And um so in in from in terms of ease of use, we felt maybe you know Argo CD and Flux CD are probably you know easier when compared to Spinnaker.

Speaker: yeah um So the other you know criteria that we had was... um So ah we ah like our team was like know heavily invested into Kubernetes.

Speaker: ah you know There were like ah if there were like many Kubernetes enthusiasts on the team who were super passionate about it. And um so ah Kubernetes was the only supported you know compute infrastructure at the company. And so we did a lot of things in a Kubernetes native way. like you know We had lots of Kubernetes operators, you know controllers for different know parts of the software development We even used it for infrastructure as code. And so we had heavily invested in Kubernetes. And so we're kind of a little biased towards Kubernetes native tools.

Speaker: ah so in that point I host co-host a podcast on Kubernetes by the time. Yeah, same here. So um in that sense, ah like Argo and Flux were Kubernetes native. And... um Spinnaker wasn't. so Spinnaker is um it's a multi-cloud, multi-environment enterprise CICT tool. um so You can use it to you know deploy to not just Kubernetes, but even outside of it. right like If you have like a VM where you're running your service or maybe AWS Lambda or like some serverless you know service, so that you can use it to deploy your code to any of those as well. It's super flexible ah that way.

Speaker: But it's just that we didn't have the need to deploy to you know any of these other environments outside of Kubernetes. And we also didn't have the need to deploy to any cloud providers outside of GCP. So we were a GCP shop and everything was in GCP. So we didn't really have to deploy you know to AWS or Azure. So um in that sense, ah there were like...

Speaker: And Spinnaker like you know had a lot of extra features that we didn't really know need at the time. so we favored Argo CD and Flux CD in the Kubernetes native you know aspect of it.

Speaker: So the other thing was the push versus the pull model. ah So with Belkite, I mean, we obviously wanted to solve the issues that we were seeing with using Belkite as a CD. And so one of the big problems was that Belkite had a push model and because of that, it had to have like the cluster permissions and um so We did consider that amongst the different tools that we are evaluating. ah With Argo CD and Flux CD, it used a pull model.

Speaker: um so The way it worked was, you have the controller running within the cluster. so Instead of like having you know an external system push your code into the cluster, ah the controller essentially runs on the cluster. and up When there is a change, it keeps checking the kit repository for changes and whenever there's a change, it automatically pulls in the change. So you don't really need to provide your cluster permissions to an external CI, CD tool. It lives within like the CD tool essentially lives inside the cluster. um

Speaker: So that is the pull model and um Argo and Flux provided this pull model. But Spinnaker again, didn't have the pull model. ah you know but yeah where the permissions to a Spinnaker so that that was a minor disadvantage for Spinnaker um Okay, so you listed three things, right? Like one was um like how how easy it is to learn or skill up on because you had to make make sure everybody in the organization is comfortable using the tool.

Speaker: ah There was a bias or preference towards Kubernetes native, native being Kubernetes native. And then the third one was pull versus push. Definitely that sounds like an interesting thing. Were there any major, ah any other major requirements? um Was cost or anything else? ah Yeah. So, I mean, there were a few other criteria that we had and most of like all the three systems did check those boxes. But I think one, ah when it came to progressive rollout, we did have some strong, you know, preferences. And so that was where we felt Argo CD and Flux. so Flux has like Flagger as the progressive rollout mechanism. And we felt that, you know, Argo CD and Flagger were,

Speaker: a little more superior when compared to Spinnaker, especially when it comes to, you know, the Kubernetes aspects of it. We felt that it was probably easier to, you know, like Argo and like Flagger had like advanced deployment progressive rule of strategies. um And some of it, we could still do it with Spinnaker, but it was just a little harder.

Speaker: And yeah, so in in that sense, ah yeah, like we, ah like we favored Argo and Flagger over Spinnaker. And yeah, this is the cost of you know maintaining Spinnaker over time. We kind of felt might be a little higher. um So yeah, we yeah we kind of then we narrowed in but on Argo and Flagger.

Speaker: ah The two were super close. In terms of features that they offer, like you know it's it's very similar. um and There were like you know some minor differences like Argo you know CD has a great UI and Argo Rollout has a great UI as well. But with Flagger, there's this new core UI that it offers and you have to use third-party UIs. And so there were a few minor differences like that. And ultimately, we ended up choosing Argo CD and Argo Rollout for our, you know,

Speaker: noiness department that also I think that's a great way to lay out the entire journey, right? Like what was important and and how did you evaluate all of these things? But then i think the next question is around like getting this inside your organization and convincing both the leadership and your the engineering teams or dev teams that you were supporting, right? To switch up such a core piece of their delivery infrastructure, right? Like it's not something that you do every week. So can you talk about like, how did you go through that ah whole transition starting with like, how did you convince your leadership team ah to to switch up this tool?

Speaker: Yeah. yeah yeah um so First of it was like a team effort. um so you know like A lot of people played like you know a key role in convincing the leadership to do this.

Speaker: um like The primary data point that we used was like the ops concern, right like the OE burden that we were facing. so you know It was like not just a burden for the platform team, but for the individual service teams as well. and There were like a few strong voices on the other side on the service you know ah teams as well side where you know they were um didn't really like the current system setup with BuildKite and you know they had ah concerns so you So we did have, um you know, um a lot of people who wanted a change. And um so that definitely helped a lot.

Speaker: And um yeah, so we kind of wrote up a doc, like, you know, laid out the pros and cons of the current approach versus like, you know, like um any future alternatives. And yeah, Also, the company was in a bit of a growth phase at that point of time and you know a lot of different services were being built. and so We felt the need to have like a better you know tooling in place that would cause less friction and know help um accelerate our development. and um

Speaker: so that That helped as well. so Given that like you know the current setup was ah causing a bit of hindrance, um ah that provided you know some data points to help convince the leadership that you know the current model wasn't working out and we wanted to do something different.

Speaker: and um ah there were There were a lot of questions about the different options that we had chosen, and um so we had to you know ah convince them, but it actually wasn't you know really hard to convince the leadership about this change. It's more about like you know getting the individual teams to then switch over.

Speaker: I think that that was the hardest part because it's it's not like a ah a single day of... like you know like It's not like you spend like a day of effort or maybe a week of effort and you switch over, right? um So, yeah. um How long was that transition, though? like ah did Did you guys have it time-boxed into like a quarter, a couple of quarters? And how long did it actually take?

Speaker: So um we we took a quarter to you know first off build out the tooling. And um so after that was done, like um we we didn't have nio set target in mind that, you know, hey, we have to move everything off of like BuildKite by, yeah say, like, you know, in in like a quarter or six months or like and something like that. Also, ah we didn't re completely, you know, get rid of like BuildKite because we still continue to use BuildKite for the CI aspect of it. It was only the CD aspect that, you know, we wanted to switch over to Argo CD. And, um,

Speaker: So we didn't really have like a set target in mind. And the strategy that we followed was we essentially targeted like there were a lot of new services that were being built at the time. And so we first have made it mandatory for any new service that was created to use, you know, this new framework that we had built.

Speaker: um So because of that, like, ah yeah. A lot of teams had to you know start using Argo CD. And also there were like a few other teams that were like you know super frustrated with the current setup. So they were motivated enough to adopt this new framework that we had built.

Speaker: yeah And there were some teams where you know they built new services and they had this new framework, but they had older services with the older you know setup. And up so they were in this you know mixed state ah where they had to you know kind of maintain two different sorts of pipelines. And so those teams were incentivized to you know move to the new platform because they didn't like the setup that they had, like they didn't like the mixed state and they wanted to you know completely move over to the new system which was supported, you know which had like a lot of great features and was supported in a better way.

Speaker: um So yeah, it it was easier to get those sorts of service teams to move over and it it took time. right it Again, it wasn't like single quarter or something like that. It took some time to And i I don't think like, so there are still services that, you know, haven't completely moved over. So we have the long tail of, you know, customers, right, who don't really want to spend the time. So there are using mainframes today. So there is every week of probably definitely has a long tail associated with it. So no, but it's great, right, that ah you at least got all the newer services

Speaker: or service teams or newer teams that were building something on the new new CD platform or the new rollout platform. And then anybody who's frustrated um to show success, right? Like with any of these transitions, it's important to have like a pilot team or something to to show that, yes, what we claimed in the proposal is actually true. We do see benefits at a smaller scale. And then if more and more teams join that, that scale can expand. So um No, it's great. Like, okay. So um what else do we need to know about this transition? Like anything else that you would want to highlight?

Speaker: No, I mean, yeah, like we didn't really face any like major issues. We took a very cautious approach. And yeah, so the transition actually, you know, worked out great. Like we didn't, like I said, see any major issues. And we were like happy with it. And yeah, there were some, you know, gotchas, especially with when you have like Argo CD as well as Argo Rollout, that we kind of, you know,

Speaker: faced when we had Argo rollout. So we first of like started off with Argo CD, we introduced Argo CD and then we introduced Argo rollout. And there were some interesting scenarios where you know you have two different controllers like Argo CD and Argo rollout you know fighting against the same resource.

Speaker: So we had that that sort of a scenario, but nothing you know major. none i think that ah I don't recall any production incidents. Did you guys ever have to like contribute code back to the community to make any enhancements in the ecosystem or just happy with with whatever was coming out from the community?

Speaker: um we we were We were happy. like We didn't really have to contribute back. um There were some cases where, like especially with Argo rollout, we wished we you know those ah those features were supported. But um i mean even even today, like some of the features are not supported natively, like multi-cluster for progressive rollout or like progressive rollout for like ah stateful sort of a workload. um um Yeah, so those those features are um not very well supported with Argo rollout.

Speaker: Gotcha. Okay, no, I didn't know um stateful workloads are not supported. Come on, Argo. like We need to do that. ah So basically, in that A-B testing or blue-green deployment scenario, right like the PVC can still be the same. like That's what it sounds like if you have read-read many. You can have multiple versions of the app ah still accessing the same data. So that is that good enough? Or are are you thinking something extra when it comes to Argo rollouts?

Speaker: inka Yeah, yeah. Yeah, yeah, definitely. So ah with Argo Rollout, it's much more than just the blueg greenen or like you know the just blue-green deployment. right yeah um So maybe just taking a step back to understand like ah you know the basic Kubernetes deployment and it would help to understand like the issues with the the basic deployment. um So that way people appreciate like what Argo Rollout offers. um so The basic Kubernetes deployment, right you have very simple rollout strategies. You have like rolling update and the recreate strategy. yeah ah With rolling update, essentially you know new pods are created, all pods are removed and it does so in a steady state until ah all the old pods are replaced with the new pods.

Speaker: And during when when a new pod comes up, ah like you know Kubernetes runs some basic health checks and if the health checks succeed, like you know it turns the success healthy and brings it up. yeah The problem with this approach is that um There could be a lot of cases where the pod itself looks healthy, but the state of the application might not be. right like ah You can't really evaluate like all the different corner cases that meant that the service would run into, especially with you know certain kinds of customer traffic you know would exercise certain code paths and wouldn't really be able to validate all of that with a basic health check.

Speaker: um So there could be cases where like you know um some of the APS could have 5X, etc. or the latency of the APS could like no be high and there could be different sorts of issues issues and we can't you know determine those issues with a basic Kubernetes deployment. and um and so Kubernetes doesn't automatically roll back when you know something happens, it just enters a crash loop back-off state and you don't get the automated rollback.

Speaker: so That is where like ah you know Argo rollout or any progress rollout helps a lot. um so With Argo rollout, you get the ability to ah look at application-specific metrics to determine like how healthy are the service is functioning.

Speaker: um and Based on that, you can decide to you know roll out or roll back. ah so That is you know one of the advantages of Argo rollout. The other thing is like with the the basic Kubernetes rollout, ah you can't really control the rate at which ah you know the deployment happens.

Speaker: I mean, you can configure like the the max surge and the max unavailable values, but beyond that, you can't you know say that, hey, I want maybe the roller to happen like very slowly initially and then wait for a while and maybe you know run some ah ah smoke tests to see how the you know the new version is and And once, you know, if they also look at the the customer ah metrics for for the service, see, you know, the business metrics and see how it's doing. And if you're satisfied, then gradually, you know, ramp up. And so keep doing this until we are much more confident with the service. And finally, maybe towards the end, you want to roll up aggressively. Right. So that sort of a fine grained control is not available with basic Kubernetes. Okay. climate

Speaker: And Argo rollout, right like can this progressive delivery thing be automated and manual? like I get the automated part. You already spoke about having those metrics in place beforehand.

Speaker: But if you are looking at, i don't know real-time information, and then you can you update that or override that with real-time traffic-shaping rules that hey, yes, this looks great and it meets my criteria, but I still am not happy with it and I do want to force more of my users to go back to the older version or a Accelerate yeah the newer version, either one of those cases.

Speaker: Yeah, you can have manual gates in your deployments. You can have like manual pause steps to you know ah for a human to look at how things are and then automatically roll forward or roll back. so It does definitely support that. okay and yeah so It definitely supports that.

Speaker: and you know Additionally, like it has the traffic shaping like you mentioned. um so With the basic Kubernetes deployment, you don't really get traffic shaping during a rollout. right so ah The amount of traffic that a new version of the server gets during the rollout essentially depends on how many pods you have in the service. right so If you have, say, a service for 10 pods and you replace ah one pod with you know a new version, ah essentially each pod gets like one tenth of the traffic. so like When you replace a pod, like um the new version will start getting 10% of your traffic. But, um you know, sometimes you might want to start very slow, right? You might want to have, say, 1% of the traffic or like something like smaller than 10%. You can't do that with the basic Kubernetes deployment. That's where the canary, you know, deployment helps that Argo Roller provides.

Speaker: Okay, no that's awesome. Sorry, go ahead. ah Yeah. So yeah, Argo has like a lot of you know interesting features, not not just with the Gennady Rollout, but even with Blue Green deployments. It's both Blue Green deployments, so you can have like a completely you know independent version of your entire service, and you can run some pre-promotion analysis, you can run some smoke test, look at metrics, see how inner things are, and then switch over and then even run like some post analysi analysis ah as well um to make sure you know things are working fine. and If you're satisfied, you can let it succeed or you can even roll back. so ah yeah It has a lot of you know great features and it has like a good UI as well where people can look at the state of their deployment and

Speaker: So yeah, it has a lot of great features. And Pushkar, when we were talking about this on the intro call, you had spoken that it also integrates with the ingress controller. Can you talk about how that integration works as well, please?

Speaker: Yeah, yeah, yeah. So it integrates with like service mishas and ingress controllers. And um so you can essentially um you know um divert like a certain portion of your traffic to the canary. Okay.

Speaker: so that That is where it helps. right like um so You can add the ingress layer, like determine like what percentage of your traffic goes to the Canary versus the active, the stable service. so That way, you're not restricted to like the pod level granularity, but you get more fine-grained control on how much you know what portion of your traffic should go to the Canary. and You can even have header-based routing. so You can have like you know certain kind of traffic with like certain header be routed to the Canary. and the rest of it can be routed too. That's where the ingress and service mesh integration comes in. One last question on on the on rollout. outside um

Speaker: Who is configuring these rules? like in in With your time at Cruise, was it the the dev teams that were specifying rules and pushing them along with code for Argo to read, or was it the the platform team that was responsible for production and didn't want to push all of the code at some time. It was both. There was cluster level analysis templates that we had put in place, which looked at basic latency and error rate of the different APIs. So, developers didn't have to bother about adding these basic metrics.

Speaker: and But the service team responsible for like the business aspects of it or if they wanted more fine-grained metrics about how their individual APIs were doing, they had to configure ah those metrics. So it was a mix of both.

Speaker: Gotcha. Okay. no Awesome. Thank you so much for walking us through that. I think one question I selfishly have because you you do come from that ah dev background. like How do you use AI in your day-to-day workflow? um Can you talk about what that looks like, your tools of choice,

Speaker: Anything, like if our listeners can just pick one or two things from what your your day-to-day workflow, I think that would be helpful as well. Yeah, yeah, for sure. um Yeah, so we obviously, you know, use quite a lot of AI, like Cloud is my primary yeah tool of choice and I'd say like 99% of my code is written through Cloud. Yeah. So we use like, yeah. Yeah. Plan-driven development quite a lot nowadays. so Essentially, any change that we have to make with any substantial change that needs to maybe go through this plan-driven development where we write the plans and or RFCs or designs and now other reviewers review plans instead of code. They do review a code as well, but as long as the plan conforms to what we're expecting, um it's usually approved.

Speaker: So, planetary development is being used quite a lot. ah We have a lot of you know AI like ah skills and plugins that we use for like a lot of different things. So, obviouslyly yeah like the company in general has built a lot of plugins like ah you know in terms of ops, like deployment, like on-call. So, if um there is like an on-call bot which can ah do you know traing the first level of tragging and It's gotten like really good. It's gotten to a point where um um you know it's able to do the analysis, figure out like even when you know the on-call gets paged, it's able to do the analysis itself and maybe do 90% of the debugging and know pinpoint the exact issue in in a lot of different cases. and it's It's getting you know really good that way. And so Pushkar, for the plan-driven development, right are there any public resources that you would want ah that you would let me link to both from your current work experience or things that you referred to ramp up on this? like um This is definitely something new in the ecosystem right now. So any any pointers on one where people can get started?

Speaker: um this ah We kind of read through like some companies like tech blogs, and that was where we kind of you know learned more about this. So I'm sure if you you know read up some tech blogs, maybe Snap might have some ah tech blog as well about it. ah But you should be able to find material online about it Awesome. No, this this this has been awesome, man. A great, it has been a great conversation with you. ah yeah ah Coming to my last question, right? Like where can people find, find you on the social medias and and how can they learn more from you?

Speaker: um um Honestly, I'm not um very active online. um But yeah, I mean, I'm based out of Seattle, like the greater Seattle area. So I work in Bellevue. So, you know, um if anyone wants to meet up for you know coffee or something, i'm I'm happy to. And in the Bellevue area.

Speaker: Awesome. Cool. no Thank you so much, Pushkar. This has been an amazing conversation. Thank you. Walking us through everything and and teaching us something new about about Argo and Argo rollouts.

Speaker: No worries. Thank you. Thank you for having me on the show. And it's it's been a great experience for me as well. Thank you so much for you know inviting me and having me on the show. Okay. That was a great interview. Hope you guys liked it as well. It it is always great to hear from a practitioner who has gone through the trials and tribulations and want to discuss and is looking is looking to share their war stories. So I'm hoping you guys learned something new about Argo CD and Argo rollouts. I know I did about Argo rollouts for sure. um But with that, we do have similar episodes lined up with practitioners from our evolving ecosystem. So I would really appreciate if you share a link to this episode on that one Slack channel or that one text chain that that you you were talking about last time.

Speaker: And with that pitch, it brings us to the end of another episode. I'm Bhavan, and thank you for listening to the Kubernetes Bytes podcast.

Speaker: Thank you for listening to the Kubernetes Bytes podcast.

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Recommended