Zencastr
00:00:00
00:00:01
Speed1x
Format▸
Share
Embed
Report

What if Kubernetes Could Resume Instead of Restart? | Checkpoint / Restore Explained

Kubernetes Bytes
Kubernetes Bytes

388 plays · Oct 2, 2026

Transcript

Speaker: You are listening to Kubernetes Bytes, a podcast bringing you the latest from the world of cloud native data management. My name is Ryan Wallner and I'm joined by Babin Shah coming to you from Boston, Massachusetts.

Speaker: We'll be sharing our thoughts on recent cloud native news and talking to industry experts about their experiences and challenges managing the wealth of data in today's cloud native ecosystem.

Speaker: Good morning, good afternoon and good evening wherever you are. We're coming to you from Boston, Massachusetts. Today is October 2nd, 2026. Hope everyone is doing well and staying safe.

Speaker: It has been a crazy couple of weeks ah with PTO and work travel for me, but I did not want to end our streak of getting good episodes out in front of you.

Speaker: So, keeping with the trend, we have another great episode lined up for you ah today, an interview with Dardustin Stoyanov, a PhD student at University of Oxford.

Speaker: We're going to talk about ah ah with about distributed model training and the open source project Creo, which helps Kubernetes users checkpoint and restore a a pod or a container running on Kubernetes to enable or to provide that infrastructure level checkpointing capabilities ah for the eventual goal of ah increasing GPU utilization across the board.

Speaker: So without further delay, let's get Radisteen on the board. Hey Radisteen, welcome to the Kubernetes Bytes podcast. Thank you so much for joining me today. Can you please take a moment and introduce yourself and talk about what you do?

Speaker: Yeah, thanks so much for inviting me. in So my my name is Raul Sinsvianov. I'm ie um so i um one of the maintainers of an open source project called Checkpoint Restore in User Space. And I am also leading in helping to lead the Checkpoint Restore working group in the Kubernetes community.

Speaker: um in as As a background about myself, I have been working on Checkpoint Restore for about 10 years now as part of my undergraduate studies and more recently as as part of my PhD in Oxford. And I've been just finishing that as well.

Speaker: So yeah, I think this is a short summary about myself. Yeah, no, thank you so much. And and i'm I'm excited for this conversation, right? Like this is ah purely from an open source project perspective, right? Like, hey, you're doing research and this is, we are talking about your research. It is not that we are trying to sell you a product or something in this episode for sure.

Speaker: ah Okay, so ah let me ah let me ask about like, for somebody who knows Kubernetes, right? Because if you look at our our listener base, All of them are from the Kubernetes background. Most of them, if not all. but ah So we are familiar with with with Kubernetes, but not with Checkpoint Restore. So can you ah help us understand like what actually happens when you Checkpoint or you perform a restore of a running container? Can you walk us through that basic workflow, workflow please?

Speaker: Yeah. um So I guess as as a starting point is, um ah there is a little bit of history how the Checkpoint restore functionality evolved and how it got you know relevant for the Kubernetes project. yeah But I guess it's an alternative of restarting workload. So for example, if you have a stateless application that doesn't have any state and it takes like ah one second to start like Using the restart functionality makes sense, but ah in many cases, like for example, with machine learning workloads, it takes several minutes to initialize a model for training workloads that run for long periods of time. If something fails, you can just restart and start from scratch, mainly because GPUs are very expensive. So essentially wasting a compute time becomes um you know a problem, ah especially if you scale to large clusters.

Speaker: So um these are some of the workloads that we have been investigating. But over the years, people have applied the Checkpoint Restore functionality, for example, to Java workloads ah to accelerate the startup time.

Speaker: And um the way the Checkpoint Restore Retrovers was introduced initially was as part of container runtimes, container engines. like For example, Docker was the first to adopt this and later Podman. And when we introduced um the first version, I guess, of the Checkpoint Restore support in Kubernetes, we also followed this model.

Speaker: Essentially, how do we check on containers and how do we them in Kubernetes? and um But essentially what it does is it saves the current state of the application, the execution, and then it allows you to resume from this point ah from the point in time when the snapshot was created. And the way this happens is very similar to how you can create a checkpoint of a virtual machine and then restore the virtual machine later. But this works at process level. So it, yeah.

Speaker: It works for containers. so so if like when i Again, I've been in this ecosystem for a while. right like Before this, whenever we wanted a snapshot or or or a checkpoint, we are are we wanted to move from ephemeral containers or stateless containers to stateful containers. like Those are PVCs that get mounted and data is written to them.

Speaker: Whenever we wanted to like capture the state of a board, we basically took ah took a snapshot of the PVC. but I believe the challenge that you're solving for is, it yes, that still might be needed and we'll talk more about it later. But ah the the problem here is whenever you're thinking about these ML workloads or AI workloads, both from a training perspective maybe even from an inferencing perspective, the startup times of these containers is too huge.

Speaker: So if they have to go and say get to this that same state, it takes a lot of time and that those are wasted GPU resources. Is that the case? And that's why we are... ah taking this ah container level or or we are checkpointing the container itself, container runtime itself.

Speaker: But where does this checkpoint get stored in Radistream? Yeah, I can. So we essentially, we have two different types of checkpointing. One is application level where, for example, you have PyTorch and PyTorch can create someot of the model state and then you can use this to restore. So this is still needed. um the I guess the way CRIU works, this we we call this infrastructure level checkpoint. So for example, if you're a Kubernetes administrator and you have some pods that are running um in the cluster. You don't know what exactly the application is doing inside of the containers, but you want to essentially stop these pods and, for example, resume them later. So this is sort of this called transparent checkpointing. You can essentially checkpoint the

Speaker: the containers and the ports without modifying the application anything inside of them. and For example, application level checkpoints, they they allow you for example to save the application state so that the application itself can use it later. um so You you sort of need both technologies um and they they essentially solve different problems as well.

Speaker: Okay, gotcha. Can you talk more about like what problem um infrastructure level checkpointing as i understand solves? Yeah. so I mean, one of the use cases that has become quite popular like over the past one, two years is essentially accelerating the cost starts of inference workloads.

Speaker: So for example, inference engines like VOM and HCLANG, especially for large module with large models like GMIK3 and G1F2, it takes like 10 minutes to initialize or 10, 15 minutes, it would be like up to half an hour, it depends on how it's configured. But essentially these workloads are very big.

Speaker: So they they will be running for example on eight B300 GPUs, which can have about terabyte of GPU memory and then have about two terabytes of cost memory. So essentially just in utilizing this for close, takes long time. And the way Checkpoint Restore works is once the model state is in the list, once the CUDA graphs, for example, have been computed, you can essentially save this state and next time you can just restore the state onto the GPU.

Speaker: Okay. So where are we saving the state? Is it in memory on the GPU or somewhere else? Yeah. So um the the way this works is when you um you have multiple phases during the checkpoint operation ah the first space is when you have GPU applications is to checkpoint the GPU state.

Speaker: And the way this works is ah we essentially move the GPU state into host memory and then we essentially save the CPU state. So this will include both CPU and GPU content and this is saved to persistent storage. So for example, to disk.

Speaker: Okay. Okay. Gotcha. And then like we like a couple of episodes back, right? I was talking to somebody from CastRI and we were talking about how these, I don't know. I know that discussion was more around model routers, but I've seen smart provisioning decisions taken by Kubernetes and and other open source projects that helps ah Kubernetes take orchestration decisions, right? Hey, orchestrate this specific pod on this node because that has the GPU that you need or that has a time slice of a GPU that that you have requested for.

Speaker: if I'm storing a snapshot state right on on on memory somewhere, or I know you said you can also push it down to disk. Does that help or does that influence the orchestration decision whenever that that new pod is coming up as well?

Speaker: um So I'll give you an example. one One of the projects that is integrat integrating this functionality is called Dynamo. It's an open source project from NVIDIA. And essentially the way it works is you you define how When you want to schedule a model, for example, you have a planner that will essentially allocate the GPUs, node resources, but then the snapshot functionality itself is just used to accelerate the code start time. okay And essentially you have something called snapshot job that would essentially create this artifact that will be used later when you want to start the model to start from a checkpoint instead of a code start.

Speaker: Okay. And then, like, I know, like, in in Kubernetes, the atomic unit is is a pod, right, not a container, like there can be multiple containers in a pod. So is this checkpointing functionality working or operating at the container level or even at the pod level, right, when when I have to capture state of multiple containers?

Speaker: Yes, so this is quite interesting. so when we started to and we I've been working with several people in the community and Adrian Reber has been one of the main people starting the initial work on Checkpoint Restore and Kubernetes and proposing the changes. And the way the way we kind of we talked with Adrian about this was um Essentially, we already have this container checkpoint functionality in Podman, which was essentially the starting point. And then we thought how we can introduce to this to Kubernetes. So we essentially follow this container checkpointing um mechanism. we We try to introduce Checkpoint Resolve functionality in ContainerD Cryo that that can be used in Kubernetes. And then at Kubernetes, we essentially added an API that allows you to trigger the container checkpoint essential mechanism. But um we essentially had a few talks. And the feedback we we got from the community was that um essentially in Kubernetes, the ah smallest of

Speaker: unit of deployment is bold. So yeahp you can't really, um you know, it doesn't make a lot of sense for Kubernetes to check but individual containers, maybe because um you can have containers that share, for example, namespaces like namespaces and ah you have to also ah consider, for example, shared volumes. And yeah um so they we once we receive this feedback and We also discussed with a few other people. We proposed the creation of the Checkpoint Restore working group. And the first thing that we outlined as essentially as um main goal that we are working towards is introducing this pod level Checkpoint Restore mechanism that essentially replaces the current container level Checkpoint Restore in Kubernetes. And this is the building block for future work as well.

Speaker: And is this something that's already in alpha or in active development? Like, has it reached that milestone yet? Yeah. So ill we we started the working group in December last year. And over the past few months, we essentially went through a few iterations of the design, how pod level checkpoint restore should look like, what APIs we need, and what is the best interface for users. And we introduced, um I guess, there is some when When we want to introduce new features in Kubernetes, we have this problem where we first need the functionality to be implemented in the container runtime. But before it's implemented in the container runtime, we need it in the Kubernetes. So essentially, we have this challenge. And the way we solve this is by first introducing the minimal set of APIs at the container runtime interface, essentially at the CRI. level. And this is what we did in 1.37 in the alpha version.

Speaker: And now the next step is to introduce essentially a bot level checkpoint restore in container D and cryo. And once this is available in the container, ramset and we can essentially introduce the um the remaining APIs in Kubernetes.

Speaker: There's like a few phases and steps to to make sure that this is done correctly. Yeah, I didn't know, honestly, that there was that chicken and egg problem that before Kubernetes, it has to be in the runtime. And for runtime, it has to be in the Kubernetes. That's an interesting workflow for sure.

Speaker: thank you for working through it. And like I think features or changes that that that have this huge of an impact do need to be rolled out in phases and it makes sense to get feedback from the community which and ah you are already doing. So that that's great news. So today, right if I wanted to use the Checkpoint restore capabilities and functionalities, it does work at the container level. How do I set it up? right like Can I deploy it on a Kubernetes cluster and and use it for my containers that I'm already running or do I need to install? How do I get this up and running? right How do I get access to it?

Speaker: Yeah, I mean, it depends on the ah the exact use case. For example, with um inference workloads, the animal project, the way they do is essentially they implement their own snapshotting mechanism that is using, for example, it's it's creating new CRDs that will essentially create a snapshot of of the workload.

Speaker: um And we are trying to introduce a core API mechanism that will allow them to essentially do the same thing, but without all older i guess um different things that they have to do outside Kubernetes to make this work.

Speaker: um I guess what we currently have been it was introduced as a better version in one thirty and Kubernetes 1.30 is this container level checkpoint functionality.

Speaker: yu But hopefully we can introduce the bot level checkpoint restore functionality in 1.38. There is a more detailed example on the Creeo.org website.

Speaker: So it's Creeo.org slash Kubernetes. It has um essentially more information on how it works, what is the API, and how to use it. Gotcha. Okay. Now we'll definitely include that in the show notes for for our listeners to to check it out, right? Okay. um So let's let's talk about like distributed applications. right So is that the next evolution that you think about or strategize about? like we We started with containers, now we are moving to pods. How about distributed applications? can Can I orchestrate a checkpoint across multiple containers or multiple pods? How does that work? Is there a component in the project that that helps me do that as well?

Speaker: Yeah. so um isn' Distributed applications have been one of the i guess main use cases that we considered when we started working on Kubernetes. are This is also the main difference between you know using Docker and Podman versus using Kubernetes. like Most of the applications deployed in Kubernetes clusters are usually distributed running across multiple nodes, multiple pods.

Speaker: um so the way i guess In terms of research, and there has been huge amount of research in in the HPC community, high-performance computing, for creating snapshots of distributed applications. um But we essentially have to find a way to make this work with Creo. So it doesn't have native support for um essentially distributed applications. And the way we are trying to solve this problem is with another project called CRIO coordinator. So this is um essentially works as an action script in CRIO. operation at different phases and allows to synchronize across multiple workloads and across multiple nodes.

Speaker: um But yeah, this is still work in progress. um So the idea is once we have the initial both-level check-out restore functionality, we can essentially extend this to distribute workloads. And in addition to handling the cpu <unk> synchronization across CPU workloads, we also need to consider, for example, NICO and different types of GPU communication.

Speaker: um Then there is, um enough I guess, work happening in parallel from the NVIDIA team working on this. so they I think they're introducing currently a support in Nico to do essentially check-out restore of more GTP workloads.

Speaker: ah This initial version of this work was released, I think about a month ago. I think someone sometime in July, August, and i currently the NVIDIA team are working on this to production state where we can essentially support large language models and inference workloads.

Speaker: And this works for with any of the inference projects like VLLM or like and yeah like is it compact? Okay. So we are actually like ah the VLLM community actually are has an RFC to integrate native support for Checkpoint Restore, but we have a few people from the Dynamo team that that are helping us and also a few people from DLM and you know um other projects like HGLN, for example, that i'll keep track of dysfunctionality. Essentially, the checkpoint restore mechanism is very helpful for reducing the call start times.

Speaker: Gotcha. Okay. And ah like if we talk about, again, the the point that you made a while back, right like ah the difference between application-level checkpointing and then infrastructure-level checkpointing,

Speaker: Do you see, like once we have pod level support in 138, do you see both of these approaches needed or I can pick and choose which one I want to use? Like how do, if I am a user, right? if of How do I think about these things together in conjunction?

Speaker: Yeah, I mean, to be honest, we discussed this in the working group, in the Checkpoint Restore working group, and I think most people lean towards deprecating the container-level Checkpoint Restore in favor of both-level Checkpoint Restore, only because they're very similar, and it in general, it will be easier to use only and to support only one functionality. um So the container-checkpoint-restore functionality is currently clean in Betas, and yeah essentially once we have the pod-level checkpoint restore functionality, merge, we can discuss this and start deprecating the pod.

Speaker: And then what about pod-level checkpointing with like PyTorch level model checkpointing, right? some of these frameworks, as you described, already have model checkpointing and those kinds of features available.

Speaker: Would we use them together or either or? um It depends on the use case. So for example, um when you When you have a training workload that will be running, let's say, for weeks or months, um you want the final result to be a PyTorch checkpoint. this is something This is essentially the artifact that you need to essentially distribute the model weights and to do inference.

Speaker: um The infinite infrastructure that our checkpoint restore is mainly for fault tolerance So for example, if something fails, you can recover from this state. And yeah, that's the main idea.

Speaker: Okay. so So both might be needed, right? Like yeah as you described. Okay. Okay. Got it. So actually you brought up a great point, right? Fault tolerance. Can you talk about like without the Creo project, what happens today when a GPU fails?

Speaker: ah What are the consequences to the training workloads or the inferencing workloads? And then obviously i want to follow up that follow that up with a question around after Creo, what would that look like? So can you talk about like what happens when a GPU fails, please?

Speaker: Yeah, so um I guess a good reference is Microsoft published a paper called Singularity in 2022. So they described how they essentially use Checkpoint Restore to solve this problem. And also they published another paper in 2024 called Just-in-Time Checkpointing that essentially describes how they implement error recovery more efficiently, again and with Checkpoint Restore. But the the main idea is that if you have a distributed workload, let's let's say running on thousands of GPUs and if one GPU fails and you don't really want to restart everything because this will be a huge amount of waste of your compute time. um So essentially they they're trying to figure out how to um essentially preserve the state. ah For example, you have different ways of doing this but the most common approach is periodic snapshots. So you have periodic checkpoints being created so that if you have to restart, you can essentially restart from last checkpoint. In the just-in-time checkpointing paper, they introduced like a new approach to this, this just-in-time approach. The idea is that you can essentially use the healthy state of other workloads to recover the GPU state if, for example, one replica first.

Speaker: ah It's essentially reducing the amount of time it takes to i'll create checkpoints and the overhead of checkpointing. Okay, but so you you said redundancy at the application level. Do I need like multiple replicas of that training workload running across all of the GPUs or is it like parity? Like is it is it replica replicas as in, hey, there is an exact copy or is it like parity where even if one node goes down, I can recreate it based on everything else that that I have access to?

Speaker: Yeah, and so there are different types of parallelism, for example, for workloads. And one of those methods is called data parallel not training. So in this case, you you just need to replicate, the for example, the model weights across multiple GPUs to be able to run parallel computations. So if one GPU fails, you have the same model weights on another GPU. And you have Moji, essentially you have NICO communication that is more efficient to transfer the at the GPU state.

Speaker: Okay, so yeah, if i if I have a training job across, let's say 100 GPUs, one or two failing might not impact my training times. Like obviously it might make it extended by five, 10 minutes. I don't know, I'm just throwing out a number there. But I don't have to go through the whole process again and recover from the latest application level checkpoint that I might have taken. Right, okay.

Speaker: Gotcha. um So your your work also uses checkpointing for like the LLM inferencing, right? So can you talk about like how how that would work with when I'm swapping models, right? If going back to the model router decisions now, people have figured out that they can choose the right model for the right job.

Speaker: Does checkpoint research help me in in that scenario or or any other part of your research? does does this Does it help me with cold start problem around different models? Yeah, so um we last year we published one paper on hot-swapping. The main idea here is um if you have... Essentially, when you create a snapshot of inference workflow, the bottleneck that we identified is essentially saving and reading the data to disk.

Speaker: So the idea is that um you can use the um GPU Checkpoint Restore functionality to essentially move the GPU state from and the gpo and into host memory and then this allows you to switch between different models very quickly.

Speaker: um So you can avoid essentially stopping the inference workload and um essentially you can switch between different models by moving the GPU state into host memory and then moving it back. So it's kind of similar to using disk as a swap, swapping space, but in this case we use the host memory to swap between different GPU workloads.

Speaker: Okay, so do I need to then right-size my CPU-based Kubernetes worker nodes so that they have enough host memory? or Sorry, that's a really basic question, but I thought the the reason we deployed it on GPUs was obviously the the parallel processing, but there is more memory available on GPUs compared to CPUs. ah can Do i need now do i now need more CPU-based memory or CPU workload-based memory for this to work? Yes.

Speaker: ah So in in general, um the way GPU Checkpoint Restore works is we want to ah essentially we um we need to increase, for example, the um the memory limits on the pods so that they can also fit the GPU snapshots.

Speaker: So um because this happens at low level, the GPU driver doesn't know what are the Kubernetes resource limits. So it it has to be configured by either the user or, for example, the inference framework that is deploying the models.

Speaker: But um yeah, essentially this is more like, um um yeah, one of the things that we have to consider. Okay, and as part of this paper that you published, right,

Speaker: did you do any benchmarking, right? Like the the having these warm models, quote unquote, ready on that in the host memory and bringing them back into active state.

Speaker: How much time savings are we talking about, right? Especially with the newer, larger models, um can you do you have any reference points? Like, hey, you can save up to X amount of time or make it X amount faster. Like, was there an outcome that that you were able to find through your work?

Speaker: Yeah, I mean, um essentially what you can achieve is you can reduce the initialization time of of inference workloads from several minutes, like for example, from 10 minutes to a few seconds.

Speaker: So you can get a huge benefit from using this snapshot functionality. But there's like more up-to-date information on d ten and the project themselves itself has a set of benchmarks and benchmark results. So essentially, you can find a lot of information on that. OK, OK, perfect. no we will if you I don't know you already shared that information. If not, send it over and we'll include it so that people can follow up on it.

Speaker: OK, perfect. um So once we have like moving on to the next question, right like once we are checkpointing these AI workloads, ah these snapshots from the discussion, it seems like they themselves can be enormous.

Speaker: um how does the newer on-the-fly memory compression work um address this problem, right? Like, can can we do any sort of compression to make sure that ah ah this is more efficient? And then what are the trade-offs, if you can talk a bit more about that?

Speaker: Yeah, so... um ah We have people in the CREO community asking for compression support for a long time. And essentially, the the first use case we had was Java applications, mainly because the Java runtime has essentially garbage collection mechanism that when you checkpoint the application, there could be many empty pages. And essentially, um when you compress the GPU um the Java memory, we can get huge amount the huge benefits from but reducing the checkpoint size.

Speaker: um there There is a project actually that integrates the Creo checkpoint restore mechanism with the Java runtime and introduces um like a few phases that will, for example, release the memory. But this is this was essentially the starting point. They implemented an initial version of the compression mechanism.

Speaker: But the way they did that was essentially it was very naive. So it would just save all the memory pages, and then it would read back the data to compress it and then save the compressed files. And then during Crystal, it would decompress them in a temporary file system and then read it back. Essentially, it wasn't very efficient. So in the Crypt community, we discussed a few different approaches and we took as an example the Z-swap mechanism in the Linux kernel.

Speaker: and Essentially, the idea is that you can compress individual memory pages. um And we discussed this at the Linux kernel last year in December. And one of the optimization mechanisms that essentially was proposed was to um essentially be smart about what data is, doesn't compress very well. For example, the model weights of an inference workup, they could look like random data. And this um this data um

Speaker: There is no need to compress to compress this data because it wouldn't give us the benefits that we want. So what we did is we essentially distinguished between empty pages, between pages that compress well and memory pages that don't compress very well.

Speaker: and We essentially, during checkpointing, we can identify these are three types of memory. And then we essentially keep only the essentially the memory pages that compress well. And we don't save any data for empty pages. We just mark them as empty. And then ah pages that don't compress very well, we just save them as the raw data so that we can essentially restore them more efficiently without the overhead of decompression.

Speaker: And all of the this analysis, right, like to identify these three types of pages and actually save that state, is that done on an ongoing basis? that like are Are you keeping track of all the different memory pages or is this happening whenever a snapshot is triggered or or whenever Trio takes a snapshot?

Speaker: Yes, so um we implemented this compression mechanism within CREU itself, within the checkpointing pipeline. So it happens when the checkpoint is created. But in addition to that, we also um we we noticed when running evaluation experiments, we noticed that because we decompress individual pages and memory regions, what we can do is we can use parallelism. so we can During restore, we can have multiple threads ranked in parallel to essentially accelerate the amount of time it takes to restore the memory. This is essentially the main, I guess, the most time-consuming phase during the restore operation. So for example, if you have an inference workload and you want to reduce the cost of time, using this um decompression parallelism, it allows you to reduce the amount of time it takes to restore, in addition to reducing the amount of data that you need to read on disk.

Speaker: OK. And so one basic question, right? like Let's assume we are in the 138 release. Who is orchestrating all of these capabilities, right? I understand that the actual checkpoint is is done through the runtime.

Speaker: But who tells that runtime? like Will we have additional pods or stateful sets or demon sets running on the Kubernetes cluster that go and and like somewhere in a YAML specification, somebody specifies that, hey, this is the period at which I want to take a a checkpoint.

Speaker: And then it goes and triggers those. like how does Who is orchestrating all of these things? Yeah, so we discussed this actually. um This is more like a design and architecture question that you know we have there are different opinions. And definite of the best solution is, um you know it's I guess, longer discussion that we had. So what we currently ended up as a design choice is to introduce a controller within essentially within the core Kubernetes. Oh, wow. Okay.

Speaker: That essentially allows you to create snapshots for pods. And then ah the restore functionality is very similar to pod creation. so it's essentially, the it's exactly the same way you would create a normal pod. But the only thing is that you would specify the the snapshot that you want to use to... essentially restore the runtime state.

Speaker: So in terms of Kubernetes, like from kubernetes Kubernetes perspective, everything else looks like creating a normal port. But the only difference comes to the container runtime. The container runtime, instead of starting container, it will restore the container from a checkpoint.

Speaker: Gotcha. OK. And that's awesome, right? Like it if it's part of the Kubernetes controller itself rather than like a second set of oh components and controllers that we have to install on the Kubernetes cluster. Yeah, this was actually something we spent a lot of time debating because if it was an external controller, it would make it easier for us to you know to iterate. Essentially, there were some arguments and essentially the conclusion was that being part of the core Kubernetes API is the right decision.

Speaker: Gotcha. know And then that also helps with adoption and then usage. right like If it's an external controller, yeah the pace of those innovations might be faster, but then if it's not being used, i think that's a definitely a ah a trade-off decision right there. One thing that I guess we also considered during the working group discussions was that in addition to the preview project, which runs with ah which can be used with RunC and CROM, the OCI container runtimes,

Speaker: um And we also Geviser and CAD containers. So both Geviser and CAD containers have similar snapshot to the Checkpoint Restore functionality. And so when we introduced the Kubernetes API, we want to make sure that it works with Preview, with Geviser, with CAD containers. And so we are defining more like a standard API for Checkpoint Restore in Kubernetes.

Speaker: Gotcha. Okay. And like, is any of this work also applicable to, I know with HPC workloads, like Slurm is the orchestrator, right? Like is Checkpoint Restore a capability that already exists in other orchestration frameworks or this is being built in Kubernetes right now and then there might be similar other projects for other orchestrators?

Speaker: um Well, Checkpoint Restore functionality for hpc workloads has existed for a long time okay okay lia There are many different frameworks that already handle this, but in Kubernetes is something relatively new that we're introducing.

Speaker: um and yeah it' Essentially getting feedback from different people in the community and just making sure that when we introduce something, it's something that people will be able to use in the future is what we're trying to do.

Speaker: Okay. okay oh Good to know, right? Like I didn't know HPC already did this. Again, that's ignorance on my part to just ignore the HPC side of the house. But okay, let's let's give we come back to Kubernetes, right? and And talk about checkpoints. Are there any security trade-offs that we need to be thinking about or aware of, right? Like if, does it, do we have to worry about the checkpoint including things like containers, ah sorry, credentials, keys, or any other sensitive data that might be available in in raw memory now?

Speaker: um And like, Are there any implications security implications to to taking checkpoints? Yeah, so we we had of many discussions on this topic. um I guess one question is, um once you create a checkpoint, once you create the snapshot of the running state of an application, the Snapshot also includes a sensitive data, so it includes a complete memory ah the complete content of the application, including secrets, passwords, user credentials, so everything. and we We need to make sure that, um for example, this is saved in in a way that is secure. and one One approach that we have been working towards is introducing

Speaker: um essentially encryption support within Creo. In the same way we introduced compression, it's just the compression was the initial state, at least each so that ah once the data is compressed, we can encrypt it. And um yeah, that's something we have been working on. um The other thing is, during restore, we need to make sure that essentially that um essentially attackers cannot be um able to compromise or like one thing that, for example, we have seen in the past is if if you have untrusted checkpoint and it specifies, for example, um

Speaker: a mount point on the host, we don't want essentially the checkpoint to be able to restore this mount point. So um in the the way this is handled today is in the container. So it does additional validation that essentially the checkpoint matches what the pod specification is asking for.

Speaker: okay Okay, interesting. okay I'm glad that you guys have already thought about this. okay let's Let's come back to to Kubernetes. right so As you had described the workflow of how things are are done today, right like without the Checkpoint Restore functionality, is if like Kubernetes deploys a pod or orchestrates a pod on a specific node, if for some reason there is a a node failure or something, at the application level the pod is basically restarted if on a different node that's available right so it's basically kill reschedule and restart framework or workflow um do you see that once we have let's say we are in the 140 release of kubernetes right do you see kubernetes eventually adopting this as the way to perform these um operations or are transfer nodes right like before it

Speaker: it reschedules, it actually does a checkpoint or before it kills, it takes a checkpoint, reschedules, and then instead of restarting, it's it's resuming from those snapshots. Are you having discussions to influence Kubernetes behavior as well?

Speaker: Yeah, so we had discussions with different SIGs and working groups, um special interest groups and in Kubernetes. So scheduling is one of the I guess, interesting use cases because in this Kubernetes scheduler has a functionality called preemption. So for example, the way this works is if you want to make space for a new workload, you can essentially preempt a running pod and this will just kill the pod today. you know yes that Using checkpoints, you can essentially move the current pod from one node to another. But it's also application specific. So this would make sense only if, for example, restarting the application is expensive.

Speaker: So if it is, for example, stateless web server like Node.js application, then just restarting it might be faster. faster um And also depends on whether you have, for example, if you create the checkpoint, how large this checkpoint will be. Like if it is, for example,

Speaker: um If it has huge amount of memory that needs to be saved to disk and you don't have enough disk space to create the checkpoint, then essentially it would make sense to you know to use the restart mechanism instead.

Speaker: There are things like that. But yes, um and the integration with the scheduler is one thing that we have been discussing in the the Kubernetes community. Gotcha. And I think that those use cases that you mentioned around preemption, right, like node evictions or or node drains, all of those things do make sense. um And maybe there is a fallback mechanism, right? Like, again, it can be a flag that the customer sets that says, hey, this is a resume application or a restart application. And then, ah obviously, if you're out of memory or out of this space, we can fall back on on the restart approach. So,

Speaker: that that's awesome right like that's good to know um i think this this brings me to my i don't know second last or last question but like uh looking uh three to five years ahead right like i know you're at oxford right now and so you're always thinking about the next thing not trying to ah sell what's on the truck basically uh where do you think this ultimately goes right does uh checkpoint restore remain like a specialized technology for just AI HPC kind of workloads or model training inferencing kind of workloads or this becomes that underlying primitive for of our communities workloads and and scheduling decisions?

Speaker: Yeah, I mean, to be honest with AI, main things are moving very quickly and it's kind of difficult to predict what is going to happen, but through but um one of the projects that is emerging at the moment is called a agent substrate. So okay it It's a project that allows you to run large amounts of the genetic workloads. And essentially, they use the snapshot, the chip and restore functionality to, um I guess, ah to accelerate the startup time of of essentially these sandboxes that agents are using.

Speaker: um And yeah, ah um this is something that will be very interesting. The idea is that you can um essentially reduce the startup time to a few milliseconds for essentially workloads in Kubernetes. um I guess, in for the next few years, at the integration with the core APIs with Kubernetes and extending essentially port-level checkpoints to other um objects like workloads, jobs, and this will be, um um I guess, it's on our roadmap we plan to do, but also supporting different types of storage, remote storage and

Speaker: um ah for example, S3 bucket in AWS and things like that. um Yeah, this is yeah this main the what we are currently working towards.

Speaker: And I agree, right? Like three to five years, especially in this day and age is a very... yeah People who are creating those roadmaps, they exist on paper, but things can drastically change two months down the line and all of that just goes to goes to scrap and then you have to start sat over. So I agree that yeah we we don't know where we'll be in five years according to some...

Speaker: researchers in in some of these larger companies, we might all be dead. So maybe we don't even have to worry about it in in five years. But okay, um you are already plugged into the Kubernetes community. You have a working group there. But for people that learned about this project for the first time and want to get involved, what do you recommend they should do? Like, should they go attend your working meeting calls? How do how do they reach out to you? And how do they get involved with this?

Speaker: Yeah, so we have weekly meetings in the working group where we essentially discuss um you know what the main things that we're working on and trying to get feedback on the next steps.

Speaker: But yeah, in in terms of, I guess, um it depends on the exact use case, but there are many projects that are already integrating the Checkment Restore Function Active. And sometimes, for example, using Podmon for containers, it has but container checkpoint command using this will be the easiest way to start. They have a good documentation on how this works.

Speaker: And then ah in the case of Dynamo, for example, they also have um documentation on how the snapshot mechanism works. um So the users don't necessarily need to understand how Cree works or what are all of different create options Active Face should be, ah hopefully in the future, should be easy as easy to use as possible.

Speaker: And they don't have to think about how it works. Gotcha. So you're not thinking about building like a Kubernetes, like similar to how Kelsey Hightower had done, like the Kubernetes the hard way. You're not thinking about model check pointing the hard way. You just, it it is it is something that will be built in as part of the Kubernetes stack. You can use it as an abstraction layer. You don't have to worry about knowing how it actually works under the covers. Is that the the place you're going to?

Speaker: I mean, I would try to write blog posts or like documentation pages that you know describe how everything works in detail. It's just how we want to make the essentially the are experience. Yeah.

Speaker: no makes sense um Are there any other things that you would want listeners to follow up on? How they can reach out to you? And if you are you doing talks at KubeCon North America you know in in a couple of months or month and a half? Can you talk about how people can get in touch with you as well, Larisne? Yeah, so we have a Slack channel in the Kubernetes community called Working Group Checkmate RESTOR. way to reach out. There are many people there that might also be able to answer questions. And the other way is on GitHub usually. So we we have GitHub issues and pull requests, and this is how we usually keep track of of everything. But yeah, Slack or GitHub is probably the best way. And I also have an email. People can just Google my name and find my email address.

Speaker: So yeah. Okay, perfect. No, thank you so much. And I know we We only scratched the surface on so many topics. I know we didn't even talk about Dynamo and the Nickel framework and all of those things. and We just referenced them. and I'll include all of those links in the show notes. But Radistin, thank you so much for your time today. And yeah, if if whenever we hit the 138 milestone or 140 milestone, whenever whenever this becomes part of Kubernetes, we'd love to have you back to talk about how this has evolved from from the discussion in September to a discussion when it actually goes GA.

Speaker: Yeah, thanks so much for inviting me. Thanks. Thank you so much for listening to the episode. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review or sharing it with your friends and colleagues, as that really helps us grow the podcast.

Speaker: You can find all past episodes or learn more about the show at kubernetesbytes.com. See you in the next episode.

Speaker: Thank you for listening to the Kubernetes Bytes podcast.

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Recommended