
What if Kubernetes Could Resume Instead of Restart? | Checkpoint / Restore Explained
Kubernetes usually handles failures and rescheduling by restarting workloads. But what if Kubernetes could save the exact running state of a workload, move it somewhere else, and resume from where it stopped? In this episode of Kubernetes Bytes, Bhavin sits down with Radostin Stoyanov, maintainer of CRIU (Checkpoint/Restore In Userspace) and a contributor helping lead the Kubernetes Checkpoint/Restore Working Group. They explore how checkpoint/restore works, why it is becoming increasingly relevant for Kubernetes and AI workloads, and how the community is working toward pod-level checkpoint and restore. The conversation covers how checkpointing can help reduce expensive AI inference cold starts, improve fault tolerance for large GPU training jobs, enable faster model swapping, and potentially influence future Kubernetes scheduling and preemption workflows. They also discuss the challenges of checkpointing distributed applications, GPU memory, compression, security, encryption, and integrating checkpoint/restore directly into Kubernetes APIs. Topics include: What checkpoint/restore actually saves Application-level vs. infrastructure-level checkpointing CRIU and Kubernetes Why AI inference workloads can take minutes to initialize Saving CPU and GPU execution state Pod-level vs. container-level checkpointing Kubernetes Checkpoint/Restore Working Group Distributed checkpoint/restore Fault tolerance for large GPU training jobs Accelerating model startup and model swapping Memory compression for large checkpoints Kubernetes APIs and controllers for checkpoint/restore Security risks of checkpoint files Encryption and validation Scheduler preemption and workload migration Checkpoint/restore for AI-agent sandboxes If you work with Kubernetes, GPUs, AI infrastructure, inference, distributed training, or scheduling, this episode provides a look at a Kubernetes capability that could become increasingly important as workloads become more expensive to restart. Subscribe to Kubernetes Bytes for conversations about Kubernetes, cloud native infrastructure, AI, platform engineering, security, and the CNCF ecosystem.



