01 · CPU and GPU checkpointing
Snapshot a fully warmed CPU and GPU container after model weights and kernels are loaded. New replicas restore from that snapshot instead of starting from scratch.
For vLLM, SGLang, and other GPU-heavy workloads, that turns multi-minute cold starts into seconds—and cuts wasted GPU time during scale-up.
Explore checkpointing →
|