“We were essentially able to reduce the cost of that cluster by about 75%. On AWS, DevZero demonstrated they could achieve significantly higher savings than we initially thought possible.”

Mihir Nair
Head of Architecture, Databahn
Checkpoint and restore inference workloads in seconds. Eliminate GPU idle time during model loading, survive spot preemptions, and migrate without downtime.
Every time an inference pod starts, it pulls model weights, loads them to VRAM, warms the engine, and initializes the KV cache. That process burns expensive GPU-hours while serving zero requests.
8-12 min
Average cold start for large models
$4.50/hr
Wasted per idle GPU during startup
60-90%
Of GPU time lost to initialization
0 requests
Served during cold start window
DevZero extends CRIU (Checkpoint/Restore In Userspace) with GPU-aware checkpointing that captures CUDA context, VRAM contents, and device state. We are core contributors and maintainers of the CRIU project, bringing novel improvements upstream.
Deploy on EKS, GKE, AKS, or bare metal. Checkpoints are stored in any S3-compatible object store. Migrate workloads across availability zones or even cloud providers.
Kubernetes-native CRD. Install via Helm chart. Checkpoint policies, storage backends, and migration triggers defined as custom resources. No application code changes.
Checkpoint and restore times measured on real models with full VRAM state, KV cache, and engine context. No shortcuts, no empty weights.
Benchmarks measured with vLLM serving on NVIDIA datacenter GPUs. Checkpoint stored to S3. Results vary by model architecture, GPU type, and network throughput.
| Feature | DevZero | NVIDIA Dynamo | GKE Pod Snapshots | Raw CRIU |
|---|---|---|---|---|
| GPU checkpoint/restore | ✅ | ✅ | ❌ | ❌ |
| Engine agnostic | ✅ | ❌ | N/A | N/A |
| Multi-GPU support | ✅ | ✅ | ❌ | ❌ |
| Cloud portable | ✅ | ❌ | ❌ | ✅ |
| Remote storage (S3/GCS) | ✅ | ❌ | ❌ | ❌ |
| Live migration | ✅ | ❌ | ❌ | ❌ |
| Spot preemption recovery | ✅ | ❌ | ✅ | ❌ |
| Kubernetes CRD | ✅ | ❌ | ✅ | ❌ |
| Compressed checkpoints | ✅ | ✅ | ❌ | ❌ |
Comparison based on publicly available documentation as of March 2026.
“We were essentially able to reduce the cost of that cluster by about 75%. On AWS, DevZero demonstrated they could achieve significantly higher savings than we initially thought possible.”

Mihir Nair
Head of Architecture, Databahn
See how checkpoint-restore can cut your inference startup time by 50× and reduce GPU costs.