Your GPUs are expensive. Don't waste them on cold starts.

Checkpoint and restore inference workloads in seconds. Eliminate GPU idle time during model loading, survive spot preemptions, and migrate without downtime.

GPU Cold Start Comparison
Traditional Cold Start
~8 min
Pull weights3.2m
Load VRAM2.1m
Warmup1.8m
KV init0.9m
GPU util:0%
Requests:0
Cost:$0.60 wasted
vs
DevZero Checkpoint Restore
~20s
~20s
Serving requests
GPU util:92%
Requests:1,247
Cost:$0.02
25× faster cold start

Inference cold starts are the silent GPU killer

Every time an inference pod starts, it pulls model weights, loads them to VRAM, warms the engine, and initializes the KV cache. That process burns expensive GPU-hours while serving zero requests.

8-12 min

Average cold start for large models

$4.50/hr

Wasted per idle GPU during startup

60-90%

Of GPU time lost to initialization

0 requests

Served during cold start window

Kubernetes-native. Engine-agnostic. Cloud-portable.

vLLM
SGLang
TRT-LLM

DevZero Operator

Checkpoint ControllerRestore ControllerMigration ControllerStorage Sync

Kubernetes Cluster

GPU Node 0
B200
GPU Node 1
H100
GPU Node 2
L40S

S3 / GCS

Checkpoints

Built on CRIU, extended for GPUs

DevZero extends CRIU (Checkpoint/Restore In Userspace) with GPU-aware checkpointing that captures CUDA context, VRAM contents, and device state. We are core contributors and maintainers of the CRIU project, bringing novel improvements upstream.

Cloud portable

Deploy on EKS, GKE, AKS, or bare metal. Checkpoints are stored in any S3-compatible object store. Migrate workloads across availability zones or even cloud providers.

Operator-managed

Kubernetes-native CRD. Install via Helm chart. Checkpoint policies, storage backends, and migration triggers defined as custom resources. No application code changes.

Tested at production scale

Checkpoint and restore times measured on real models with full VRAM state, KV cache, and engine context. No shortcuts, no empty weights.

Llama 3.1 405B
Config: 8×B200, tensor-parallel
Size: ~380 GB
Save: 35s
Restore: 22s
Cold start: 11 min
Llama 3.1 70B
Config: 2×H100, tensor-parallel
Size: ~130 GB
Save: 14s
Restore: 10s
Cold start: 8 min
Mixtral 8×22B
Config: 4×H100, expert-parallel
Size: ~260 GB
Save: 25s
Restore: 16s
Cold start: 9 min
Llama 3.1 8B
Config: 1×L40S
Size: ~16 GB
Save: 3s
Restore: 2s
Cold start: 4 min

Benchmarks measured with vLLM serving on NVIDIA datacenter GPUs. Checkpoint stored to S3. Results vary by model architecture, GPU type, and network throughput.

How DevZero compares

FeatureDevZeroNVIDIA DynamoGKE Pod SnapshotsRaw CRIU
GPU checkpoint/restore
Engine agnosticN/AN/A
Multi-GPU support
Cloud portable
Remote storage (S3/GCS)
Live migration
Spot preemption recovery
Kubernetes CRD
Compressed checkpoints

Comparison based on publicly available documentation as of March 2026.

Frequently asked questions

What our customers say

Databahn logo

We were essentially able to reduce the cost of that cluster by about 75%. On AWS, DevZero demonstrated they could achieve significantly higher savings than we initially thought possible.

Mihir Nair

Mihir Nair

Head of Architecture, Databahn

Stop wasting GPUs on cold starts

See how checkpoint-restore can cut your inference startup time by 50× and reduce GPU costs.