Kubernetes workload optimization
that runs on autopilot

DevZero rightsizes CPU and memory in-place using CRIU checkpointing. Unlike VPA, it never restarts pods. Workloads keep running, costs drop in minutes.

node-04

m5.2xlarge

28% used
CRIU snapshot

dakr-cache-a1b2

18% actual  /  80% allocated

+62%

dakr-cache-c3d4

12% actual  /  80% allocated

+68%

api-worker-e5f6

9% actual  /  80% allocated

+71%

svc-proxy-g7h8

21% actual  /  80% allocated

+59%

Detect waste

72% CPU idle

Checkpoint

CRIU snapshot

Restore

rightsized node

Done

0 restarts · 15s

node-05

New

rightsized

72%

CPU reclaimed

0

Pod restarts

15s

Restore time

CRIU checkpoint

Method

Companies who slashed their Kubernetes
spend
using DevZero

DATABAHN
Starburst
Fi
Outerbounds
Codilas
personality pool
Onnitech
OpenObserve
Parsimo
Dentira
DATABAHN
Starburst
Fi
Outerbounds
Codilas
personality pool
Onnitech
OpenObserve
Parsimo
Dentira

Resize CPU and memoryin-place with no restarts

The write operator adjusts CPU and memory requests based on observed usage. On Kubernetes 1.33+, changes apply in-place without restarting pods. Choose Balanced, Conservative, or Aggressive mode per workload. The operator falls back to a rolling restart automatically when in-place resize is not feasible.

Search and filter data
Clear All
WorkloadCPU RequestsMem RequestsReq. Based CostOptimizationsHealth
Poorafka-mon-kafka-exporter
30.05m / 500m
46.89 / 512 MiB
$0.1260CPU $0.1109 · Mem $0.0151Not OptimizedHealthy
retina-agent
0.51 / 34.41 cores
29.02 / 67.22 GiB
$10.5866CPU $8.3600 · Mem $0.2266Not OptimizedRecovered
microsoft-defender-publi...
0.42 / 10.25 cores
13.99 / 10.67 GiB
$2.9768CPU $2.5080 · Mem $0.4688Not OptimizedRecovered
kube-proxy
3.76 / 33.14 cores
12.41 / 0 GiB
$8.7879CPU $8.3604 · Mem $0.4276Not OptimizedHealthy
prometheus-prometheus...
1.25 / 16 cores
26.34 / 108 GiB
$6.8033CPU $3.5424 · Mem $3.2609Not OptimizedRecovered
abnormal-abuse-campai...
2.56m / 25m
16.6 / 18.77 MiB
$0.0067CPU $0.0061 · Mem $0.0006Optimized1 policy attachedN/A
CostCurrently Showing
$1,200.55

32% of workloads (72)

account for ~80% of total cost

CPU
4%request utilization

13% of workloads (30)

account for ~80% of CPU usage

Memory
19%request utilization

26% of workloads (57)

account for ~80% of memory usage

GPU
0%request utilization

50% of workloads (2)

account for ~80% of GPU usage

Cost Distribution

ml--apac-...

$24.07

ml--earth-...

$24.05

ml--emea-...

$23.97

ml--mercu...

$21.08

ml--apac-...

$20.97

ml--apac-...

$20.90

qis--apac...

$20.66

qis--previ...

$20.56

qis--previ...

$19.62

qis--nar-st...

$20.55

qis--regre...

$20.55

gossip--us...

$20.36

gossip--int...

$20.40

qis--nar-pr...

$17.62

yahoo-otel...

$16.15

qis--regre...

$17.46

gossip--u...

$15.08

node-loca...

$14.94

qis--eme...

$13.24

avengers...

$13.15

qis--apa...

$11.70

ml--nar-e...

$11.38

gossip--...

$10.19

gossip--...

$10.18

ml--eme...

$9.81

mi--eme...

$9.67

mi--nar-...

$9.67

gossip--...

$9.41

gossip-...

$9.41

gossip--...

$9.40

gossip--...

$9.39

ml--merc...

$9.39

ml--nar-...

$9.07

qis--apa...

$8.73

dz-prom...

$7.95

s1-agent

$7.97

ais--apa...

$7.57

You're paying for nodesnobody is using

DevZero breaks down infrastructure spend by cluster, namespace, workload, and team using live and historical metrics. Per-workload visibility into CPU, memory, and GPU shows exactly where waste is. Savings projections are available from day one.

Reclaim idle GPUjob keeps running

DevZero tracks GPU utilization per workload and surfaces idle capacity between training phases. Reclamation applies without interrupting active jobs. GPU waste across H100, A100, L4, and T4 instances is tracked and quantified continuously.

Projected Monthly Cost

$48.51K/mo

Current Node Count

110

Overall Underutilization

21%

Node Capacities
CPU1,002.78 / 1,245.69 cores (81%)
Memory3,356.95 / 4,816.72 GiBs (70%)
GPU22.15 / 22 devices (101%)
Instance Breakdown

110 On-Demand

$48,509.4

0 Spot

$0.0

0 Reserved

$0.0

Most Underutilized Node

ip-10-169-16-90.ec2.internal
m6i.8xlarge

99.4%

underutilized

$1000$500$200$100
Current Cost$660.00Usage Cost$390.00

Why DevZero over native Kubernetes tools

VPA, Karpenter, and manual limits each leave money on the table. DevZero covers all three gaps.

CPU USAGE - LIVEPredictive mode
api-1
api-2
worker
svc-a
svc-b
burst
ml-job
Requested (over-provisioned)
DevZero rightsized

CPU rightsizing

Adjusts CPU requests using max observed usage in Balanced mode or P90 in Aggressive mode, per container. Changes apply in-place without triggering a rolling restart.

MEMORY - RSS VS
REQUESTED
OOM watch: active
768Mi512Mi256MiOOM risk↑ scaled up
Requested (static)
Actual RSS
DevZero

Memory optimization

Sizes memory requests against actual observed usage. Scales up before OOM pressure causes eviction. Conservative mode adds 1.2x headroom for stateful or unpredictable workloads.

GPU UTIL - H100
TRAINING JOB
Reclaiming idle
LoadingTrainingValidationIdle100%10%69 GBreclaimed
GPU utilization
Actual RSS

GPU reclamation

Reclaims idle GPU between training phases in real time. Eliminates the biggest source of GPU waste on H100, A100, L4, and T4 instances.

Built for platform engineers

The controls your team actually needs, not a dashboard that requires a PhD to interpret.

cgroup v2 · live
CPU throttle
78%
Mem PSI stall
43ms
OOM events
0
sub-second
prometheus · 15s avg
CPU usage
38%
Mem usage
51%
Throttle visible
averaged out

P99 latency spike · api-server-7f6d9

throttle_pct=0.78 · caught in 180ms window

Detected

DevZero reads directly from the Linux cgroup v2 hierarchy, not Prometheus scrapes. CPU throttle microseconds, memory pressure stall data, and OOM kill counters are surfaced at sub-second resolution, making short-lived anomalies that cause P99 latency spikes visible and actionable.

EKS
841
pods managed
GKE
519
pods managed
AKS
302
pods managed
On-prem
178
pods managed
Single tenant1,840
pods · 4 clusters
$41kMonthly Savings
3Cloud Providers
4Clusters
1Tenant

One DevZero tenant manages rightsizing across all your clusters. Cost savings are aggregated per team, namespace, and workload owner, with cloud billing integrated via AWS Cost Explorer, GCP Billing Export, or Azure Cost Management.

DEVZERO

cpu.request: 2000m → 680m · conf=0.94
Triggered

SLACK · #PLATFORM-RIGHTSIZING

Advisory mode · saves $2.84/hr
Posted

GITHUB · HELM-VALUES PATCH

resources.requests.cpu: "680m"
PR open

Advisory mode sends recommendations where your team already works. GitHub PR integration patches your Helm values file directly. PagerDuty suppression prevents rightsizing changes during active incidents.

What our customers say

Databahn logo

We were essentially able to reduce the cost of that cluster by about 75%. On AWS, DevZero demonstrated they could achieve significantly higher savings than we initially thought possible.

Mihir Nair

Mihir Nair

Head of Architecture, Databahn

Frequently asked questions

Technical questions from platform engineers who've evaluated DevZero.

Most clusters are overprovisioned.Let's prove yours is.

Run a free assessment to identify overprovisioned workloads, idle capacity, and your potential savings, in minutes.

Optimize now