Inference pricing that
stops at the token
One API for 500+ models across 70+ providers. Pay provider list price, add one flat platform fee, and bring your own keys or GPUs whenever it works out cheaper.
Free
Kick the tires on open models — no card required.
$0platform fee
100 requests / day
- 20+ open models, 3 shared providers
- Playground and OpenAI-compatible API
- Latency and throughput metrics
- 7-day activity log
Most Popular
Pay-as-you-go
Every model, every provider, billed per token.
4%platform fee on inference spend
No minimum, no commitment
- 500+ models across 70+ providers
- $50K/mo of BYOK traffic with no fees
- Budgets, prompt caching and spend alerts
- Email + shared Slack support
Enterprise
Committed capacity, private routing, contractual SLAs.
Customvolume pricing
Invoicing and AWS Marketplace
- Dedicated throughput and committed-use pricing
- Policy-based routing and residency controls
- SSO/SAML, SCIM and managed policy enforcement
- Negotiated uptime SLA and a named account team
| Feature | Free | Pay-as-you-go | Enterprise |
|---|---|---|---|
| Access | |||
| Models | 20+ open models | 500+ models | 500+ models |
| Providers | 3 shared providers | 70+ providers | 70+ providers |
| Playground & OpenAI-compatible API | |||
| Bring your own provider keys (BYOK) | — | $50K/mo free, 4% after | Custom limits |
| Bring your own GPUs | — | Self-serve connect | Managed fleets |
| Rate limits | 100 req/day | High shared limits | Dedicated capacity |
| Cost control | |||
| Platform fee | None | 4% on inference spend | Negotiable / volume tiers |
| Token pricing | Free models only | At-provider list price | Committed-use discounts |
| Prompt & prefix caching | — | ||
| Budgets and spend alerts | — | ||
| Per-key and per-team cost attribution | — | ||
| Invoicing & committed spend | — | — | |
| Routing & reliability | |||
| Automatic failover between providers | |||
| Price and latency-aware routing | — | ||
| Preferred-provider ordering | — | ||
| Region and data-residency pinning | — | ||
| Policy-based routing | — | — | |
| Uptime SLA | — | — | By contract |
| Observability | |||
| Activity logs | 7 days | 90 days | Custom retention |
| Log export & webhooks | — | ||
| Latency, TTFT and throughput metrics | |||
| Evaluation and A/B traffic splits | — | ||
| Security & governance | |||
| Zero data retention on request | — | ||
| Management API keys | — | ||
| SSO / SAML & SCIM | — | — | |
| Managed policy enforcement | — | — | |
| VPC / self-hosted control plane | — | — | |
| Support | |||
| Support level | Community | Email + Shared Slack | Shared Slack + SLA |
| Onboarding & migration help | — | — | |
| Dedicated account manager | — | — | |
Token prices are set by the underlying providers and passed through at list price — the platform fee is the only markup. Ask about volume discounts.
Pricing FAQ
Billing and pricing
Usage and rate limits
Routing and latency
Privacy and security
Models and reliability
Ready to get started?
Start on free models in minutes, then scale onto committed capacity when the traffic shows up.