Pay for what
you ship.

No seats to count, no minimums. Start free, scale per request, talk to us at volume.

Free

Build

$0
Everything you need to prototype and ship a first feature.
  • 50K requests / month
  • All agents and models
  • Live playground and console
  • Community support
Start free
Most teams
Scale

Scale

$0.20 / 1K calls
Usage-based, volume discounts kick in automatically as you grow.
  • Unlimited requests
  • p50 under 100ms, priority pools
  • Render pipeline and webhooks
  • Email support, 99.9% SLA
Start on Scale
Enterprise

Enterprise

Custom
Volume commitments, residency, and a team that knows your stack.
  • Committed-volume pricing
  • Dedicated region and VPC peering
  • SSO, audit log, 99.99% SLA
  • Named solutions engineer
Talk to sales

Per-model rates

Scale pricing. Billed per 1,000 calls, metered to the request.

Model classLatency p50Price / 1K
Text and embeddings85ms$0.02 – $0.20
Agents640ms$0.55 – $0.80
Structured and rerank120ms$0.08 – $0.90
Image render2.1s$0.70 / render
Audio and voice900ms$0.15 – $1.10

Questions, answered

How does metering work?

Every call is metered to the request. You see live usage and projected spend in the console, broken down by model and key.

Do volume discounts apply automatically?

Yes. As monthly volume grows, per-call rates step down without a contract change. Enterprise commitments unlock the lowest tiers.

Where does inference run?

Default region is eu-central-1 in Frankfurt, GDPR-aligned. Enterprise plans can pin a dedicated region.

What counts against the free tier?

Any successful prediction or render. Failed calls and health checks are never billed.

Start free.
Scale when ready.

No card to start. Move to usage-based the day you ship.

Get your API keyTalk to sales