Reserved throughput, proven on your workload.
Run your workload on Impala, benchmark it on your real traffic,
then we tailor a dedicated, single-tenant reserved plan.
No rate limits, no pricing guesswork.
What you get
Single-tenant throughput
The highest throughput per GPU on open models, with no rate limits and no noisy neighbors.
Most intelligence per dollar
Lower cost than serverless alternatives at scale.
Enterprise-ready
Uptime SLAs, autoscaling, and hands-on dedicated support from our team.
Powered by the Impala Herd engine
Reserved throughput runs on Impala Herd — our inference engine that reshapes itself around your traffic: kernels, batching, decoding. Same hardware, more tokens, lower cost per token.
300%
Token throughput
Up to 300% on the same hardware
↓ PPMT
Price per million tokens
Throughput lands on unit cost
0
Changes on your side
Same models, API, and prompts


