The State of AI Inference — August 2026
The first fully measured snapshot: who leads, what a frontier API buys, and how far apart identical models sit.
- Cerebras leads the measured provider ranking at 99.0 under the General profile.
- On GPT-OSS 120B — the widest-covered model (6 providers) — identical weights span 179–474 ms median TTFT, while tool-calling reliability from 0.0% to 100.0%.
- 2 of 8 measured providers sit on the price–performance frontier; every other provider is dominated on both axes at once.
The measured ranking
Aggregated across every measured model with coverage weighting and freshness shading, Cerebras holds #1 at 99.0, ahead of Groq (98.2) and Mistral (93.2). The gap between #1 and the median measured provider is 6.9 points — the market is not close.
Frontier APIs enter the table measured like everyone else, under the same request shape, the same exclusion rules and the same window statistics as every open-weight serving specialist.
Identical weights, different products
GPT-OSS 120B is served by 7 measured providers. The same weights produce median TTFT from 179 to 474 ms (2.6×) and output price from $0.17 to $0.75 per million depending on whose infrastructure executes them — tool-calling reliability from 0.0% to 100.0% across the set. Infrastructure is not a commodity layer — it is the product.
- Model selection and provider selection are separate decisions of comparable weight — benchmark both.
- Frontier pricing is not a proxy for serving quality; measure the path you intend to ship.
InferenceBench (2026). The State of AI Inference — August 2026. InferenceBench Research.