PRICE~moonshotai/kimi-latest cached_input_per_mtok decreased from 0.80 to 0.292mPRICE~moonshotai/kimi-latest output_per_mtok decreased from 13.00 to 11.362mPRICE~moonshotai/kimi-latest input_per_mtok decreased from 0.99 to 0.752mPRICE~deepseek/deepseek-pro-latest cached_input_per_mtok increased from 0.0042 to 0.102mPRICE~deepseek/deepseek-pro-latest output_per_mtok increased from 0.40 to 4.202mPRICE~deepseek/deepseek-pro-latest input_per_mtok decreased from 0.13 to 0.132mPRICE~moonshotai/kimi-latest cached_input_per_mtok increased from 0.29 to 0.8033mPRICE~moonshotai/kimi-latest output_per_mtok increased from 11.36 to 13.0033mPRICE~moonshotai/kimi-latest input_per_mtok increased from 0.75 to 0.9933mPERFORMANCEGLM-5.2 → Scaleway: TTFT ↑ 20%1hPERFORMANCEDeepSeek V4 Flash → Scaleway: TTFT ↑ 35%1hPERFORMANCEMistral Medium 3.5 → Scaleway: TTFT ↓ 21%1hPERFORMANCEDeepSeek V4 Flash → Scaleway: throughput ↓ 20%1hPERFORMANCEGPT-OSS 120B → Scaleway: throughput ↑ 31%1hPERFORMANCEGPT-OSS 120B → Scaleway: TTFT ↓ 24%1hPERFORMANCEGPT-OSS 120B → Together AI: throughput ↑ 20%1hPERFORMANCEKimi K3 via OpenRouter: TTFT ↓ 49%1hPERFORMANCEGLM-5.2 via OpenRouter: reliability recovered 96.8% → 100.0%1hPERFORMANCEKimi K3 via Cortecs: throughput ↑ 184%1hPERFORMANCEKimi K3 via Cortecs: TTFT ↓ 42%1hPERFORMANCEGLM-5.2 via OpenRouter: throughput ↓ 49%1hPERFORMANCEGLM-5.2 via OpenRouter: TTFT ↑ 121%1hPERFORMANCEMiniMax M3 via OpenRouter: reliability recovered 96.8% → 100.0%1hPERFORMANCEDeepSeek V4 Pro via Cortecs: throughput ↑ 27%1hPERFORMANCEGPT-OSS 120B via OpenRouter: throughput ↓ 22%1hPERFORMANCEGLM-4.7 via OpenRouter: TTFT ↓ 78%1hPERFORMANCEGPT-OSS 20B → Groq: TTFT ↑ 24%1hPERFORMANCEKimi K3 → Together AI: TTFT ↓ 34%1hPERFORMANCEDeepSeek V4 Flash via OpenRouter: throughput ↓ 17%1hPERFORMANCELlama 3.3 70B via Cortecs: TTFT ↓ 51%1hPERFORMANCEQwen3 235B via OpenRouter: throughput ↑ 37%1hPERFORMANCEQwen3 235B via OpenRouter: TTFT ↓ 29%1hPERFORMANCEDeepSeek V4 Flash → Together AI: throughput ↑ 19%1hPERFORMANCEGLM-5.3 via OpenRouter: throughput ↓ 17%1hPERFORMANCEGPT-OSS 20B via Cortecs: throughput ↑ 28%1hPERFORMANCELlama 3.3 70B → Together AI: TTFT ↓ 33%1hPRICE~moonshotai/kimi-latest cached_input_per_mtok decreased from 0.80 to 0.291hPRICE~moonshotai/kimi-latest output_per_mtok decreased from 13.00 to 11.361hPRICE~moonshotai/kimi-latest input_per_mtok decreased from 0.99 to 0.751hPRICE~deepseek/deepseek-v4-flash-latest cached_input_per_mtok decreased from 0.0077 to 0.00111h
MARKET REPORT

The State of AI Inference — August 2026

The first fully measured snapshot: who leads, what a frontier API buys, and how far apart identical models sit.

Key findings
  • Cerebras leads the measured provider ranking at 99.0 under the General profile.
  • On GPT-OSS 120B — the widest-covered model (6 providers) — identical weights span 179–474 ms median TTFT, while tool-calling reliability from 0.0% to 100.0%.
  • 2 of 8 measured providers sit on the price–performance frontier; every other provider is dominated on both axes at once.

The measured ranking

Aggregated across every measured model with coverage weighting and freshness shading, Cerebras holds #1 at 99.0, ahead of Groq (98.2) and Mistral (93.2). The gap between #1 and the median measured provider is 6.9 points — the market is not close.

Frontier APIs enter the table measured like everyone else, under the same request shape, the same exclusion rules and the same window statistics as every open-weight serving specialist.

879196100$4.37$8.74$13.10$17.47Output price $/M →Score ↑Cerebras — $0.75/M out · score 99.0CerebrasGroq — $0.45/M out · score 98.2Mistral — $4.50/M out · score 93.2Scaleway — $1.81/M out · score 92.4Cortecs — $1.23/M out · score 92.1OpenAI — $15.60/M out · score 90.4Together — $1.12/M out · score 85.5OpenRouter — $1.20/M out · score 84.3
Score vs $/M output across measured providers · the frontier is where no provider is better and cheaper at once · ● MEASURED

Identical weights, different products

GPT-OSS 120B is served by 7 measured providers. The same weights produce median TTFT from 179 to 474 ms (2.6×) and output price from $0.17 to $0.75 per million depending on whose infrastructure executes them — tool-calling reliability from 0.0% to 100.0% across the set. Infrastructure is not a commodity layer — it is the product.

Scaleway179 msTogether182 msCerebras184 msCortecs240 msGroq265 msOpenRouter474 ms
Median TTFT on GPT-OSS 120B across measured providers — identical weights, best first · ● MEASURED
What this means for engineers
  • Model selection and provider selection are separate decisions of comparable weight — benchmark both.
  • Frontier pricing is not a proxy for serving quality; measure the path you intend to ship.
Dataset

Figures recompute from the live rolling window on every visit; the snapshot and versions above are the provenance of the regime that produced them. View the current benchmark →

Cite this researchInferenceBench (2026). The State of AI Inference — August 2026. InferenceBench Research.
← All research
Related research
The State of AI Inference — August 2026 · InferenceBench