Research
Independent analysis of AI inference performance, pricing and infrastructure, based on InferenceBench data.
MARKET REPORT · FEATURED
The State of AI Inference — August 2026
The first fully measured snapshot: who leads, what a frontier API buys, and how far apart identical models sit.
9,119 results8 providers15 models● MEASURED
MARKET REPORT
The European Inference Report
3 of 8 measured providers qualify as European — EU-based or offering EU data residency.
BENCHMARKSame Model, Different Infrastructure
GPT-OSS 120B spans 179–474 ms median TTFT across measured providers — a 2.6× spread on identical weights.
MARKET INTELLIGENCEThe Price–Performance Frontier
Measured output prices span $0.45–$15.60 per million tokens — a 35× range.
ENGINEERING STUDYThe P95 Problem
Measured p95/p50 TTFT ratios span 1.5× to 7.2× across providers.
ENGINEERING STUDYTool Calling Is an Infrastructure Property
On DeepSeek V4 Flash, measured tool-call reliability spans 0.0%–100.0% across 5 providers serving identical weights.
ENGINEERING STUDYStructured Output: Trust, but Validate
Measured structured-output validity spans 91.7%–100.0% across providers.