InferenceBench
EuropeProvidersModelsResearchPartners
MethodologyDEV DATA
BENCHMARKFirst benchmarks recorded: MiniMax M3 → Cortecs → ?3hBENCHMARKFirst benchmarks recorded: Kimi K3 → Cortecs → ?3hBENCHMARKFirst benchmarks recorded: DeepSeek V4 Pro → Cortecs → ?4hBENCHMARKFirst benchmarks recorded: GLM-4.7 → Cortecs → ?4hBENCHMARKFirst benchmarks recorded: GLM-5.2 → Cortecs → ?4h
BENCHMARKFirst benchmarks recorded: MiniMax M3 → Cortecs → ?3hBENCHMARKFirst benchmarks recorded: Kimi K3 → Cortecs → ?3hBENCHMARKFirst benchmarks recorded: DeepSeek V4 Pro → Cortecs → ?4hBENCHMARKFirst benchmarks recorded: GLM-4.7 → Cortecs → ?4hBENCHMARKFirst benchmarks recorded: GLM-5.2 → Cortecs → ?4h

Research

Independent analysis of AI inference performance, pricing and infrastructure, based on InferenceBench data.

AllMarket ReportBenchmarkMarket IntelligenceEngineering StudyData Note
ENGINEERING STUDY

Same Model, Different Infrastructure: How Large Is the Gap?

TTFT for Qwen3-235B spans 148–587 ms across 12 providers — a 4.0× spread.

104k runs · Aug 15, 2026
ENGINEERING STUDY

P95 Matters: The Hidden Tail Latency of AI Providers

12 of 13 providers (92%) show p95 TTFT above 2× their median.

104k runs · Aug 15, 2026
ENGINEERING STUDY

Tool Calling Is an Infrastructure Problem Too

Qwen3-235B tool success spans 94.7%–99.2% (4.5 points) across 12 providers under the Agents workload.

104k runs · Aug 15, 2026
ENGINEERING STUDY

Structured Output Reliability Across Providers

Structured-output validity spans 96.0%–99.8% across 12 providers on the same model.

104k runs · Aug 15, 2026

Support independent AI infrastructure research.

Benchmark Partners · Become a Founding Partner →
InferenceBenchIndependent intelligence for AI inference infrastructure.
Providers · Models · Europe · Research · Methodology · Legal · Privacy · Cookies · Partner Terms · Partners · Become a Partner →InferenceBench is operated by SnowStorm Solutions S.L. · © 2026 SnowStorm Solutions S.L.