Research
Independent analysis of AI inference performance, pricing and infrastructure, based on InferenceBench data.
ENGINEERING STUDY
Same Model, Different Infrastructure: How Large Is the Gap?
TTFT for Qwen3-235B spans 148–587 ms across 12 providers — a 4.0× spread.
ENGINEERING STUDYP95 Matters: The Hidden Tail Latency of AI Providers
12 of 13 providers (92%) show p95 TTFT above 2× their median.
ENGINEERING STUDYTool Calling Is an Infrastructure Problem Too
Qwen3-235B tool success spans 94.7%–99.2% (4.5 points) across 12 providers under the Agents workload.
ENGINEERING STUDYStructured Output Reliability Across Providers
Structured-output validity spans 96.0%–99.8% across 12 providers on the same model.