The P95 Problem
Median latency is marketing; the tail is what your users experience.
- Measured p95/p50 TTFT ratios span 1.5× to 7.2× across providers.
- Scaleway shows the heaviest measured tail: 172 ms median vs 1234 ms p95.
- Groq serves the tightest distribution — its p95 stays within 1.5× of median.
Distributions, not points
Two providers with similar medians can ship completely different user experiences: one in twenty requests lives at p95, and for interactive products that request defines perceived quality. The ratio, not the median, is the engineering number.
| Provider | p50 | p95 | p95/p50 |
|---|---|---|---|
| Scaleway | 172 ms | 1234 ms | 7.2× |
| OpenRouter | 656 ms | 3027 ms | 4.6× |
| Mistral | 307 ms | 1009 ms | 3.3× |
| Cortecs | 425 ms | 1254 ms | 3.0× |
| Together | 572 ms | 1625 ms | 2.8× |
| OpenAI | 1040 ms | 2453 ms | 2.4× |
| Cerebras | 184 ms | 323 ms | 1.8× |
| Groq | 262 ms | 384 ms | 1.5× |
- Set SLOs on p95, benchmark on p95, and treat median-only comparisons as incomplete by construction.
InferenceBench (2026). The P95 Problem. InferenceBench Research.