Same Model, Different Infrastructure
The identical open-weight model, measured across every provider that serves it.
- GPT-OSS 120B spans 179–474 ms median TTFT across measured providers — a 2.6× spread on identical weights.
- Measured reliability on the same model ranges 0.0%–100.0%.
- Output price for the same tokens spans $0.17–$0.75 per million.
One model, many products
GPT-OSS 120B is the widest-covered model in the measured set (6 providers). Whatever explains the differences — hardware, batching policy, routing, quantisation — the buyer experiences them as different products at different prices wearing the same name.
TTFT179 ms474 ms
Throughput68 tok/s1749 tok/s
Output price$0.17/M$0.75/M
Reliability93.5%100.0%
- Never extrapolate a model's behaviour from one provider's serving of it.
- Model benchmarks without an infrastructure axis hide the variable that dominates production experience.
InferenceBench (2026). Same Model, Different Infrastructure. InferenceBench Research.