Qwen3-235B

12 providers · 235B MoE · 128k context · 3 regions · all models →

Qwen3-235B across 12 providers

Same model. Different infrastructure. DEV DATA
TTFT spread4.0×Throughput spread28.4×Output price3.9×Tool reliability4.3 pts
TTFT148 ms587 msbest: Cerebras
Throughput74 tok/s2100 tok/sbest: Cerebras
Output price$0.31/M$1.20/Mbest: SiliconFlow
Reliability96.6%99.3%best: Fireworks
Tool success94.8%99.1%best: OpenRouter

Qwen3-235B — tool-calling benchmark

12 providers · 94.8% → 99.1% · 4.3 pt spread · DEV DATA
#ProviderToolsΔ bestRel.JSONTTFTP95Tok/s$ Out
01Fireworks99.1%99.3%99.5%312 ms918 ms186$0.88
02OpenRouter99.1%−0.098.0%98.9%384 ms1106 ms160$0.88
03Cerebras98.5%−0.698.4%98.4%148 ms501 ms2100$1.20
04Groq98.3%−0.898.5%98.2%179 ms530 ms750$0.59
05Nebius98.0%−1.198.9%99.3%288 ms1061 ms142$0.40
06Scaleway97.2%−1.998.2%97.4%356 ms1025 ms96$0.72
07Together97.0%−2.197.9%98.0%401 ms1452 ms130$0.72
08DeepInfra96.3%−2.896.8%98.0%512 ms1337 ms88$0.36
09Nextbit96.3%−2.897.4%97.0%341 ms1158 ms84$0.44
10Mistral95.2%−3.997.6%97.1%298 ms963 ms110$1.10
11OVHcloud95.1%−4.097.3%95.8%372 ms1345 ms74$0.68
12SiliconFlow94.8%−4.396.6%97.9%587 ms1986 ms92$0.31
Qwen3-235B inference benchmark: fastest & cheapest providers · InferenceBench