PERFORMANCEGLM-5.2 → Scaleway: TTFT ↑ 20%8mPERFORMANCEDeepSeek V4 Flash → Scaleway: TTFT ↑ 35%8mPERFORMANCEMistral Medium 3.5 → Scaleway: TTFT ↓ 21%8mPERFORMANCEDeepSeek V4 Flash → Scaleway: throughput ↓ 20%8mPERFORMANCEGPT-OSS 120B → Scaleway: throughput ↑ 31%8mPERFORMANCEGPT-OSS 120B → Scaleway: TTFT ↓ 24%8mPERFORMANCEGPT-OSS 120B → Together AI: throughput ↑ 20%8mPERFORMANCEKimi K3 via OpenRouter: TTFT ↓ 49%8mPERFORMANCEGLM-5.2 via OpenRouter: reliability recovered 96.8% → 100.0%8mPERFORMANCEKimi K3 via Cortecs: throughput ↑ 184%8mPERFORMANCEKimi K3 via Cortecs: TTFT ↓ 42%8mPERFORMANCEGLM-5.2 via OpenRouter: throughput ↓ 49%8mPERFORMANCEGLM-5.2 via OpenRouter: TTFT ↑ 121%8mPERFORMANCEMiniMax M3 via OpenRouter: reliability recovered 96.8% → 100.0%8mPERFORMANCEDeepSeek V4 Pro via Cortecs: throughput ↑ 27%8mPERFORMANCEGPT-OSS 120B via OpenRouter: throughput ↓ 22%8mPERFORMANCEGLM-4.7 via OpenRouter: TTFT ↓ 78%8mPERFORMANCEGPT-OSS 20B → Groq: TTFT ↑ 24%8mPERFORMANCEKimi K3 → Together AI: TTFT ↓ 34%8mPERFORMANCEDeepSeek V4 Flash via OpenRouter: throughput ↓ 17%8mPERFORMANCELlama 3.3 70B via Cortecs: TTFT ↓ 51%8mPERFORMANCEQwen3 235B via OpenRouter: throughput ↑ 37%8mPERFORMANCEQwen3 235B via OpenRouter: TTFT ↓ 29%8mPERFORMANCEDeepSeek V4 Flash → Together AI: throughput ↑ 19%8mPERFORMANCEGLM-5.3 via OpenRouter: throughput ↓ 17%8mPERFORMANCEGPT-OSS 20B via Cortecs: throughput ↑ 28%8mPERFORMANCELlama 3.3 70B → Together AI: TTFT ↓ 33%8mPRICE~moonshotai/kimi-latest cached_input_per_mtok decreased from 0.80 to 0.298mPRICE~moonshotai/kimi-latest output_per_mtok decreased from 13.00 to 11.368mPRICE~moonshotai/kimi-latest input_per_mtok decreased from 0.99 to 0.758mPRICE~deepseek/deepseek-v4-flash-latest cached_input_per_mtok decreased from 0.0077 to 0.00118mPRICE~deepseek/deepseek-v4-flash-latest output_per_mtok decreased from 1.28 to 1.048mPRICE~deepseek/deepseek-v4-flash-latest input_per_mtok decreased from 0.0077 to 0.00388mPRICEz-ai/glm-5.1 cached_input_per_mtok increased from 0.18 to 0.268mPRICEz-ai/glm-5.1 output_per_mtok increased from 3.03 to 4.408mPRICEz-ai/glm-5.1 input_per_mtok increased from 0.96 to 1.408mPRICEtencent/hy4-preview cached_input_per_mtok increased from 0.038 to 0.0428mPRICEtencent/hy4-preview input_per_mtok increased from 0.75 to 0.838mPRICEtencent/hy3 cached_input_per_mtok increased from 0.021 to 0.0338mPRICEtencent/hy3 output_per_mtok increased from 0.33 to 0.538m

OpenRouter — AI inference benchmarks

US · routed · 13 benchmarked models · 2 regions
#8 GLOBALAdd to compare
Overall85.2Rel.99.5%Tools—JSON—TTFT656 msP95—Tok/s91Price$0.30 / $1.20Δ30d↑ 3.3Excluded this window: 202 throttled · 0 client faults● last run 7 min ago · 52 runs in 24h

Head-to-head: vs Cerebras · vs Groq · vs Cortecs · vs Together · vs Scaleway · vs Mistral

Provider vs market

Percentile across 8 ranked providers
Reliability 99.5%P25
Tools 99.9%P38
Structured 96.4%P25
TTFT 656 msP25
Throughput 91 tok/sP13
Cost $1.20P63

Price vs performance

OpenRouter highlighted vs peers
889296100$4.37$8.74$13.10$17.47Output price $/M →Score ↑Groq — $0.45/M out · score 99.2Cerebras — $0.75/M out · score 99.0Mistral — $4.50/M out · score 93.2Scaleway — $1.81/M out · score 92.4Cortecs — $1.23/M out · score 92.1OpenAI — $15.60/M out · score 90.4Together — $1.12/M out · score 85.5OpenRouter — $1.20/M out · score 85.2OpenRouter

30-day trend

Measured daily series
Score85.2
TTFT480 ms
Tools100.0%
Reliability99.1%
Price$3.99

Key positioning

Derived from benchmark percentiles
Strengths
  • Balanced profile
Weaknesses
  • Reliability (P25)
  • Tools (P38)
  • Structured (P25)
  • TTFT (P25)
  • Throughput (P13)
Best suited for

OpenRouter in the registry

  • HeadquartersUS
  • API protocolsopenai chat completions
  • Catalogued endpoints1
  • Websiteopenrouter.ai

Observed models & declared pricing

Read from OpenRouter's own catalogue by continuous discovery — declared data, not benchmark results.

ModelContext$ in / M$ out / MToolsObserved
~anthropic/claude-fable-latest1000k10.0050.00yes2026-10-03
~anthropic/claude-haiku-latest200k1.005.00yes2026-10-03
~anthropic/claude-opus-latest1000k4.0020.00yes2026-10-03
~anthropic/claude-sonnet-latest1000k2.0010.00yes2026-10-03
~deepseek/deepseek-flash-latest1049k0.0150.68yes2026-10-03
~deepseek/deepseek-pro-latest1049k0.130.40yes2026-10-03
~deepseek/deepseek-v4-flash-latest1049k0.00381.04yes2026-10-03
~google/gemini-flash-latest1049k0.753.75yes2026-10-03
~google/gemini-pro-latest1049k2.0012.00yes2026-10-03
~moonshotai/kimi-latest1049k0.7511.36yes2026-10-03
~openai/gpt-astra-latest1050k10.0050.00yes2026-10-03
~openai/gpt-luna-latest1050k0.100.50yes2026-10-03
~openai/gpt-mini-latest400k0.754.50yes2026-10-03
~openai/gpt-sol-latest1050k2.0010.00yes2026-10-03
~openai/gpt-terra-latest1050k2.0012.00yes2026-10-03
~x-ai/grok-latest500k2.006.00yes2026-10-03
~z-ai/glm-flash-latest1049k0.0260.93yes2026-10-03
~z-ai/glm-latest1049k0.124.00yes2026-10-03
aion-labs/aion-2.0131k0.801.60yes2026-10-03
aion-labs/aion-3.0131k3.006.00yes2026-10-03
aion-labs/aion-3.0-mini131k0.701.40yes2026-10-03
aion-labs/aion-3.5262k3.006.00yes2026-10-03
aion-labs/aion-3.5-mini262k0.701.40yes2026-10-03
aion-labs/aion-rp-llama-3.1-8b33k0.801.60—2026-10-03
amazon/nova-2-lite-v11000k0.302.50yes2026-10-03
amazon/nova-lite-v1300k0.060.24yes2026-10-03
amazon/nova-micro-v1128k0.0350.14yes2026-10-03
amazon/nova-premier-v11000k2.5012.50yes2026-10-03
amazon/nova-pro-v1300k0.803.20yes2026-10-03
anthracite-org/magnum-v4-72b33k2.505.00—2026-10-03

All providers · How we benchmark