PRICE~deepseek/deepseek-v4-flash-latest output_per_mtok decreased from 1.04 to 0.6018mPRICE~deepseek/deepseek-v4-flash-latest input_per_mtok increased from 0.0038 to 0.01318mPRICE~deepseek/deepseek-flash-latest cached_input_per_mtok decreased from 0.0021 to 0.002118mPRICEdeepseek/deepseek-v4.1-flash cached_input_per_mtok increased from 0.0021 to 0.00618mPRICEdeepseek/deepseek-v4.1-flash output_per_mtok increased from 0.68 to 1.2018mPRICEdeepseek/deepseek-v4.1-flash input_per_mtok increased from 0.015 to 0.3018mPRICE~deepseek/deepseek-v4-flash-latest cached_input_per_mtok decreased from 0.0077 to 0.001149mPRICE~deepseek/deepseek-v4-flash-latest output_per_mtok decreased from 1.28 to 1.0449mPRICE~deepseek/deepseek-v4-flash-latest input_per_mtok decreased from 0.0077 to 0.003849mPRICE~deepseek/deepseek-flash-latest cached_input_per_mtok decreased from 0.02 to 0.002149mPRICE~deepseek/deepseek-flash-latest output_per_mtok increased from 0.60 to 0.6849mPRICE~deepseek/deepseek-flash-latest input_per_mtok decreased from 0.02 to 0.01549mPRICEdeepseek/deepseek-v4.1-flash cached_input_per_mtok decreased from 0.02 to 0.002149mPRICEdeepseek/deepseek-v4.1-flash output_per_mtok increased from 0.60 to 0.6849mPRICEdeepseek/deepseek-v4.1-flash input_per_mtok decreased from 0.02 to 0.01549mPRICEdeepseek/deepseek-v4-flash-0731 cached_input_per_mtok increased from 0.0077 to 0.01749mPRICEdeepseek/deepseek-v4-flash-0731 input_per_mtok increased from 0.0077 to 0.01749mPRICE~deepseek/deepseek-v4-flash-latest cached_input_per_mtok decreased from 0.012 to 0.00771hPRICE~deepseek/deepseek-v4-flash-latest input_per_mtok decreased from 0.012 to 0.00771hPRICEdeepseek/deepseek-v4-flash-0731 cached_input_per_mtok decreased from 0.012 to 0.00771hPRICEdeepseek/deepseek-v4-flash-0731 input_per_mtok decreased from 0.012 to 0.00771hPRICEdeepseek/deepseek-v3.1-terminus input_per_mtok decreased from 0.30 to 0.271hPRICE~deepseek/deepseek-v4-flash-latest cached_input_per_mtok increased from 0.0011 to 0.0122hPRICE~deepseek/deepseek-v4-flash-latest output_per_mtok increased from 0.60 to 1.282hPRICE~deepseek/deepseek-v4-flash-latest input_per_mtok decreased from 0.013 to 0.0122hPRICEdeepseek/deepseek-v4-flash-0731 cached_input_per_mtok decreased from 0.017 to 0.0122hPRICEdeepseek/deepseek-v4-flash-0731 input_per_mtok decreased from 0.017 to 0.0122hNEW MODELnvidia/switchyard appeared in the catalogue2hPRICE~deepseek/deepseek-v4-flash-latest output_per_mtok decreased from 1.04 to 0.602hPRICE~deepseek/deepseek-v4-flash-latest input_per_mtok increased from 0.0038 to 0.0132hPRICEnvidia/nemotron-3-ultra-550b-a55b cached_input_per_mtok decreased from 0.12 to 0.102hPRICEnvidia/nemotron-3-ultra-550b-a55b output_per_mtok decreased from 2.40 to 2.202hPRICEnvidia/nemotron-3-ultra-550b-a55b input_per_mtok decreased from 0.60 to 0.502hPRICEdeepseek/deepseek-v4-flash-0731 cached_input_per_mtok increased from 0.0051 to 0.0172hPRICEdeepseek/deepseek-v4-flash-0731 input_per_mtok increased from 0.0051 to 0.0172hPRICE~moonshotai/kimi-latest cached_input_per_mtok decreased from 0.80 to 0.293hPRICE~moonshotai/kimi-latest output_per_mtok decreased from 13.00 to 11.363hPRICE~moonshotai/kimi-latest input_per_mtok decreased from 1.39 to 1.003hPRICE~deepseek/deepseek-v4-flash-latest output_per_mtok increased from 0.95 to 1.043hPRICE~deepseek/deepseek-v4-flash-latest input_per_mtok decreased from 0.0058 to 0.00383h

Methodology

How InferenceBench measures, normalizes and ranks AI inference infrastructure.

Methodology v0.13-devUpdated Aug 24, 202610,889 benchmark runsView changelog

Overview

InferenceBench measures how the same model behaves on different inference infrastructure. The unit of observation is the execution path — Model × Provider × Region × Deployment configuration — because that is what an engineering team actually ships against. Provider rankings are an explicit aggregation over those paths, never the raw unit.

A benchmark is only as credible as its measurement process. This page specifies what is measured, how it is executed, how results are normalized and aggregated, what is excluded, and where the limits are.

Measurement model

Every number on InferenceBench belongs to exactly one of three classes. The class is part of the metric's definition and is shown in the metrics registry below.

Observed
Measured directly by InferenceBench runners against live provider endpoints: TTFT, p95 TTFT, throughput, reliability, tool-calling success, structured-output validity.
Declared
Taken from official provider documentation at observation time: input/output prices, regions, data residency, context length, supported models, hosting type. Declared data never overrides an observation.
Derived
Computed by InferenceBench from observed and declared inputs: the overall score, E2E latency, percentiles, ranks, 30-day changes, provider aggregation.

Entity model

Provider
The commercial platform operating inference endpoints (e.g. Fireworks, Nebius). Provider rows in the ranking are aggregates, one per provider.
Model
The logical model (e.g. DeepSeek-V3), independent of where it runs.
Model × Provider
The basic unit of observation: a specific model served by a specific provider through a specific endpoint and region, under a stated deployment configuration. Internally: ExecutionPath.
ModelProviderCombination = {
  modelId, providerId,
  regionId?, endpointId?, deploymentConfig?
}

Benchmark environment

Benchmarks run as raw HTTPS requests from dedicated runner locations — no provider SDKs, so client-library overhead cannot differ between providers. The first runner operates from Falkenstein (DE); additional client regions are planned and every observation records its client region.

DimensionSpecification
HTTP clientRaw HTTPS/1.1, no provider SDKs
StreamingEnabled; token timestamps captured on arrival
ConnectionTLS session and connection reuse within a run batch
Timeout60 s per network operation (connect/read) plus a 120 s wall-clock bound per request; a timed-out request is recorded as a reliability failure, never discarded
RetriesDisabled — a failed request is a data point
Warm-upOne untimed warm-up request per endpoint per batch; cold-start behaviour is tracked separately (see Limitations)
ConcurrencySequential — one request in flight at a time (interactive profile); stress testing is out of scope for the public ranking
TimestampsMonotonic clock read in the runner process; stored latencies are whole milliseconds
Rate limitsEvery provider gets a 0.5 s default inter-request gap, tightened where an account's tier requires it (currently Mistral 1.2 s, Groq 10 s); Retry-After is honored (bounded at 60 s); a per-provider lock keeps the whole fleet to one suite per provider at a time. Throttled (429) responses are excluded from published denominators as account-tier artifacts and surfaced as counts
Prompt cachingEvery request carries a unique nonce in a system line, so providers with automatic prompt caching can never serve a sweep from a warm prefix — the hash-locked case prompts themselves stay byte-identical
Gateway routingFor routed paths, the upstream provider each request actually resolved to is recorded per result and per run — a routed row's numbers name the mixture they came from. The path's routing policy is applied to the wire: its declared request parameters (e.g. OpenRouter provider-sort preferences) are merged into every request including the warm-up, and a non-default policy that declares no request parameters refuses to run rather than measure default routing under its label
Serving precisionQuantization a provider itself declares per model (Nebius publishes it in its catalogue) is synced into the registry as a model variant and shown as a chip on that provider's rows (e.g. FP8). Undeclared precision stays unknown and is never inferred

Workloads

Workloads are versioned entities, not ad-hoc prompt sets. Each defines its prompt corpus size, input-length range, target output length, sampling parameters and the capabilities it exercises. Scores always name the workload profile they were computed under.

WorkloadCasesInput tokensOutput ceilingTemp.Capabilities
General generation v0.312≤40040960general
Tool calling v0.27≤60020480tool_calling
Structured output v0.36≤400schema-bound, 20480structured_output
Realtime v0.26≤20010240general (latency-weighted)

Request parameters are normalized per provider through recorded, versioned quirks, applied identically to every run of that provider: providers that reject an explicit temperature parameter (currently OpenAI and Anthropic) run with the parameter omitted at their documented default, and providers that reject response_format run structured-output cases prompt-only. A quirk is never inferred silently — each one is recorded the day it is observed.

Corpus sizes are deliberately small and versioned. Note the interaction with the sample gates: frontier-lab paths run correctness weekly, and one weekly sweep (7 tool cases, 6 structured) sits below the 12-result rate gate — their tool-calling and structured columns publish only after a second sweep lands in the window. Statistical weight comes from repetition — every case runs on every sweep. At the published cadence a hosted-model path accumulates ~217 eligible results per 7-day window (crossing the high-confidence threshold below); frontier-lab paths run correctness suites weekly for cost reasons and accumulate ~139, which is medium confidence and is labelled as such. Suites are hash-locked: changing any case without a version bump makes the loader refuse to run. Corpus growth is a funded-capacity decision and arrives as new suite versions, never as silent edits. Responses are never deliberately cut: every stream runs to the model's natural completion, and the per-case token ceiling exists only to bound runaway generation — it sits far above any well-formed interactive answer, reasoning burn included, with the 60 s per-operation timeout and the 120 s wall-clock bound as the outer bounds. This regime dates from 2026-08-19: earlier suite versions imposed output budgets sized for terse open-weight models, which turned reasoning-class models' internal deliberation into truncations (v0.2 of General generation raised those budgets the same day before budgets were dropped entirely). Reasoning burn is never scored as a reliability failure: a case whose entire budget is consumed by reasoning fails only that case's own capability metric, and it surfaces where a caller actually feels it — in measured TTFT and in token cost. Runs recorded under earlier versions keep citing them.

Model selection is operator-curated and never automatic, under three published rules: (1) first-party labs we measure are represented by their current flagship API models — a successor release enters, its predecessor stays measured while providers still serve it; (2) an open-weight model enters when at least two covered providers serve it, because same-model cross-provider comparison is the point of this observatory; (3) every credentialed provider keeps at least two measured models where its catalogue allows. Freshness is enforced by machine, decided by human: catalogues are re-discovered daily, new model ids surface as events the day they appear, and a daily coverage report cross-checks every tracked model against every credentialed provider's newest catalogue — naming the exact missing ids and flagging providers below the two-model minimum. Admission of a new model spends budget daily, so it stays an explicit operator decision; the mirror guarantees the decision is prompted the same day, never missed.

TTFT

Time to first token
TTFT = t(first_streamed_delta) − t(request_sent)

TTFT measures when the model demonstrably starts responding: the first streamed delta of any channel, reasoning included — which keeps reasoning and non-reasoning models comparable on the same scale. The wait until the first user-visible output (time-to-visible) is captured separately on every run; for non-reasoning models the two are usually identical (a leading whitespace delta can separate them by one frame), for reasoning models the difference IS the thinking phase, and it is never silently mixed into either number. TTFT includes DNS/TLS (amortized by connection reuse), provider queueing, prompt processing and the first delta's generation; it excludes benchmark-client processing. The table shows p50 as the primary metric; p95 is reported alongside because tail latency determines how a provider feels under production load. Percentiles are computed over the rolling 7-day sampling window per path, never pooled across models. All latency figures are measured from the canonical European vantage (Falkenstein, DE): for providers serving from outside Europe, TTFT therefore includes genuine transatlantic network distance. This is by design, stated rather than hidden — the benchmark answers “what does this path feel like from Europe”, which is the question its audience is asking. Additional vantage regions are on the roadmap; when they ship, every figure will carry its vantage explicitly.

StatisticRole
p50 TTFTPrimary table metric
p95 TTFTTail metric, shown as P95

End-to-end latency

E2E for a normalized 400-token response
E2E = TTFT + (400 / throughput) × 1000 [ms]

End-to-end latency is only comparable when output length is controlled, so InferenceBench normalizes it to a 400-output-token response using each path's observed sustained throughput. It is a derived metric and is labelled as such.

Throughput

Sustained output throughput
throughput = usage_completion_tokens / (t(last_delta) − t(first_delta))

Measured over the full delta span — TTFT is excluded, so the metric captures generation speed rather than queueing. The numerator is the provider-reported completion token count (reasoning tokens included) and the denominator spans every streamed delta, reasoning included: the same population on both sides. A response without reported usage contributes no throughput sample — SSE chunk counts undercount batched tokens and are never used as a token count — and responses shorter than 20 output tokens are excluded as degenerate speed readings. Spans under 250 ms and computed speeds above 5,000 tok/s are burst-delivery measurement artifacts, discarded at capture and retroactively excluded from aggregates. Unit: tokens per second.

Reliability

reliability = valid_completed_requests / eligible_requests
Counts against the providerExcluded from the denominator
HTTP 5xx responses, and 4xx bodies matching no client-fault patternAuth/account failures (401/403)
TimeoutsThrottles (429) and balance exhaustion (402/quota) — our account's tier and funding, not the provider's serving
Connection failures, malformed streams, mid-stream cutsRequest-shape rejections recorded as provider quirks
Empty responses / invalid API responsesMisregistered paths (wrong model id, non-serverless deployment)
—Invalid case or validator configuration

Hitting the per-case token ceiling (finish_reason=length) is a generation outcome, not a serving defect: the response was delivered, so it never lowers reliability — it fails the case, which for the tool-calling and structured suites shows in their published rates; the health suites record the failure but have no published pass rate yet. Excluded throttles stay visible: dossiers surface each provider's throttle count so the signal is published without becoming a score. Every exclusion persists a human-readable reason alongside the result, and when classification rules change, stored verdicts are re-classified under the new rules with the correction itself recorded. A run aborted by a runner or scheduler error is never written at all. Retries are disabled in the runner, so reliability reflects first-attempt behaviour.

Structured output

Structured output is not reduced to 'the JSON parses'. Each run is validated in three mechanical stages; a run only passes if every applicable stage passes.

JSON validity
The output parses as JSON.
Schema compliance
The parsed object validates against the requested JSON schema.
Constraint compliance
Stated constraints hold (enums, ranges, required fields, formats).

Tool calling

Tool-calling runs present a tool inventory and a task. A run passes only when all of the following hold:

  • the correct tool is selected (including a correct no-tool decision when no tool applies);
  • the argument object validates against the tool's schema;
  • all required arguments are present;
  • no hallucinated arguments are added;

Pricing

Input and output prices are tracked separately, per million tokens, from official provider documentation for the public serverless tier, and recorded in the currency the provider declares. For comparison — tables, rankings, the price dimension of scores and price/performance charts — non-USD prices are converted to USD at the ECB daily euro reference rate; the rate is fetched daily, stored dated with its source, and travels with each converted figure so the conversion is citable, never silent. Dossiers keep the declared figure alongside. A currency without a stored official rate is not converted and stays out of cross-currency comparisons. The score's cost dimension is the effective cost of the benchmarked exchange — the workload's published input size plus the same normalized 400-token response the E2E metric uses, at both declared prices — and it enters a score only when both prices are recorded:

Effective workload cost (the score's cost dimension)
effectiveCost = (workload_input_tokens / 1e6) × price_in + (400 / 1e6) × price_out

Input sizes per workload: General 400 tokens, Tool calling 600, Structured output 400, Realtime 200. Once measured per-path token mixes are materialized, the cost dimension will move to them as a versioned methodology change — that is what will make reasoning burn visible in cost.

Pricing is enforced as an invariant, not an aspiration: the scheduler refuses to plan new measurements for any path without a current declared price, a pricing gap surfaces loudly on every operations cycle until the catalogue or an operator resolves it, and operator-recorded prices are flagged for re-verification after 30 days.

Dedicated, batch and cached-input pricing are recorded when published but do not enter the public serverless ranking.

Scoring

The overall score is a weighted sum of market-normalized metric values under the active workload profile, shaded by measurement confidence. The weights below are read from the exact configuration the engine executes — the General profile is the homepage default; each workload has its own published profile.

score = Σ ( weight_i × normalize_market(metric_i) ) × confidence

Normalization runs against a fixed reference population — every measured path in the index — using the market's winsorized band: each dimension's range is the reference population's p5–p95, so a single outlier path clamps to the floor or ceiling instead of defining the whole scale. Published minimum spans per dimension (reliability 5 pp, tool-calling 5 pp, structured output 10 pp, latency 1,000 ms of end-to-end time, throughput 50 tok/s, effective workload cost $0.0002 per benchmarked exchange) keep a saturated market from amplifying sampling noise; the span anchors at the good extreme and extends only the weak side. Within the band, values map linearly onto 55–100 — the weak edge reads 'weak', not 'zero'. Only measured values shape a range; an unmeasured field never injects a fixture number into the reference. Weights renormalize over the dimensions actually measured for a path. An optimization tilt shifts 12 points of weight toward the user's selected priority; confidence applies the factors in the Sampling section.

Provider aggregation

A provider aggregate that averages path scores inherits the provider's catalogue: a lab serving only its own expensive reasoning model would get that model's properties billed as provider quality. The provider ranking's Overall is therefore Serving Quality — the provider's effect measured within the same model, which is the only place a provider effect is identifiable. These are the rules the engine runs:

  • Serving Quality: for every model measured on two or more providers, each serving dimension (reliability, latency as the completed exchange, throughput, tool calling, structured output) normalizes inside that model's own cross-provider band — same 55-floor grammar and published minimum spans as path scores. Price never enters (see Value below); the composite quality dimension never enters. A provider's Overall is the mean of its within-model effects across shared models × confidence factor × freshness factor.
  • Sole servers: a provider whose models nobody else serves (currently the frontier labs) has no counterfactual — its serving effect is unidentifiable. The row is flagged non-comparable and is never ranked on this axis; its paths remain fully measured, priced and comparable in Model × Provider mode.
  • Coverage: displayed on the row (models benchmarked vs the index's best coverage) and multiplies nothing. Cherry-picking resistance comes from the within-model design itself: every shared model the provider serves enters its mean.
  • Value: Serving Quality per effective dollar of the use-case's benchmarked exchange (workload input tokens + the normalized 400-token response, both prices USD-converted at the ECB reference rate). Shown as a percentile across comparable providers; exact prices live in Model × Provider mode.
  • quality metrics (reliability, tools, structured): mean across the provider's included paths; TTFT, P95, E2E, prices: medians; throughput: mean;
  • freshness factor = 1.0 (≤30 min), 0.995 (≤60 min), 0.99 (older); confidence factor = 1.0 / 0.988 / 0.968 (high / medium / low).

Sampling & confidence

Every path carries a run count and a confidence level over the rolling 7-day window. Confidence shades the score multiplicatively and gates winner eligibility.

LevelScore factorEffect
High×1.000Fully eligible, including #1 badges
Medium×0.988Ranked; shown with medium confidence
Low×0.968Visible but flagged; not eligible for #1 badges

Confidence levels map to eligible results in the rolling 7-day window: high at 200 or more, medium at 50 or more, low below 50 — computed per path from that path's own window evidence, never inherited from the provider's other models. Run counts include only canonical, finished runs. Below minimum window samples a statistic is not published at all: p50 needs 10 samples, p95 needs 30, p99 needs 100, rates need a 12-result denominator (two full sweeps of the smallest suite — the excluded failure mode is single-sweep publication), throughput needs 10 — the surface renders '—' instead of a number whose variance would swamp its meaning. Measured paths report real counts; in zero-fixture production an unmeasured pair is excluded entirely, while development fixtures stay labelled DEV DATA.

Outlier handling

Bad results are never silently deleted. Provider-attributable failures stay in reliability. Latency samples are winsorized at the 99th percentile within each path's window before p50/p95 computation — extreme values are capped, not removed. The stored p99 deliberately reads the uncapped tail (the p99 of a p99-winsorized list is the cap itself). Every exclusion under the client-fault rules of the Reliability section is logged with a reason.

Geography

'Europe' is not one property. InferenceBench tracks six distinct geographic dimensions per path and never collapses them:

Provider HQ
Where the operating company is headquartered.
Ownership region
Region of the controlling entity — an EU endpoint of a US-owned provider is not EU-owned.
Endpoint region
Where the inference request is processed.
Benchmark client region
Where the measurement originated (currently Falkenstein, DE).
Data residency
Where the provider guarantees request data remains.
Legal jurisdiction
The law governing the service contract.

Freshness & scheduling

Scheduling is tiered and published: health suites run daily on every measured path; correctness suites run daily on hosted-model paths and weekly on frontier-lab paths (a cost decision, stated as such); heavy suites run weekly. The stalest due (path, suite) pairs are enqueued first on every operations cycle, and detected provider changes (price edits, new endpoints, new model ids) surface as events the same day.

Two hard guards bound every cycle: a daily spend ceiling checked before any billable call, and the never-unpriced invariant — a path without a current declared price is not scheduled at all.

Every surface shows freshness: global ('updated 4m ago'), provider ('last benchmark 8m ago') and path level (run count and last-run age in the score breakdown).

Fairness

The following principles are structural, not aspirational — the pipeline has no code path that lets any of them be violated:

  • Providers cannot pay for ranking position.
  • Sponsorship does not change scores.
  • All providers are measured with the same methodology, workloads and runner locations.
  • Provider-supplied benchmark numbers never replace InferenceBench observations.
  • Corrections are versioned and published in the changelog.

Limitations

Honest limits of the current design — none of these are hidden from readers of the numbers:

  • Internet variability: measurements traverse the public internet; regional network effects are partially controlled by fixed runner locations.
  • Hidden routing: some providers route internally across clusters or vendors; InferenceBench observes the endpoint, not the internals.
  • Hardware opacity: provider hardware is not always disclosed and is never guessed.
  • Silent model updates and aliasing: providers may update weights behind a stable model name; re-benchmarks catch drift with a delay.
  • Quantization differences: declared serving precision is published as a variant chip on the provider's rows; unless declared, quantization is unknown and is not inferred.
  • Load and batching: provider-side batching and load vary; sampling windows smooth but cannot eliminate this.
  • Cold starts: the current design measures warm behaviour after one warm-up request; cold-start latency is not yet a published metric.
  • Workload representation: four workload profiles cannot represent every production pattern.
  • Coverage: the measured corpus is deliberately dev-scale and grows as an explicit budget decision; a pair not yet measured is excluded from rankings rather than estimated.
  • History: the 30-day score-change column stays empty until a path accumulates 30 days of measured history — it is never filled from fixtures in production.
  • Transport: benchmark traffic runs over HTTP/1.1; enabling HTTP/2 would change measured latencies and will only ever arrive as a versioned methodology change.

Reproducibility

The methodology is specified so that a third party could reimplement it: raw HTTPS, streaming timestamps, the formulas above, published workload parameters and the scoring configuration. A public runner CLI is planned but not yet available — no claim of current public tooling is made.

# planned public runner (not yet available)
inferencebench run --model deepseek-v3 --provider nebius --suite general

Versioning

Every observation and score records its methodology version. Historical rankings remain reconstructable under the methodology that produced them; a methodology change never silently rewrites the past. Research articles pin both a methodology version and a dataset snapshot.

Commercial independence

InferenceBench is funded through commercial partnerships that do not influence benchmark results. Partners support the engineering, infrastructure and research behind the benchmark; they cannot purchase ranking position, influence methodology, change scores, suppress results or control editorial conclusions.

  • Scores, rankings, winners and badges are computed by the published pipeline — there is no commercial input anywhere in it.
  • Partner identity lives on separate commercial surfaces (the partner rail, /partners); it never appears inside benchmark tables, winners, the ticker or Compare.
  • Research supported by a partner is disclosed on the article, and the partner has no control over methodology, provider selection, results or conclusions.
  • Conflicts are handled by disclosure and versioned corrections, never by silent edits.

InferenceBench is operated by SnowStorm Solutions S.L. and may be funded through commercial partnerships and related services.

Metrics registry

One definition per metric. The benchmark table tooltips and detail pages read from this registry — no surface redefines a metric locally.

Time to first token (p50)OBSERVEDTTFT · ms · lower is better · enters the score

Median elapsed time between sending the request and the first streamed delta of any channel — reasoning included, so reasoning and non-reasoning models compare on the same scale. Measured with a monotonic clock from the canonical vantage (Falkenstein, DE) over the rolling 7-day window. The wait until first visible output is captured separately.

TTFT = t(first_streamed_delta) − t(request_sent)
Time to first token (p95)OBSERVEDP95 · ms · lower is better

95th percentile of TTFT across the sampling window. Reveals tail latency that a median hides; a provider with a low p50 but high p95 will stall a meaningful share of production requests.

P95 = percentile(TTFT_samples, 0.95)
End-to-end latencyDERIVEDE2E · ms · lower is better

Total time for a normalized 400-output-token response, derived once over the whole window from its median TTFT and median sustained throughput, and published only when both inputs cleared their sample gates. Derived by construction — the two medians come from overlapping but not identical run sets.

E2E = TTFT_p50 + (400 / throughput_p50) × 1000
Output throughputOBSERVEDTok/s · tokens/s · higher is better · enters the score

Sustained output tokens per second over the full delta span: provider-reported completion tokens (reasoning included) divided by the first-to-last-delta interval. TTFT is excluded. Responses without reported usage contribute no sample, and burst-delivery artifacts (sub-250 ms spans, speeds above 5,000 tok/s) are discarded.

throughput = usage_completion_tokens / (t(last_delta) − t(first_delta))
ReliabilityOBSERVEDRel. · % · higher is better · enters the score

Share of eligible requests that complete with a valid, well-formed response, with client retries disabled. Provider-attributable failures (5xx, timeouts, connection and stream failures, empty responses) count against it. Benchmark-client faults — auth, throttles from our account's tier, balance exhaustion, request-shape rejections, misregistered paths — are excluded from the denominator, and hitting the token ceiling is a generation outcome scored by the case's own metric, never here.

reliability = valid_completed_requests / eligible_requests
Tool-calling successOBSERVEDTools · % · higher is better · enters the score

Share of tool-calling runs where the correct tool is selected with a schema-valid argument object, all required arguments present, no hallucinated arguments, and a correct no-tool decision when no tool applies.

Structured output validityOBSERVEDJSON · % · higher is better · enters the score

Share of structured-output runs that produce valid JSON, comply with the requested schema, and satisfy the stated constraints. JSON that parses but violates the schema counts as a failure.

Input priceDECLARED$ In · USD/M tokens · lower is better · enters the score

Provider-published price per million input tokens for the public serverless tier, as listed in official documentation at observation time. Enters the score through the effective workload cost.

Output priceDECLARED$ Out · USD/M tokens · lower is better · enters the score

Provider-published price per million output tokens for the public serverless tier, converted to USD at the stored ECB reference rate when declared in another currency. Together with the input price it forms the score's cost dimension: the effective cost of the benchmarked exchange, normalized to a 400-token response. The dimension enters a score only when both prices are recorded.

effectiveCost = (workload_input_tokens / 1e6) × price_in + (400 / 1e6) × price_out
InferenceBench scoreDERIVEDScore · points · higher is better

Weighted sum of market-normalized metric values under the active workload profile, shaded by measurement confidence. The normalization reference is the entire measured index — never the filtered view — so the same provider reads the same score under every filter and on every page; filters change who is listed, never the numbers.

score = Σ(weight_i × normalize_market_i) × confidence
30-day score changeDERIVEDΔ30d · points · higher is better

30-day change in the path's daily measured overall (reliability, latency and structured output — price, being declared, never enters the historical series). Only days measured under the current scoring regime participate: a window that would straddle a scoring-version change is suppressed, and the column stays empty until 28+ days of homogeneous history exist — never filled from fixtures.

Score weights

Read directly from the scoring configuration the engine executes. The General profile is the homepage default; every workload page shows its own profile.

Workload profileWeights
General (default)Reliability 33% · Latency 27% · Structured output 20% · Cost 20%
AgentsTool calling 31% · Structured output 25% · Reliability 19% · Latency 13% · Cost 13%
CodingTool calling 38% · Structured output 25% · Latency 19% · Cost 19%
RealtimeLatency 56% · Reliability 17% · Throughput 17% · Cost 11%
VoiceLatency 60% · Reliability 15% · Throughput 15% · Cost 10%
RAGStructured output 33% · Cost 27% · Latency 20% · Reliability 20%
Structured outputStructured output 55% · Reliability 15% · Latency 15% · Cost 15%
BatchThroughput 38% · Cost 38% · Reliability 25%
Long contextCost 36% · Throughput 36% · Reliability 29%

The deterministic pipeline that consumes these weights — market-reference normalization, optimization tilt and confidence shading — is specified in Scoring and Provider aggregation.

Changelog

VersionDateChanges
0.13-devSep 4, 2026The score's latency dimension is now the completed exchange, not the first token (operator-approved with a production before/after). The site was publishing two definitions of speed at once: every Fastest crown had already been re-based to end-to-end latency (entry below), while the score's latency dimension still read TTFT p50 alone — so the homepage could name Cerebras the fastest provider at 0.53 s and simultaneously score its latency below a provider averaging 4.4 s end to end, because generation speed reached the score through no dimension at all (throughput carries no weight in the General profile). Crown and score now read the same value from the same function and cannot diverge again. TTFT p50 and p95 keep their own columns and definitions as reported, unscored metrics. Effect at adoption across the eleven ranked providers: Cerebras 97.7 to 99.0 (#2 to #1), Groq 97.0 to 98.3, OpenAI 90.6 to 93.5 (#8 to #4), Anthropic 82.9 to 87.7 (#11 to #9), Mistral 98.5 to 99.0 (#1 to #2), Scaleway 94.3 to 90.5, Fireworks 92.3 to 88.4. Providers whose strength was a fast first token followed by slow generation fall; providers that deliver the whole answer quickly rise.
0.13-devAug 27, 2026The Fastest crown is re-based from bare TTFT to the completed benchmarked exchange (median TTFT + the normalized 400-token response at sustained throughput — the same E2E the table publishes). A provider's speed is a curve, not a scalar; a benchmark's "fast" must be indexed to its declared workload. TTFT remains a first-class column, and first-token behaviour remains visible per path.
0.13-devAug 27, 2026Scoring v3 — the provider Overall becomes Serving Quality: the provider effect measured within the same model (the only place it is identifiable), over serving dimensions only. Price leaves the Overall and becomes the Value column (serving quality per effective dollar, percentile display; exact prices in Model × Provider mode). Coverage stops multiplying the score and is displayed as information. Providers with no shared models (frontier labs) are flagged non-comparable and unranked on this axis rather than scored against a field they cannot be compared to; their paths remain fully measured in Model × Provider mode, whose score keeps the deployment semantics (absolute anchors, cost included) under its own name. Prompted by external reviewers unable to reconstruct the ranking from the visible columns — the Overall is now reconstructible from what the row shows.
0.12-devAug 24, 2026Route policies applied and declared precision published. A gateway path's routing policy now shapes the wire: its declared request parameters (OpenRouter provider-sort for cheapest/fastest/throughput) merge into every request, warm-up included, and a non-default policy that declares no request parameters refuses to run rather than measure default routing under its label — previously policies were validated at registration but never sent, so every 'cheapest' or 'fastest' label would have measured the same default routing. Provider-declared quantization (Nebius publishes it per model) is synced from discovery into the registry as model variants and shown as a chip (e.g. FP8) on the affected rows; undeclared precision remains unknown and uninferred. Retired measurement paths now render their stored retirement reason instead of the generic 'not yet measured' state. Counts labelled 'runs' that were actually eligible-result counts now say results; the partner report gained the same per-path exclusion accounting the dossiers publish.
0.12-devAug 24, 2026Publication discipline. Market-event shifts now require persistence: an INCIDENT must hold across two complete days measured against the preceding stable day; the partial current day never participates; per-day sample floors match the published gates (rates 12, latency 10) — one flipped 6-case sweep can no longer headline as a provider incident. Re-aggregation purges a day's orphaned metric rows when a suite version is retired, instead of letting a dead regime's numbers age out of the window. The immutable daily dataset snapshot freezes at the UTC day close, not at whichever mid-day cycle ran first. Provider sparklines come from the provider's best-evidenced path, chosen deterministically; the displayed E2E is the measured window value when present (the derived fallback mixes runs and is labelled as derived); the cited scoring version is the one for the displayed vantage.
0.12-devAug 24, 2026Robust market normalization (operator-approved with a production before/after). Score ranges now come from the reference population's winsorized p5–p95 band with published minimum spans per dimension, instead of raw min–max: one 8-second reasoning path or one $30/M price no longer defines the scale every other path is measured against (outliers clamp at the band edge), and a saturated dimension cannot turn sampling noise into double-digit swings. The one-benchmark invariant is untouched — same reference on every page. Winner badges now enforce the long-published confidence gate: low-confidence paths and unmeasured fields cannot take a #1 crown.
0.12-devAug 24, 2026Comparability hardening. (1) A per-request cache-busting nonce (unique system line, case prompts untouched) removes the structural TTFT advantage of providers with automatic prompt caching. (2) Request-shape quirks render as visible row properties ('default temp', 'prompt-only JSON') instead of prose footnotes. (3) Gateway paths record the per-request resolved upstream on every result and run — a routed row's numbers now name the mixture they came from. (4) Model-identity evidence: registering a path without citing where the provider model id was verified leaves the alias PROVISIONAL instead of silently CONFIRMED, and catalogue matching prefers the base model id over decorated variants. Also: two more misregistered Together paths corrected (llama-3.3-70b moved to the catalogue-verified serverless Turbo id; qwen3-235b retired — no serverless deployment exists), and clean correctness sweeps were run against every path whose window held only excluded transport failures, restoring measured tool-calling and structured-output values for all active providers. The rate gate was recalibrated to two full sweeps of the smallest suite (12) after the 20-result gate briefly blanked columns whose active-version history was three sweeps deep.
0.12-devAug 23, 2026Reasoning-aware timing and physical throughput (validators 1.3). TTFT is now the first streamed delta of ANY channel — reasoning included — so reasoning models are comparable instead of carrying their whole thinking phase as 'latency' (GLM-4.7 published 11.5 s TTFT and scored 0 on the latency dimension, dragging every other path's scale with it). The wait to first visible output is captured separately on every run and will be published as its own column. Throughput is provider-reported completion tokens over the full delta span — the same population on both sides; without reported usage there is no sample (SSE chunk counts undercount batched tokens and were the silent fallback), and burst-delivery artifacts (sub-250 ms spans, speeds above 5,000 tok/s — production had stored samples up to 9.5M tok/s) are discarded at capture and excluded retroactively from aggregates. A 120 s wall-clock bound closes the loophole where a stream emitting one frame every 59 s ran forever. Whitespace-only deltas and empty tool-call scaffolds no longer stop the first-token clock.
0.12-devAug 23, 2026Window statistics are now what the site publishes. A materialised rolling-window table replaces 'the most recent daily row' as the source of every pair metric: percentiles are computed over the whole 7-day window (winsorized within the window, as this page states), rates use window denominators, and the window is exactly 7 days. Sample-size gates apply before publication (p50 ≥ 10 samples, p95 ≥ 30, p99 ≥ 100, rates ≥ 12 — two full sweeps of the smallest suite, throughput ≥ 10); below a gate the surface shows '—'. Confidence and run counts are per path from that path's own window evidence — a new path no longer inherits its provider's pooled counts — and run counts include only canonical, finished runs. Suite provenance lists every suite behind a window's numbers, including the health suites behind TTFT and reliability. Previously each published metric was its most recent daily value — a ~31-sample day for percentiles, a 6–7-case sweep for rates — wearing the window's label.
0.12-devAug 23, 2026Failure-attribution corrections from the adversarial accuracy audit (scoring 0.3.0, validators 1.3). (1) Throttles (429) and balance exhaustion (402/insufficient-quota) are benchmark-client faults, excluded from every published denominator: they reflect the operator account's tier and funding, not provider serving — the audit measured Groq at 52.5% published reliability against 100% infrastructure-only, entirely from our on_demand tier's token-per-minute limits. Dossiers surface throttle counts so the signal stays visible without becoming a score. (2) The client-fault classifier now covers misregistered paths (nonexistent model ids, non-serverless deployments) and more request-shape families; three misregistered paths were retired with published exclusion reasons. (3) Hitting the per-case token ceiling no longer lowers reliability: it is a generation outcome that fails only the case's own suite metric; a short-but-complete answer is a content miss, not a truncation; JSON cut mid-object by the ceiling classifies as truncation, not as a JSON-capability failure. (4) Stored verdicts from before these rules were re-classified under them, with the correction counted per provider and rule. (5) Every provider now gets a default inter-request gap, Retry-After is honored, and a per-provider lock serializes the fleet to one suite per provider at a time.
0.11-devAug 22, 2026Effective workload cost implemented as the score's cost dimension: the workload's published input size plus the normalized 400-token response, at both declared prices (USD-converted), entering a score only when both prices are recorded. This is the formula originally published and withdrawn earlier today as unimplemented; it is now the code the engine executes. Provider-level impact at adoption: ≤0.1 points, no rank changes. Measured per-path token mixes (making reasoning burn visible in cost) remain a planned versioned change.
0.11-devAug 22, 2026Methodology audit corrections — the page now states exactly what the pipeline executes. Transport corrected to HTTP/1.1 (HTTP/2 had been published in error; enabling it would change measured latencies and arrives only as a versioned change). The unimplemented per-batch DNS-pinning claim removed. Timeout semantics stated precisely: 60 s per network operation. Throttled (429) responses documented as always counting against reliability. The reliability exclusion list reduced to the client-fault classes the pipeline actually expresses (auth/account failures, request-shape quirks, invalid case/validator configuration; aborted runs are never written). The score's cost dimension documented as the declared output price — the effective-workload-cost formula was never implemented and is withdrawn. Confidence window stated as 7 days. Per-provider request-parameter quirks disclosed (temperature omitted where rejected; prompt-only structured output on Anthropic). Also published: the never-unpriced invariant — the scheduler refuses to plan any path without a current declared price, gaps surface every operations cycle, and operator-recorded prices are re-verified after 30 days.
0.11-devAug 22, 2026One benchmark, one score: score normalization now runs against a fixed reference population — every measured path in the index — instead of the candidate set left by the page's filters. The 1.5% cross-region deployment-fit multiplier is removed from the score (cross-region execution stays visible as an explicit trade-off line), and provider coverage shading divides by the best coverage across the measured index instead of within the filtered scope. The same provider therefore reads the same score on the global ranking, on /europe and in every dossier; filters change who is listed, never the numbers. Weights renormalize over the dimensions actually measured for a pair, and only measured values shape a normalization range, so fixture values never leak into a published score. Previously the same provider could score differently per page because normalization ranges, the fit multiplier and the coverage denominator all depended on which competitors were in scope.
0.11-devAug 21, 2026Cross-currency price comparability: non-USD declared prices (EUR providers) convert to USD at the stored ECB daily reference rate for tables, rankings, scores and charts; declared figures and the exact rate remain visible as provenance. Previously EUR-priced providers were excluded from price comparisons entirely.
0.11-devAug 20, 2026Measurement-regime audit corrections (scoring 0.2.0, validators 1.2). (1) Published aggregates now include only currently-active suite versions: superseded versions — including their output-budget truncation artifacts and pre-quirk request-shape failures — remain stored and citable but no longer mix into current metrics. (2) HTTP 400 responses whose body rejects our request shape (unsupported parameter/value classes that later become recorded provider quirks) are classified as benchmark-client faults and excluded from every published denominator, exactly like auth failures. (3) Correctness suites moved to daily cadence for hosted-model paths (~217 eligible results per 7-day window, crossing the high-confidence threshold); frontier-lab paths keep weekly correctness (~139 per window, medium confidence) for cost reasons, stated as such.
0.10-devAug 18, 2026Measured overall score: for measured pairs the headline score is recomputed exclusively from measured dimensions — reliability (25), latency (20), structured output (15), output price (15) — with weights renormalized over what was actually observed. The Quality dimension is not measured in this version and never contributes to a measured score. Latency maps 200 ms–2 s TTFT p50 onto 100–0 linearly; output price maps $0.10–$10 per MTok onto 100–0 on a log scale. Per-provider daily budget ceilings and per-tier scheduling cadence (health daily, correctness every 2 days, heavy weekly; frontier labs weekly outside health checks) published.
0.10-devAug 17, 2026Measurement engine live: monotonic-clock runner, hash-locked suites (corpus sizes now published as measured, with repetition-based sampling), scoring pipeline with winsorized percentiles, change detector, dataset snapshots. Confidence thresholds published (high ≥200, medium ≥50 eligible results per window). First measured paths render real numbers; unmeasured rows stay labelled.
0.9-devAug 15, 2026Added P95 TTFT, E2E latency and split input/output pricing to the public table; provider aggregation formalized (coverage + freshness factors, medians); geography taxonomy expanded to six dimensions; metrics registry published.
0.8-devAug 15, 2026Provider ranking mode introduced as an aggregation layer; awards separated from provider properties; confidence gating for #1 badges.
0.7-devAug 14, 2026Deterministic decision engine: hard filters with published exclusion reasons, min–max normalization with floor 55, optimization tilt, deployment fit, confidence shading.
0.5-devAug 12, 2026Initial workload profiles (General, Agents, Structured, Realtime) and metric set; registry-driven catalogue (providers, models, endpoints) went live.
Questions about a specific number? Every score popover links here, and every research article pins the methodology version that produced it.Back to the live benchmark →