Tool Calling Is an Infrastructure Property
The same model's function-calling reliability varies wildly with the provider serving it.
- On DeepSeek V4 Flash, measured tool-call reliability spans 0.0%–100.0% across 5 providers serving identical weights.
- OpenRouter leads the measured set on this model; the weakest path fails roughly 100 of every 100 tool calls.
- Identical weights, different serving stacks, different tool-call outcomes: the split is infrastructure-side.
Where tool calls break
A failed tool call is rarely the model refusing — it is malformed JSON, truncated arguments, or a serving stack that mangles the function-calling protocol under load. That is infrastructure behaviour, and it only becomes visible when the same model is measured across providers under an identical harness.
| Provider | Tool reliability |
|---|---|
| OpenRouter | 100.0% |
| Cortecs | 100.0% |
| Together | 100.0% |
| Scaleway | 100.0% |
| Fireworks | 0.0% |
- For agent stacks, provider selection is a first-order reliability decision — test the exact path, not the model in the abstract.
InferenceBench (2026). Tool Calling Is an Infrastructure Property. InferenceBench Research.