Structured Output: Trust, but Validate
Schema-conformance rates across measured providers, one mechanical validator, no partial credit.
- Measured structured-output validity spans 91.7%–100.0% across providers.
- Cerebras leads; Mistral trails — a pipeline parsing its output re-requests or crashes on the difference.
Conformance, measured
Each structured-output case demands schema-valid JSON and validates the result mechanically — no partial credit. The native-schema column records each provider's declared enforcement mode; whether it separates the table is for the measured column to say. Anthropic runs this suite prompt-only by recorded request quirk.
| Provider | Validity | Native schema mode |
|---|---|---|
| Cerebras | 100.0% | no |
| Groq | 100.0% | yes |
| Scaleway | 100.0% | yes |
| OpenAI | 100.0% | yes |
| Cortecs | 98.3% | yes |
| Together | 96.8% | yes |
| OpenRouter | 96.4% | yes |
| Mistral | 91.7% | yes |
- Treat schema validity as an SLO with a measured baseline per provider — and always validate downstream regardless.
InferenceBench (2026). Structured Output: Trust, but Validate. InferenceBench Research.