AgentsBenchmarked for tool reliability, structured output, multi-step execution, latency and production reliability.Tool reliability · Structured output · Multi-step execution · Production reliabilityCodingBenchmarked for tool-driven editing loops, structured patches, long-context handling and latency.Tool reliability · Structured edits · Long context · LatencyRealtimeBenchmarked for time-to-first-token, p95 tail latency, streaming stability and availability.TTFT · p95 latency · Streaming stability · AvailabilityVoiceBenchmarked for first-token latency, inter-token pacing, stream integrity and availability under load.TTFT · Inter-token latency · Stream integrity · AvailabilityRAGBenchmarked for structured grounding, context handling, caching behaviour and cost at retrieval scale.Structured output · Context handling · Prompt caching · CostStructured outputBenchmarked for JSON validity, schema adherence, enum and nested-object correctness and truncation behaviour.Schema adherence · JSON validity · Reliability · LatencyBatchBenchmarked for throughput, cost per token, batch API support and sustained reliability.Throughput · Cost · Batch API · ReliabilityLong contextBenchmarked for context handling, caching, throughput on large prompts and price at high input volume.Context handling · Prompt caching · Input price · Throughput