Effective output throughput
UTC · pointer/touch · Left/Right and Home/End keys
A time series chart. Left and Right arrow keys move the shared active UTC bucket. Home moves to the first bucket and End moves to the latest.
ModelsLoading selection… Filter by provider or model
| Model | Route | TPS | Duration | Output | Source | Current status | Measurement observed at |
|---|
How these measurements work
Aggregate tokens_per_second is the arithmetic mean of output_tokens divided by the arithmetic mean of duration_seconds; it is not the arithmetic mean of per-attempt output_tokens / duration_seconds ratios. Effective output throughput includes proxy, network, provider queue, reasoning, and generation time; it is not raw decoder speed.
Each successful observation summarizes three sequential attempts. The four Codex subscription GPT routes are post-response bounded at the declared 4,096 billed-output-token envelope. The other six routes enforce a 512-token request output cap. Different output limits and routes affect comparability; this is observability, not a model-quality ranking.
Only validated schema-v6 observations are shown. Current failures, missed samples, budget pauses, and never-observed routes remain distinct. A prior successful measurement is historical evidence, not proof that a route works now.