Cloud LLM benchmarks
- Models
- 283
- Providers
- 9
- Samples
- 79,843
- Median tok/s
- 30
- p90 tok/s
- 73
- Max tok/s
- 298
- Median spread
- 188%
- Median TTFT
- 1.25s
Delivered throughput · tok/s
64 visible tokens ÷ time to the 64thNo Delivered TPS data yet
Throughput distribution · tok/s
By provider
| Provider | Models | Median | p90 | Best | Spread | TTFT |
|---|---|---|---|---|---|---|
| groq | 6 | 177 | 298 | 298 | 150% | — |
| cerebras | 1 | 133 | 133 | 133 | 418% | 1.03 |
| bedrock | 22 | 56 | 96 | 117 | 110% | 0.39 |
| together | 21 | 52 | 82 | 107 | 226% | — |
| openai | 38 | 38 | 62 | 84 | 199% | 1.93 |
| openai via OpenRouter | 6 | 37 | 47 | 47 | 71% | 0.92 |
| fireworks | 22 | 36 | 74 | 99 | 192% | — |
| together via OpenRouter | 1 | 35 | 35 | 35 | 111% | 0.77 |
| deepinfra | 126 | 25 | 50 | 171 | 211% | — |
| anthropic | 10 | 25 | 28 | 44 | 159% | 1.47 |
| deepinfra via OpenRouter | 19 | 24 | 42 | 55 | 147% | 1.00 |
| anthropic via OpenRouter | 6 | 22 | 37 | 37 | 69% | 1.70 |
| google via OpenRouter | 1 | 3 | 3 | 3 | 40% | 1.09 |
| 4 | 2 | 160 | 160 | 145% | 0.93 |
Throughput × spread · 283 models
mean tok/s →↑ spread %
Full results
283 of 283 modelsState
min·mean·max
Trend
Throughput over time · shared scale
Method. A cron job calls each model's live API endpoint on a schedule and records what came back. Mean, min and max are over completed samples in the selected window, and n is how many there were. A missing value means the endpoint returned an error or the model was not yet in the catalogue.
Spread is (max − min) ÷ mean, so it measures run-to-run variation rather than absolute speed. TTFT is seconds to the first visible token; runs that emit only reasoning tokens are left out of that average. Mean counts visible output tokens where a provider reports them and falls back to generated throughput where it does not, which is why Gen is higher for models that think before answering.