Cloud LLM benchmarks
- Models
- 507
- Providers
- 10
- Samples
- 83,291
- Median tok/s
- 30
- p90 tok/s
- 61
- Max tok/s
- 283
- Median spread
- 133%
- Median TTFT
- 1.36s
Delivered throughput · tok/s
64 visible tokens ÷ time to the 64thNo Delivered TPS data yet
Throughput distribution · tok/s
By provider
| Provider | Models | Median | p90 | Best | Spread | TTFT |
|---|---|---|---|---|---|---|
| groq | 6 | 170 | 283 | 283 | 136% | — |
| cerebras | 1 | 136 | 136 | 136 | 409% | 1.02 |
| bedrock | 21 | 61 | 96 | 117 | 110% | 0.41 |
| together | 23 | 59 | 81 | 102 | 232% | — |
| openai | 38 | 39 | 63 | 86 | 202% | 1.97 |
| fireworks | 25 | 38 | 78 | 99 | 194% | — |
| openai via OpenRouter | 6 | 37 | 47 | 47 | 71% | 0.92 |
| openrouter | 209 | 30 | 53 | 87 | 70% | 1.35 |
| deepinfra via OpenRouter | 19 | 26 | 41 | 55 | 154% | 0.98 |
| deepinfra | 126 | 25 | 50 | 171 | 211% | — |
| anthropic | 10 | 25 | 29 | 45 | 163% | 1.45 |
| anthropic via OpenRouter | 7 | 22 | 34 | 34 | 101% | 1.54 |
| fireworks via OpenRouter | 3 | 16 | 42 | 42 | 139% | 2.28 |
| together via OpenRouter | 6 | 15 | 48 | 48 | 116% | 1.18 |
| groq via OpenRouter | 2 | 11 | 25 | 25 | 0% | 0.99 |
| google via OpenRouter | 1 | 3 | 3 | 3 | 40% | 1.09 |
| 4 | 3 | 160 | 160 | 145% | 0.95 |
Throughput × spread · 507 models
mean tok/s →↑ spread %
Full results
507 of 507 modelsState
min·mean·max
Trend
Throughput over time · shared scale
Method. A cron job calls each model's live API endpoint on a schedule and records what came back. Mean, min and max are over completed samples in the selected window, and n is how many there were. A missing value means the endpoint returned an error or the model was not yet in the catalogue.
Spread is (max − min) ÷ mean, so it measures run-to-run variation rather than absolute speed. TTFT is seconds to the first visible token; runs that emit only reasoning tokens are left out of that average. Mean counts visible output tokens where a provider reports them and falls back to generated throughput where it does not, which is why Gen is higher for models that think before answering.