Cloud LLM benchmarks
- Models
- 283
- Providers
- 9
- Samples
- 80,632
- Median tok/s
- 30
- p90 tok/s
- 74
- Max tok/s
- 301
- Median spread
- 188%
- Median TTFT
- 1.26s
Delivered throughput · tok/s
64 visible tokens ÷ time to the 64thNo Delivered TPS data yet
Throughput distribution · tok/s
By provider
| Provider | Models | Median | p90 | Best | Spread | TTFT |
|---|---|---|---|---|---|---|
| groq | 6 | 181 | 301 | 301 | 134% | — |
| cerebras | 1 | 133 | 133 | 133 | 418% | 1.01 |
| bedrock | 22 | 56 | 96 | 116 | 110% | 0.39 |
| together | 21 | 54 | 83 | 105 | 223% | — |
| openai | 38 | 38 | 61 | 84 | 199% | 1.96 |
| openai via OpenRouter | 6 | 37 | 47 | 47 | 71% | 0.92 |
| fireworks | 22 | 36 | 75 | 99 | 192% | — |
| together via OpenRouter | 1 | 34 | 34 | 34 | 116% | 0.85 |
| deepinfra | 126 | 25 | 50 | 171 | 211% | — |
| anthropic | 10 | 25 | 28 | 44 | 152% | 1.47 |
| deepinfra via OpenRouter | 19 | 24 | 42 | 55 | 147% | 1.00 |
| anthropic via OpenRouter | 6 | 21 | 35 | 35 | 78% | 1.70 |
| google via OpenRouter | 1 | 3 | 3 | 3 | 40% | 1.09 |
| 4 | 2 | 160 | 160 | 145% | 0.94 |
Throughput × spread · 283 models
mean tok/s →↑ spread %
Full results
283 of 283 modelsState
min·mean·max
Trend
Throughput over time · shared scale
Method. A cron job calls each model's live API endpoint on a schedule and records what came back. Mean, min and max are over completed samples in the selected window, and n is how many there were. A missing value means the endpoint returned an error or the model was not yet in the catalogue.
Spread is (max − min) ÷ mean, so it measures run-to-run variation rather than absolute speed. TTFT is seconds to the first visible token; runs that emit only reasoning tokens are left out of that average. Mean counts visible output tokens where a provider reports them and falls back to generated throughput where it does not, which is why Gen is higher for models that think before answering.