Cloud LLM benchmarks
- Models
- 802
- Providers
- 10
- Samples
- 99,668
- Median tok/s
- 26
- p90 tok/s
- 55
- Max tok/s
- 283
- Median spread
- 125%
- Median TTFT
- 1.59s
Throughput distribution · tok/s
legacy · not comparableBy provider
legacy · not comparable| Provider | Models | Median | p90 | Best | Spread | TTFT |
|---|---|---|---|---|---|---|
| anthropic | 17 | 25 | 34 | 45 | 132% | 1.54 |
| bedrock | 21 | 60 | 96 | 117 | 111% | 0.41 |
| cerebras | 1 | 136 | 136 | 136 | 409% | 1.02 |
| deepinfra | 147 | 25 | 49 | 171 | 197% | 0.98 |
| fireworks | 28 | 34 | 78 | 98 | 191% | 2.28 |
| 5 | 3 | 159 | 159 | 180% | 1.09 | |
| groq | 6 | 146 | 283 | 283 | 136% | 0.99 |
| openai | 49 | 38 | 61 | 83 | 195% | 1.72 |
| openrouter | 499 | 23 | 46 | 122 | 76% | 1.69 |
| together | 29 | 47 | 81 | 102 | 182% | 1.18 |
Throughput × spread · 802 models
mean tok/s →↑ spread %
Full results
legacy · not comparable802 of 802 modelsState
min·mean·max
Trend
Throughput over time · shared scale
legacy · not comparableMethod. A cron job calls each model's live API endpoint on a schedule and records what came back. Mean, min and max are over completed samples in the selected window, and n is how many there were. A missing value means the endpoint returned an error or the model was not yet in the catalogue.
Spread is (max − min) ÷ mean, so it measures run-to-run variation rather than absolute speed. TTFT is seconds to the first visible token; runs that emit only reasoning tokens are left out of that average. Mean counts visible output tokens where a provider reports them and falls back to generated throughput where it does not, which is why Gen is higher for models that think before answering.