Cloud LLM benchmarks
- Models
- 807
- Providers
- 10
- Samples
- 105,395
- Median tok/s
- 26
- p90 tok/s
- 53
- Max tok/s
- 283
- Median spread
- 131%
- Median TTFT
- 1.62s
Throughput distribution · tok/s
legacy · not comparableBy provider
legacy · not comparable| Provider | Models | Median | p90 | Best | Spread | TTFT |
|---|---|---|---|---|---|---|
| anthropic | 17 | 25 | 34 | 45 | 132% | 1.54 |
| bedrock | 21 | 60 | 96 | 117 | 108% | 0.41 |
| cerebras | 1 | 136 | 136 | 136 | 409% | 1.02 |
| deepinfra | 147 | 25 | 49 | 171 | 197% | 0.98 |
| fireworks | 28 | 34 | 78 | 97 | 191% | 2.28 |
| 5 | 3 | 158 | 158 | 181% | 1.09 | |
| groq | 6 | 145 | 283 | 283 | 137% | 0.99 |
| openai | 49 | 38 | 59 | 83 | 197% | 1.74 |
| openrouter | 504 | 23 | 47 | 122 | 101% | 1.69 |
| together | 29 | 47 | 81 | 102 | 182% | 1.18 |
Throughput × spread · 807 models
mean tok/s →↑ spread %
Full results
legacy · not comparable807 of 807 modelsState
min·mean·max
Trend
Throughput over time · shared scale
legacy · not comparableMethod. A cron job calls each model's live API endpoint on a schedule and records what came back. Mean, min and max are over completed samples in the selected window, and n is how many there were. A missing value means the endpoint returned an error or the model was not yet in the catalogue.
Spread is (max − min) ÷ mean, so it measures run-to-run variation rather than absolute speed. TTFT is seconds to the first visible token; runs that emit only reasoning tokens are left out of that average. Mean counts visible output tokens where a provider reports them and falls back to generated throughput where it does not, which is why Gen is higher for models that think before answering.