Cloud LLM benchmarks
- Models
- 283
- Providers
- 9
- Samples
- 79,358
- Median tok/s
- 31
- p90 tok/s
- 73
- Max tok/s
- 297
- Median spread
- 183%
- Median TTFT
- 1.24s
Throughput distribution · tok/s
- GPT-oss-safeguard-20b groq297 tok/s
- Qwen3.6-27B groq219 tok/s
- llama-3.1-8b groq215 tok/s
- nvidia/NVIDIA-Nemotron-3.5-Lightning deepinfra191 tok/s
- qwen-3-32b groq175 tok/s
- Google: Nano Banana (Gemini 2.5 Flash Image) google160 tok/s
- llama-3.3-70b groq147 tok/s
- gemma-4-31b cerebras132 tok/s
- nova-micro bedrock118 tok/s
- llama-4-scout groq117 tok/s
- llama-4-maverick bedrock107 tok/s
- LFM2.5-8B-A1B together107 tok/s
- qwen-2-1.5b-instruct together99 tok/s
- GPT-oss-120b fireworks99 tok/s
- llama-3.1-8b bedrock96 tok/s
- llama-4-scout bedrock96 tok/s
- qwen-3.5-35b-a3b deepinfra94 tok/s
- nova-lite bedrock91 tok/s
- nova-pro bedrock89 tok/s
- llama-3.3-70b bedrock85 tok/s
- GPT-5.1-codex-mini openai84 tok/s
- Nemotron-Lightning-3.5-30B-A3B fireworks83 tok/s
Visible-token throughput, same basis as the table · curves normalised per model · 2 models' tails run past the axis · 261 slower models not drawn, all 283 are in the table below
By provider
| Provider | Models | Median | p90 | Best | Spread | TTFT |
|---|---|---|---|---|---|---|
| groq | 6 | 175 | 297 | 297 | 152% | — |
| cerebras | 1 | 132 | 132 | 132 | 421% | 1.01 |
| bedrock | 22 | 56 | 96 | 118 | 110% | 0.38 |
| together | 21 | 52 | 79 | 107 | 212% | — |
| openai | 38 | 38 | 61 | 84 | 189% | 1.92 |
| openai via OpenRouter | 6 | 38 | 47 | 47 | 72% | 0.90 |
| fireworks | 22 | 36 | 74 | 99 | 193% | — |
| together via OpenRouter | 1 | 34 | 34 | 34 | 105% | 0.84 |
| deepinfra | 126 | 25 | 50 | 191 | 210% | — |
| deepinfra via OpenRouter | 19 | 25 | 43 | 52 | 141% | 1.01 |
| anthropic | 10 | 24 | 28 | 43 | 152% | 1.48 |
| anthropic via OpenRouter | 6 | 21 | 37 | 37 | 51% | 1.60 |
| google via OpenRouter | 1 | 3 | 3 | 3 | 40% | 1.09 |
| 4 | 2 | 160 | 160 | 145% | 0.93 |
Throughput × spread · 283 models
Full results
283 of 283 modelsThroughput over time · shared scale
Method. A cron job calls each model's live API endpoint on a schedule and records what came back. Mean, min and max are over completed samples in the selected window, and n is how many there were. A missing value means the endpoint returned an error or the model was not yet in the catalogue.
Spread is (max − min) ÷ mean, so it measures run-to-run variation rather than absolute speed. TTFT is seconds to the first visible token; runs that emit only reasoning tokens are left out of that average. Mean counts visible output tokens where a provider reports them and falls back to generated throughput where it does not, which is why Gen is higher for models that think before answering.