Cloud LLM benchmarks
- Models
- 244
- Providers
- 9
- Samples
- 66,403
- Median tok/s
- 32
- p90 tok/s
- 74
- Max tok/s
- 296
- Median spread
- 141%
- Median TTFT
- 1.46s
Throughput distribution · tok/s
- GPT-oss-safeguard-20b groq296 tok/s
- llama-3.1-8b groq223 tok/s
- Qwen3.6-27B groq217 tok/s
- qwen-3-32b groq174 tok/s
- Google: Nano Banana (Gemini 2.5 Flash Image) google159 tok/s
- llama-3.3-70b groq152 tok/s
- gemma-4-31b cerebras151 tok/s
- nova-micro bedrock122 tok/s
- llama-4-scout groq114 tok/s
- LFM2.5-8B-A1B together110 tok/s
- qwen-3.5-35b-a3b deepinfra110 tok/s
- qwen-2-1.5b-instruct together109 tok/s
- llama-4-maverick bedrock107 tok/s
- GPT-oss-120b fireworks99 tok/s
- llama-3.1-8b bedrock96 tok/s
- llama-4-scout bedrock95 tok/s
- nova-pro bedrock92 tok/s
- nova-lite bedrock91 tok/s
- GPT-5.1-codex-mini openai84 tok/s
- GPT-5 Nano openai84 tok/s
- llama-3.3-70b bedrock84 tok/s
- mistral-7b bedrock82 tok/s
Visible-token throughput, same basis as the table · curves normalised per model · 1 model's tail runs past the axis · 222 slower models not drawn, all 244 are in the table below
By provider
| Provider | Models | Median | p90 | Best | Spread | TTFT |
|---|---|---|---|---|---|---|
| groq | 6 | 174 | 296 | 296 | 145% | — |
| cerebras | 1 | 151 | 151 | 151 | 134% | 0.62 |
| together | 20 | 50 | 79 | 110 | 188% | — |
| bedrock | 23 | 50 | 96 | 122 | 110% | 0.40 |
| fireworks | 19 | 40 | 74 | 99 | 155% | — |
| openai | 38 | 38 | 60 | 84 | 153% | 1.86 |
| deepinfra | 123 | 27 | 48 | 110 | 146% | — |
| anthropic | 10 | 24 | 30 | 43 | 106% | 1.47 |
| 4 | 2 | 159 | 159 | 112% | 0.95 |
Throughput × spread · 244 models
Full results
244 of 244 modelsThroughput over time · shared scale
Method. A cron job calls each model's live API endpoint on a schedule and records what came back. Mean, min and max are over completed samples in the selected window, and n is how many there were. A missing value means the endpoint returned an error or the model was not yet in the catalogue.
Spread is (max − min) ÷ mean, so it measures run-to-run variation rather than absolute speed. TTFT is seconds to the first visible token; runs that emit only reasoning tokens are left out of that average. Mean counts visible output tokens where a provider reports them and falls back to generated throughput where it does not, which is why Gen is higher for models that think before answering.