Cloud LLM benchmarks
- Models
- 1,432
- Providers
- 82
- Samples
- 99,634
- Median tok/s
- 20
- p90 tok/s
- 48
- Max tok/s
- 277
- Median spread
- 83%
- Median TTFT
- 2.56s
Throughput distribution · tok/s
legacy · not comparableBy provider
legacy · not comparable| Provider | Models | Median | p90 | Best | Spread | TTFT |
|---|---|---|---|---|---|---|
| aion-labs | 4 | 6 | 25 | 25 | 33% | 4.66 |
| aionlabs | 4 | 7 | 29 | 29 | 288% | 8.17 |
| akashml | 15 | 12 | 26 | 29 | 55% | 5.49 |
| alibaba | 85 | 14 | 47 | 124 | 40% | 4.77 |
| ambient | 3 | 8 | 47 | 47 | 96% | 8.33 |
| anthropic | 32 | 22 | 34 | 43 | 106% | 1.66 |
| arcee-ai | 1 | 35 | 35 | 35 | 33% | 1.58 |
| arceeai | 1 | 37 | 37 | 37 | 108% | 1.47 |
| atlas-cloud | 24 | 8 | 25 | 36 | 53% | 9.86 |
| atlascloud | 9 | 7 | 26 | 26 | 23% | 12.50 |
| azure | 17 | 24 | 38 | 39 | 61% | 1.66 |
| baidu | 13 | 21 | 26 | 29 | 50% | 2.12 |
| baseten | 8 | 22 | 52 | 52 | 45% | 2.74 |
| bedrock | 43 | 35 | 94 | 116 | 99% | 1.20 |
| cerebras | 3 | 98 | 123 | 123 | 57% | 0.82 |
| chutes | 7 | 2 | 10 | 10 | 52% | 33.70 |
| claude-on-aws | 8 | 20 | 23 | 23 | 29% | 1.76 |
| claudeplatformonaws | 6 | 20 | 25 | 25 | 93% | 1.50 |
| cloudflare | 20 | 26 | 51 | 83 | 46% | 0.86 |
| cohere | 8 | 31 | 53 | 53 | 99% | 0.77 |
| coreweave | 32 | 14 | 50 | 57 | 58% | 3.17 |
| crusoe | 11 | 30 | 70 | 89 | 75% | 1.28 |
| darkbloom | 7 | 6 | 39 | 39 | 30% | 11.40 |
| decart | 6 | 4 | 20 | 20 | 10% | 17.90 |
| deepinfra | 260 | 22 | 44 | 162 | 140% | 1.53 |
| deepseek | 3 | 8 | 11 | 11 | 94% | 8.65 |
| digitalocean | 26 | 6 | 27 | 74 | 86% | 10.10 |
| fireworks | 27 | 32 | 72 | 98 | 164% | 3.63 |
| friendli | 7 | 29 | 71 | 71 | 134% | 3.23 |
| gmicloud | 23 | 13 | 20 | 26 | 43% | 5.85 |
| 56 | 15 | 67 | 159 | 77% | 2.18 | |
| groq | 12 | 75 | 232 | 277 | 69% | 0.87 |
| inception | 3 | 61 | 62 | 62 | 107% | 1.45 |
| inceptron | 8 | 6 | 29 | 29 | 112% | 7.34 |
| io-net | 3 | 6 | 28 | 28 | 94% | 11.10 |
| ionet | 1 | 24 | 24 | 24 | 131% | 2.03 |
| liquid | 1 | 1 | 1 | 1 | 44% | 71.20 |
| makora | 1 | 4 | 4 | 4 | 0% | 17.10 |
| mancer | 7 | 13 | 45 | 45 | 22% | 4.35 |
| mancer2 | 2 | 16 | 35 | 35 | 61% | 0.71 |
| mara | 3 | 34 | 87 | 87 | 120% | 4.30 |
| minimax | 13 | 14 | 35 | 35 | 91% | 5.52 |
| mistral | 30 | 42 | 54 | 68 | 52% | 0.98 |
| modal | 2 | 23 | 24 | 24 | 0% | 2.78 |
| modelrun | 5 | 27 | 97 | 97 | 72% | 2.31 |
| moonshotai | 4 | 5 | 14 | 14 | 15% | 11.30 |
| morph | 9 | 18 | 29 | 29 | 81% | 0.99 |
| nebius | 21 | 22 | 41 | 66 | 74% | 2.48 |
| nex-agi | 2 | 17 | 18 | 18 | 61% | 3.32 |
| nexagi | 2 | 19 | 26 | 26 | 0% | 1.96 |
| nextbit | 16 | 14 | 31 | 36 | 18% | 2.39 |
| novita | 100 | 14 | 40 | 54 | 41% | 2.74 |
| nvidia | 1 | 13 | 13 | 13 | 367% | 9.00 |
| open-inference | 2 | 2 | 25 | 25 | 0% | 0.98 |
| openai | 68 | 34 | 47 | 80 | 125% | 1.48 |
| openinference | 2 | 2 | 5 | 5 | 1% | 0.78 |
| parasail | 60 | 25 | 44 | 74 | 71% | 1.65 |
| perceptron | 2 | 27 | 30 | 30 | 15% | 0.88 |
| perplexity | 9 | 24 | 35 | 35 | 91% | 2.13 |
| phala | 21 | 5 | 28 | 52 | 73% | 8.75 |
| poolside | 5 | 21 | 39 | 39 | 79% | 3.38 |
| reka | 5 | 23 | 40 | 40 | 41% | 2.42 |
| relace | 4 | 25 | 65 | 65 | 58% | 0.95 |
| sail-research | 5 | 14 | 50 | 50 | 74% | 6.30 |
| sailresearch | 2 | 6 | 17 | 17 | 88% | 3.37 |
| sakanaai | 2 | 7 | 19 | 19 | 30% | 3.32 |
| sambanova | 7 | 39 | 55 | 55 | 50% | 2.35 |
| sambanova-turbo | 1 | 53 | 53 | 53 | 57% | 1.13 |
| seed | 4 | 2 | 9 | 9 | 18% | 6.75 |
| siliconflow | 47 | 10 | 23 | 31 | 64% | 9.83 |
| stealth | 2 | 4 | 5 | 5 | 11% | 12.40 |
| streamlake | 41 | 11 | 35 | 39 | 55% | 6.30 |
| tencent | 1 | 2 | 2 | 2 | 28% | 30.20 |
| together | 50 | 31 | 75 | 97 | 123% | 2.86 |
| upstage | 4 | 27 | 34 | 34 | 87% | 1.46 |
| venice | 40 | 8 | 34 | 47 | 58% | 6.71 |
| wafer | 2 | 9 | 10 | 10 | 23% | 6.39 |
| xai | 14 | 4 | 35 | 35 | 98% | 15.60 |
| xiaomi | 4 | 8 | 14 | 14 | 20% | 3.23 |
| z-ai | 11 | 3 | 9 | 12 | 67% | 25.60 |
| zai | 2 | 2 | 14 | 14 | 0% | 2.68 |
Throughput × spread · 1432 models
mean tok/s →↑ spread %
Full results
legacy · not comparable1432 of 1432 modelsState
min·mean·max
Trend
Throughput over time · shared scale
legacy · not comparableMethod. A cron job calls each model's live API endpoint on a schedule and records what came back. Mean, min and max are over completed samples in the selected window, and n is how many there were. A missing value means the endpoint returned an error or the model was not yet in the catalogue.
Spread is (max − min) ÷ mean, so it measures run-to-run variation rather than absolute speed. TTFT is seconds to the first visible token; runs that emit only reasoning tokens are left out of that average. Mean counts visible output tokens where a provider reports them and falls back to generated throughput where it does not, which is why Gen is higher for models that think before answering.