Cloud LLM benchmarks
- Models
- 781
- Providers
- 70
- Samples
- 48,566
- Median tok/s
- 16
- p90 tok/s
- 47
- Max tok/s
- 116
- Median spread
- 64%
- Median TTFT
- 3.32s
Delivered TPS · 64-token, end-to-end
64 visible tokens ÷ time to the 64th, reasoning counted as time, per endpointLoading Delivered TPS
Throughput distribution · tok/s
legacy · not comparableThroughput × spread · 781 models
legacy · not comparable- steady ≤110% 562
- variable 110–180% 176
- unstable >180% 43
- 5 above 266% drawn at the ceiling
By provider
legacy · not comparable| Provider | Models | Median | p90 | Best | Spread | TTFT |
|---|---|---|---|---|---|---|
| aionlabs | 4 | 6 | 25 | 25 | 33% | 4.66 |
| akashml | 8 | 9 | 29 | 29 | 47% | 5.49 |
| alibaba | 49 | 13 | 47 | 58 | 33% | 5.53 |
| ambient | 3 | 8 | 47 | 47 | 96% | 8.33 |
| anthropic | 10 | 20 | 34 | 40 | 107% | 1.72 |
| arceeai | 1 | 35 | 35 | 35 | 33% | 1.58 |
| atlascloud | 25 | 10 | 25 | 36 | 55% | 9.86 |
| azure | 10 | 23 | 27 | 39 | 27% | 1.48 |
| baidu | 8 | 13 | 29 | 29 | 50% | 2.32 |
| baseten | 6 | 22 | 49 | 49 | 45% | 1.71 |
| bedrock | 29 | 43 | 96 | 116 | 100% | 0.55 |
| cerebras | 2 | 86 | 107 | 107 | 51% | 0.59 |
| chutes | 4 | 2 | 10 | 10 | 52% | 19.50 |
| claudeonaws | 9 | 20 | 23 | 23 | 33% | 2.17 |
| cloudflare | 14 | 26 | 52 | 83 | 49% | 0.86 |
| cohere | 4 | 34 | 53 | 53 | 41% | 0.58 |
| coreweave | 20 | 25 | 50 | 57 | 58% | 2.00 |
| crusoe | 7 | 34 | 70 | 70 | 75% | 1.28 |
| darkbloom | 4 | 5 | 61 | 61 | 12% | 11.40 |
| decart | 3 | 4 | 20 | 20 | 141% | 17.90 |
| deepinfra | 69 | 14 | 36 | 69 | 98% | 2.81 |
| deepseek | 4 | 3 | 11 | 11 | 94% | 8.65 |
| digitalocean | 17 | 7 | 26 | 74 | 86% | 18.20 |
| fireworks | 4 | 17 | 23 | 23 | 76% | 3.63 |
| friendli | 5 | 24 | 71 | 71 | 116% | 3.31 |
| gmicloud | 16 | 10 | 20 | 21 | 29% | 6.12 |
| 32 | 12 | 69 | 92 | 38% | 3.48 | |
| groq | 8 | 62 | 110 | 110 | 46% | 0.89 |
| inception | 2 | 33 | 62 | 62 | 0% | 1.45 |
| inceptron | 5 | 11 | 29 | 29 | 165% | 6.74 |
| ionet | 3 | 6 | 28 | 28 | 94% | 10.90 |
| makora | 1 | 4 | 4 | 4 | 0% | 17.10 |
| mancer | 7 | 13 | 45 | 45 | 22% | 4.35 |
| mara | 4 | 34 | 84 | 84 | 31% | 0.83 |
| minimax | 8 | 8 | 35 | 35 | 56% | 5.52 |
| mistral | 15 | 42 | 49 | 52 | 122% | 1.00 |
| modal | 1 | 23 | 23 | 23 | 76% | 3.42 |
| modelrun | 3 | 27 | 97 | 97 | 73% | 2.31 |
| moonshotai | 3 | 5 | 14 | 14 | 131% | 13.70 |
| morph | 6 | 18 | 29 | 29 | 72% | 0.83 |
| nebius | 14 | 22 | 55 | 66 | 41% | 2.56 |
| nexagi | 2 | 17 | 19 | 19 | 49% | 3.07 |
| nextbit | 10 | 14 | 31 | 31 | 12% | 2.39 |
| novita | 67 | 13 | 44 | 54 | 45% | 3.38 |
| openai | 25 | 23 | 45 | 52 | 150% | 2.07 |
| openinference | 2 | 2 | 25 | 25 | 10% | 0.99 |
| parasail | 40 | 24 | 46 | 59 | 72% | 1.81 |
| perceptron | 1 | 27 | 27 | 27 | 15% | 0.98 |
| perplexity | 5 | 24 | 33 | 33 | 65% | 2.13 |
| phala | 19 | 5 | 28 | 52 | 73% | 8.75 |
| poolside | 2 | 21 | 39 | 39 | 49% | 2.91 |
| reka | 4 | 6 | 40 | 40 | 0% | 1.59 |
| relace | 2 | 25 | 65 | 65 | 0% | 0.83 |
| sailresearch | 5 | 14 | 35 | 35 | 143% | 6.30 |
| sakanaai | 2 | 7 | 19 | 19 | 30% | 3.32 |
| sambanova | 6 | 39 | 55 | 55 | 24% | 1.57 |
| sambanovaturbo | 1 | 53 | 53 | 53 | 57% | 1.12 |
| seed | 2 | 2 | 9 | 9 | 41% | 6.75 |
| siliconflow | 36 | 8 | 23 | 31 | 76% | 7.95 |
| stealth | 1 | 5 | 5 | 5 | 67% | 12.40 |
| streamlake | 25 | 11 | 37 | 38 | 57% | 6.30 |
| tencent | 1 | 2 | 2 | 2 | 28% | 30.20 |
| together | 21 | 16 | 38 | 58 | 65% | 3.86 |
| upstage | 2 | 27 | 34 | 34 | 37% | 1.38 |
| venice | 31 | 8 | 41 | 47 | 58% | 6.71 |
| wafer | 2 | 9 | 10 | 10 | 23% | 6.39 |
| xai | 7 | 4 | 35 | 35 | 98% | 18.00 |
| xiaomi | 2 | 8 | 13 | 13 | 31% | 3.51 |
| zai | 11 | 3 | 9 | 12 | 67% | 25.60 |
Full results
legacy · not comparable781 models of 1,442 lanes · 360 folded behind their current laneThroughput over time · shared scale
legacy · not comparableMethod. A cron job calls each model's live API endpoint on a schedule and records what came back. Mean, min and max are over completed samples in the selected window, and n is how many there were. A missing value means the endpoint returned an error or the model was not yet in the catalogue.
Lanes. A lane is one way of reaching a model: a provider, a catalogue id and a transport (a direct key, or served through OpenRouter). Since 2026-08-16 most measurement runs through OpenRouter; the direct lanes it replaced are retired, not deleted. By default the page shows live lanes and one row per model at each provider, spoken for by the lane measured most recently; All shows every lane in the window, retired ones tagged off.
Spread is (max − min) ÷ mean, so it measures run-to-run variation rather than absolute speed. TTFT is seconds to the first visible token; runs that emit only reasoning tokens are left out of that average. Mean counts visible output tokens where a provider reports them and falls back to generated throughput where it does not, which is why Gen is higher for models that think before answering. Decode is the generation rate once tokens are flowing, with time to first token and reasoning excluded; only streams that deliver a few tokens per chunk can carry it, so a dash means the stream was too coarse to time.