Measured over the last 30 days
- Models Tracked
- 32
- Avg Tokens / Second
- 19.19
- Avg Time to First Token (ms)
- 8827.61
- Last Updated
- Sep 1, 2026
All coreweave models
32 tracked · fastest first| Model | Avg Toks/Sec | Min | Max | Avg TTF (ms) |
|---|---|---|---|---|
| Llama 3.1 8B Instruct | 57.10 | 44.60 | 78.00 | 700.00 |
| Granite 4.1 8B | 55.50 | 33.10 | 74.50 | 500.00 |
| Granite 4.1 8B | 50.60 | 19.00 | 110.00 | 870.00 |
| Llama 3.3 70B Instruct | 50.40 | 44.80 | 55.60 | 560.00 |
| Qwen3 30B A3B Instruct 2507 | 49.60 | 35.90 | 64.40 | 480.00 |
| Llama 3.1 70B Instruct | 47.00 | 38.70 | 51.20 | 580.00 |
| DeepSeek V3.1 | 38.40 | 31.90 | 45.30 | 630.00 |
| Kimi K2.7 Code | 35.80 | 19.10 | 47.30 | 1820.00 |
| MiniMax M3 | 28.30 | 9.85 | 61.10 | 2140.00 |
| DeepSeek V4 Flash 0423 | 25.00 | 19.60 | 37.60 | 920.00 |
| Gemma 4 31B | 24.90 | 9.52 | 32.10 | 660.00 |
| Gemma 4 31B | 24.60 | 13.90 | 30.20 | 620.00 |
| Nemotron 3.5 Lightning | 22.40 | 16.40 | 26.70 | 2720.00 |
| MiniMax M3 | 18.40 | 6.93 | 24.60 | 3920.00 |
| Kimi K2.6 | 16.10 | 6.77 | 27.30 | 4970.00 |
| gpt-oss-120b | 13.60 | 4.20 | 22.90 | 2630.00 |
| gpt-oss-120b | 13.50 | 4.77 | 26.00 | 3170.00 |
| Granite 4.2 8B | 13.40 | 13.40 | 13.40 | 3760.00 |
| gpt-oss-20b | 12.40 | 1.89 | 34.80 | 5440.00 |
| GLM 5.2 | 11.80 | 5.80 | 19.10 | 6390.00 |
| gpt-oss-20b | 11.10 | 6.38 | 13.80 | 5830.00 |
| Qwen3.5-35B-A3B | 11.10 | 9.83 | 12.10 | 5580.00 |
| gpt-oss-120b | 9.83 | 4.72 | 16.70 | 5300.00 |
| Auto Router (Beta) | 9.19 | 9.19 | 9.19 | 6660.00 |
| DeepSeek V4 Flash 0731 | 4.74 | 1.83 | 7.53 | 19300.00 |
| GLM 5.2 | 4.53 | 4.53 | 4.53 | 10700.00 |
| Qwen3.6 35B A3B | 4.36 | 4.09 | 4.80 | 14200.00 |
| Nemotron 3.5 Lightning | 4.00 | 4.00 | 4.00 | 15600.00 |
| GLM 5.2 | 3.08 | 3.08 | 3.08 | 1690.00 |
| DeepSeek V4 Pro 0813 | 2.60 | 2.60 | 2.60 | 17300.00 |
| Kimi K2.6 | 1.71 | 1.71 | 1.71 | 35400.00 |
| DeepSeek V4 Flash 0731 | 1.57 | 1.57 | 1.57 | 42500.00 |
Frequently Asked Questions
Which coreweave model is fastest?
Based on recent tests, Llama 3.1 8B Instruct shows the highest average throughput among tracked coreweave models.
How many recent measurements feed this dashboard?
This provider summary aggregates 3553 individual prompts measured across 3095 monitoring runs over the past month.