Measured over the last 30 days
- Models Tracked
- 21
- Avg Tokens / Second
- 21.98
- Avg Time to First Token (ms)
- 9078.65
- Last Updated
- Sep 1, 2026
All nebius models
21 tracked · fastest first| Model | Avg Toks/Sec | Min | Max | Avg TTF (ms) |
|---|---|---|---|---|
| Nemotron 3 Nano 30B A3B | 65.90 | 65.90 | 65.90 | 780.00 |
| Hermes 4 70B | 55.00 | 50.90 | 57.80 | 570.00 |
| Qwen3 30B A3B Instruct 2507 | 41.30 | 37.30 | 46.20 | 510.00 |
| Hermes 4 70B | 36.90 | 8.44 | 61.00 | 1240.00 |
| gpt-oss-120b | 35.80 | 25.10 | 42.10 | 1570.00 |
| Qwen3 235B A22B Instruct 2507 | 30.40 | 5.75 | 51.50 | 3680.00 |
| Hermes 4 405B | 29.20 | 24.10 | 36.00 | 530.00 |
| Hermes 4 405B | 26.40 | 6.72 | 33.40 | 760.00 |
| Qwen2.5 VL 72B Instruct | 25.00 | 6.41 | 35.20 | 1020.00 |
| Qwen2.5 VL 72B Instruct | 24.50 | 9.16 | 32.10 | 600.00 |
| Nemotron 3 Super | 21.90 | 20.00 | 24.40 | 2800.00 |
| Qwen3 30B A3B Instruct 2507 | 19.60 | 9.47 | 29.00 | 2480.00 |
| Gemma 3 27B | 18.60 | 7.69 | 40.10 | 2560.00 |
| Gemma 3 27B | 18.10 | 5.47 | 48.70 | 1360.00 |
| Nemotron 3 Super | 13.70 | 4.89 | 22.40 | 7780.00 |
| Qwen3 Next 80B A3B Thinking | 7.82 | 6.80 | 8.83 | 7660.00 |
| Llama 3.3 70B Instruct | 6.59 | 6.59 | 6.59 | 5740.00 |
| MiniMax M2.5 | 6.27 | 3.57 | 7.75 | 10900.00 |
| Qwen3 32B | 3.98 | 2.09 | 6.64 | 18200.00 |
| Llama 3.3 70B Instruct | 1.46 | 1.46 | 1.46 | 11700.00 |
| GLM 5.1 | 1.45 | 1.16 | 1.66 | 43400.00 |
Frequently Asked Questions
Which nebius model is fastest?
Based on recent tests, Nemotron 3 Nano 30B A3B shows the highest average throughput among tracked nebius models.
How many recent measurements feed this dashboard?
This provider summary aggregates 2706 individual prompts measured across 2602 monitoring runs over the past month.