Measured over the last 30 days
- Models Tracked
- 68
- Avg Tokens / Second
- 17.78
- Avg Time to First Token (ms)
- 9621.32
- Last Updated
- Sep 24, 2026
All deepinfra models
68 live · one row per model · fastest first · 30 folded behind their current lane| Model | Avg Toks/Sec | Min | Max | Avg TTF (ms) |
|---|---|---|---|---|
| Nemotron 3.5 Lightning | 96.40 | 96.40 | 96.40 | 510.00 |
| Qwen3 Next 80B A3B Instruct | 55.70 | 55.70 | 55.70 | 720.00 |
| Llama 3 8B Lunaris | 54.40 | 48.30 | 62.00 | 440.00 |
| Llama 4 Maverick | 44.70 | 38.60 | 50.90 | 610.00 |
| Gemma 3 12B | 41.40 | 26.80 | 56.00 | 560.00 |
| Mistral Small 3 | 40.90 | 39.30 | 42.40 | 470.00 |
| Phi 4 | 40.20 | 31.20 | 49.10 | 550.00 |
| Qwen3 Coder 480B A35B | 37.80 | 29.60 | 46.00 | 510.00 |
| Gemma 4 31B | 37.30 | 5.37 | 73.70 | 930.00 |
| MythoMax 13B | 36.90 | 35.50 | 38.40 | 480.00 |
| Llama 3.1 Euryale 70B v2.2 | 34.00 | 33.70 | 34.30 | 470.00 |
| Mistral Small 3.2 24B | 33.00 | 15.30 | 50.60 | 550.00 |
| Hermes 3 70B Instruct | 30.90 | 30.70 | 31.20 | 520.00 |
| Llama 3.1 70B Instruct | 30.40 | 27.30 | 33.50 | 720.00 |
| Nemotron 3 Nano 30B A3B | 29.40 | 29.40 | 29.40 | 1400.00 |
| gpt-oss-120b | 28.00 | 6.71 | 49.20 | 2830.00 |
| Llama 4 Scout | 27.40 | 27.40 | 27.40 | 850.00 |
| Granite 4.2 8B | 25.90 | 1.73 | 44.50 | 3240.00 |
| DeepSeek V4 Flash 0423 | 25.10 | 23.70 | 27.00 | 770.00 |
| Mistral Nemo | 24.20 | 24.00 | 24.40 | 560.00 |
| Llama 3.1 8B Instruct | 24.00 | 16.70 | 31.30 | 750.00 |
| Gemma 4 26B A4B | 20.80 | 11.70 | 33.90 | 870.00 |
| GLM 5.2 | 19.30 | 7.69 | 30.90 | 4460.00 |
| DeepSeek V3 0324 | 19.10 | 14.80 | 23.40 | 1340.00 |
| Llama 3.3 70B Instruct | 18.80 | 13.50 | 27.30 | 660.00 |
| MiMo-V2.5-Pro | 18.00 | 18.00 | 18.00 | 2300.00 |
| MiMo-V2.5 | 17.90 | 14.70 | 21.20 | 2270.00 |
| Hermes 3 405B Instruct | 17.30 | 11.20 | 23.50 | 1370.00 |
| Gemma 3 4B | 17.00 | 10.00 | 23.90 | 1630.00 |
| Gemma 3 27B | 15.80 | 5.53 | 26.20 | 6180.00 |
| Muse Glimmer 30B | 15.70 | 12.20 | 18.10 | 3600.00 |
| Qwen3 VL 30B A3B Instruct | 15.20 | 4.65 | 25.80 | 5270.00 |
| Qwen3 235B A22B Instruct 2507 | 13.60 | 5.82 | 21.20 | 1070.00 |
| DeepSeek V3.2 | 13.20 | 4.64 | 24.10 | 1530.00 |
| Qwen3 14B | 11.70 | 9.44 | 14.10 | 4950.00 |
| Qwen3.8 2.4T A95B | 11.50 | 11.50 | 11.50 | 4630.00 |
| Qwen3 VL 235B A22B Instruct | 11.30 | 11.30 | 11.30 | 1340.00 |
| gpt-oss-20b | 11.20 | 8.09 | 16.40 | 5530.00 |
| DeepSeek V3 | 10.50 | 4.16 | 26.60 | 5690.00 |
| Ling 3.0 Flash | 10.30 | 10.30 | 10.30 | 5340.00 |
| MiniMax M2.7 | 9.84 | 1.57 | 24.90 | 12100.00 |
| Qwen2.5 72B Instruct | 8.62 | 6.82 | 10.40 | 3900.00 |
| DeepSeek V3.1 | 8.44 | 7.98 | 8.90 | 920.00 |
| Kimi K2.5 | 7.87 | 7.87 | 7.87 | 7390.00 |
| Nemotron 3 Super | 7.70 | 7.13 | 8.28 | 7560.00 |
| Qwen3 30B A3B | 7.20 | 7.20 | 7.20 | 8850.00 |
| Qwen3.5-122B-A10B | 7.07 | 4.65 | 9.50 | 9860.00 |
| MiniMax M3 | 5.69 | 2.49 | 10.00 | 10700.00 |
| DeepSeek V4 Flash 0731 | 5.66 | 1.17 | 11.70 | 24300.00 |
| Llama Guard 4 12B | 5.23 | 4.59 | 5.57 | 520.00 |
| Kimi K2.7 Code | 4.74 | 2.37 | 6.89 | 13500.00 |
| Qwen3.5-35B-A3B | 4.34 | 2.91 | 5.42 | 14900.00 |
| Inkling Small | 4.28 | 3.21 | 5.55 | 15700.00 |
| GLM 4.7 Flash | 4.25 | 4.25 | 4.25 | 14500.00 |
| Qwen3.6 35B A3B | 3.12 | 0.59 | 6.10 | 44400.00 |
| Qwen3.5-27B | 2.94 | 2.56 | 3.31 | 21100.00 |
| Qwen3.8 27B | 2.90 | 0.81 | 11.90 | 46200.00 |
| GLM 5.1 | 2.67 | 0.71 | 4.38 | 36800.00 |
| R1 0528 | 2.60 | 1.69 | 3.33 | 23500.00 |
| Qwen3.5-9B | 2.54 | 2.54 | 2.54 | 24400.00 |
| DeepSeek V4.1 Flash | 2.39 | 2.39 | 2.39 | 25900.00 |
| Qwen3.5 397B A17B | 2.23 | 1.63 | 2.95 | 28300.00 |
| Kimi K2.6 | 2.10 | 1.02 | 2.93 | 33800.00 |
| GLM 5 | 2.06 | 1.54 | 2.50 | 31100.00 |
| GLM 4.7 | 1.83 | 1.58 | 2.07 | 33100.00 |
| Qwen3 32B | 1.76 | 1.66 | 1.85 | 34100.00 |
| GLM 4.6 | 1.31 | 0.95 | 1.67 | 580.00 |
| Hy3 | 1.30 | 1.30 | 1.30 | 47800.00 |
Frequently Asked Questions
Which deepinfra model is fastest?
Based on recent tests, Nemotron 3.5 Lightning shows the highest average throughput among tracked deepinfra models.
How many recent measurements feed this dashboard?
This provider summary aggregates 4178 individual prompts measured across 3555 monitoring runs over the past month.