Measured over the last 30 days
- Models Tracked
- 123
- Avg Tokens / Second
- 28.80
- Avg Time to First Token (ms)
- 0.00
- Last Updated
- Aug 9, 2026
| Model | Avg Toks/Sec | Min | Max | Avg TTF (ms) |
|---|---|---|---|---|
| qwen-3.5-35b-a3b | 110.00 | 10.70 | 156.00 | 0.00 |
| GPT-oss-120b-Turbo | 65.80 | 5.07 | 180.00 | 0.00 |
| Nemotron-3-Nano-30B-A3B | 64.50 | 2.87 | 93.10 | 0.00 |
| Qwen3-30B-A3B | 64.00 | 26.60 | 98.20 | 0.00 |
| NVIDIA-Nemotron-Nano-9B-v2 | 62.90 | 3.66 | 103.00 | 0.00 |
| qwen-3.5-122b-a10b | 61.90 | 33.50 | 80.00 | 0.00 |
| MiMo-V2.5-Pro | 60.00 | 3.96 | 120.00 | 0.00 |
| Inkling | 58.60 | 5.36 | 102.00 | 0.00 |
| GPT-oss-20b | 58.10 | 17.10 | 112.00 | 0.00 |
| Inkling-Small | 55.90 | 16.80 | 87.40 | 0.00 |
| MiniMax-M2.7-Turbo | 50.40 | 4.02 | 91.20 | 0.00 |
| nvidia/Llama-3.1-Nemotron-70B-Instruct | 50.10 | 45.90 | 54.40 | 0.00 |
| Llama-3.3-Nemotron-Super-49B-v1.5 | 47.90 | 8.19 | 73.80 | 0.00 |
| qwen-3.5-27b | 47.60 | 10.00 | 70.90 | 0.00 |
| qwen-3.5-397b-a17b | 47.20 | 26.20 | 97.00 | 0.00 |
| GLM-4.7-Flash | 46.60 | 2.23 | 129.00 | 0.00 |
| Qwen3.6-27B | 45.70 | 2.65 | 74.10 | 0.00 |
| Qwen3-Next-80B-A3B-Instruct | 45.30 | 10.00 | 77.70 | 0.00 |
| phi-4 | 45.20 | 3.92 | 53.00 | 0.00 |
| nemotron-3-ultra-550b-a55b | 44.40 | 9.13 | 71.40 | 0.00 |
| Nemotron-3-Nano-Omni-30B-A3B-Reasoning | 44.00 | 6.46 | 79.90 | 0.00 |
| llama-3.2-90b | 43.10 | 19.00 | 57.60 | 0.00 |
| Qwen 2.5 Coder 32B | 42.70 | 19.90 | 62.60 | 0.00 |
| gemma-4-E4B-it | 42.40 | 12.60 | 66.20 | 0.00 |
| llama-4-maverick | 42.00 | 5.08 | 64.20 | 0.00 |
| Qwen3-32B | 41.00 | 13.40 | 63.60 | 0.00 |
| NVIDIA-Nemotron-Nano-12B-v2-VL | 40.40 | 2.31 | 78.30 | 0.00 |
| NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | 39.90 | 1.61 | 67.40 | 0.00 |
| NVIDIA-Nemotron-3-Super-120B-A12B | 39.00 | 11.90 | 64.70 | 0.00 |
| llama-3-8b | 36.40 | 17.40 | 69.40 | 0.00 |
| mixtral-8x7b | 36.20 | 17.00 | 52.00 | 0.00 |
| DeepSeek-V3.1-Terminus | 36.10 | 5.00 | 71.70 | 0.00 |
| qwen-3.5-2b | 36.00 | 5.76 | 89.00 | 0.00 |
| GLM-5.2 | 36.00 | 11.30 | 78.00 | 0.00 |
| Mistral-Small-24B-Instruct-2501 | 35.90 | 17.30 | 56.30 | 0.00 |
| gemma-4-31B-it-turbo | 35.50 | 6.65 | 59.60 | 0.00 |
| qwen-3.5-4b | 35.40 | 8.40 | 86.90 | 0.00 |
| Kimi-K2.7-Code | 34.50 | 4.49 | 87.90 | 0.00 |
| gemini-2.5-flash | 34.30 | 14.50 | 42.00 | 0.00 |
| GLM-4.6 | 33.80 | 11.70 | 65.00 | 0.00 |
| Qwen3.6-35B-A3B | 33.50 | 1.24 | 75.00 | 0.00 |
| gemini-3.1-flash-lite | 33.30 | 11.70 | 43.10 | 0.00 |
| gemma-3-12b-it | 33.00 | 10.70 | 45.30 | 0.00 |
| MiniMax-M2.1 | 32.70 | 23.40 | 40.50 | 0.00 |
| Hy3 | 32.70 | 0.99 | 62.50 | 0.00 |
| claude-4-opus | 32.00 | 25.40 | 37.30 | 0.00 |
| Ornith-1.0-35B | 32.00 | 3.43 | 89.40 | 0.00 |
| claude-opus-4-7 | 31.10 | 23.90 | 36.10 | 0.00 |
| claude-sonnet-5 | 30.60 | 19.00 | 36.60 | 0.00 |
| L3.3-70B-Euryale-v2.3 | 30.40 | 17.90 | 34.60 | 0.00 |
| Qwen3-Coder-480B-A35B-Instruct | 30.30 | 2.24 | 54.60 | 0.00 |
| llama-3.1-70b | 30.00 | 2.77 | 45.00 | 0.00 |
| claude-haiku-4.5 | 29.70 | 12.90 | 37.30 | 0.00 |
| claude-opus-5 | 29.70 | 28.90 | 30.50 | 0.00 |
| llama-3-70b | 29.70 | 2.82 | 44.70 | 0.00 |
| Qwen3-Coder-480B-A35B-Instruct-Turbo | 29.60 | 1.76 | 57.40 | 0.00 |
| qwen-3.5-0.8b | 29.20 | 3.48 | 96.00 | 0.00 |
| MiMo-V2.5 | 28.60 | 7.42 | 55.60 | 0.00 |
| llama-2-70b | 28.40 | 2.04 | 44.10 | 0.00 |
| gemini-3.5-flash | 27.90 | 10.20 | 34.80 | 0.00 |
| claude-opus-4-8 | 27.60 | 16.10 | 33.00 | 0.00 |
| Qwen/Qwen2.5-VL-32B-Instruct | 27.40 | 23.80 | 30.90 | 0.00 |
| GLM-5.1 | 27.40 | 10.30 | 50.40 | 0.00 |
| GLM-4.7 | 26.90 | 8.25 | 55.30 | 0.00 |
| MiniMax-M2 | 25.20 | 22.00 | 28.30 | 0.00 |
| gemma-3-4b-it | 25.20 | 10.10 | 51.30 | 0.00 |
| DeepSeek-V4-Flash | 24.70 | 5.37 | 38.30 | 0.00 |
| DeepSeek-V4-Flash-0731 | 24.30 | 23.40 | 25.20 | 0.00 |
| Qwen3-Max | 23.40 | 17.00 | 29.40 | 0.00 |
| Qwen3-VL-30B-A3B-Instruct | 23.40 | 3.30 | 38.50 | 0.00 |
| MiniMax-M2.5 | 22.70 | 19.90 | 27.30 | 0.00 |
| mistral-7b | 22.30 | 1.38 | 53.30 | 0.00 |
| devstral-small | 22.10 | 2.93 | 51.50 | 0.00 |
| gemini-3.1-pro | 21.80 | 17.60 | 26.80 | 0.00 |
| Mistral-Small-3.2-24B-Instruct-2506 | 21.70 | 3.30 | 51.90 | 0.00 |
| GPT-oss-120b | 21.20 | 7.69 | 46.40 | 0.00 |
| Mistral-Nemo-Instruct-2407 | 21.00 | 3.36 | 69.00 | 0.00 |
| gemma-3-27b-it | 20.70 | 4.28 | 31.00 | 0.00 |
| gemma-4-26B-A4B-it | 20.70 | 5.24 | 34.00 | 0.00 |
| Kimi-K2.5 | 20.10 | 1.15 | 78.00 | 0.00 |
| GLM-5 | 19.60 | 4.15 | 47.40 | 0.00 |
| qwen-3-235b | 19.10 | 2.41 | 34.60 | 0.00 |
| Qwen3-Max-Thinking | 18.90 | 12.90 | 24.80 | 0.00 |
| GLM-4.6V | 18.70 | 1.37 | 38.40 | 0.00 |
| qwen-3-14b | 18.00 | 2.81 | 54.70 | 0.00 |
| claude-fable-5 | 17.90 | 11.30 | 23.60 | 0.00 |
| DeepSeek-V4-Pro | 17.60 | 1.52 | 41.00 | 0.00 |
| qwen-2.5-72b | 17.30 | 2.32 | 32.00 | 0.00 |
| deepseek-ai/DeepSeek-R1-0528-Turbo | 17.30 | 9.49 | 29.60 | 0.00 |
| claude-4-sonnet | 17.20 | 10.00 | 23.70 | 0.00 |
| DeepSeek-R1-0528 | 17.20 | 9.94 | 30.80 | 0.00 |
| claude-sonnet-4.6 | 16.90 | 7.23 | 21.30 | 0.00 |
| DeepSeek-V3 | 16.80 | 1.76 | 44.80 | 0.00 |
| claude-3-7-sonnet-latest | 16.70 | 8.69 | 19.80 | 0.00 |
| llama-3.1-8b | 16.70 | 8.71 | 28.20 | 0.00 |
| llama-3.1-405b | 16.60 | 3.01 | 24.00 | 0.00 |
| Kimi-K2.5-Turbo | 16.30 | 2.24 | 64.70 | 0.00 |
| deepseek-ai/DeepSeek-R1-Distill-Llama-70B | 16.10 | 11.80 | 20.40 | 0.00 |
| Kimi-K2-Instruct-0905 | 16.10 | 2.74 | 62.90 | 0.00 |
| Olmo-3.1-32B-Instruct | 16.00 | 1.47 | 62.90 | 0.00 |
| MiniMax-M2.7 | 14.80 | 7.75 | 33.40 | 0.00 |
| Kimi-K2-Thinking | 14.20 | 2.07 | 62.00 | 0.00 |
| Kimi-K2.6 | 14.00 | 2.19 | 34.70 | 0.00 |
| llama-3.3-70b | 13.20 | 3.68 | 27.60 | 0.00 |
| Llama-3.3-70B-Instruct-Turbo | 13.20 | 1.20 | 24.90 | 0.00 |
| Qwen3-VL-235B-A22B-Instruct | 13.10 | 1.05 | 40.20 | 0.00 |
| gemma-4-31B-it | 12.80 | 1.06 | 51.20 | 0.00 |
| llama-3.2-11b | 12.60 | 1.18 | 56.70 | 0.00 |
| llama-3.2-1b | 12.40 | 1.20 | 50.10 | 0.00 |
| microsoft/WizardLM-2-8x22B | 12.40 | 5.84 | 18.90 | 0.00 |
| deepseek-v3.2 | 12.10 | 1.02 | 32.30 | 0.00 |
| llama-3.2-3b | 11.60 | 1.13 | 40.20 | 0.00 |
| MiniMax-M3 | 10.50 | 0.87 | 61.10 | 0.00 |
| DeepSeek-V3.1 | 8.80 | 2.28 | 17.60 | 0.00 |
| Qwen3-235B-A22B-Thinking-2507 | 7.87 | 1.38 | 28.30 | 0.00 |
| deepseek-ai/DeepSeek-OCR | 7.15 | 4.29 | 10.00 | 0.00 |
| PaddlePaddle/PaddleOCR-VL-0.9B | 5.90 | 4.26 | 7.53 | 0.00 |
| allenai/olmOCR-2-7B-1025 | 5.70 | 3.08 | 8.32 | 0.00 |
| Qwen3.8-Max | 4.72 | 2.66 | 6.78 | 0.00 |
| Llama-Guard-4-12B | 2.76 | 0.82 | 4.25 | 0.00 |
| Seed-2.0-pro | 2.31 | 1.54 | 3.06 | 0.00 |
| Qwen3.7-Max | 2.29 | 2.04 | 2.63 | 0.00 |
| Seed-2.0-code | 0.64 | 0.53 | 0.81 | 0.00 |
Frequently Asked Questions
Which deepinfra model is fastest?
Based on recent tests, qwen-3.5-35b-a3b shows the highest average throughput among tracked deepinfra models.
How many recent measurements feed this dashboard?
This provider summary aggregates 4284 individual prompts measured across 2726 monitoring runs over the past month.