Measured over the last 30 days
- Models Tracked
- 47
- Avg Tokens / Second
- 15.37
- Avg Time to First Token (ms)
- 10911.85
- Last Updated
- Sep 1, 2026
All siliconflow models
47 tracked · fastest first| Model | Avg Toks/Sec | Min | Max | Avg TTF (ms) |
|---|---|---|---|---|
| Gemma 4 26B A4B | 31.30 | 7.40 | 47.60 | 2680.00 |
| Hunyuan A13B Instruct | 29.10 | 5.12 | 42.90 | 1870.00 |
| Hunyuan A13B Instruct | 28.40 | 5.25 | 37.80 | 3760.00 |
| Gemma 4 31B | 22.90 | 18.40 | 27.60 | 1950.00 |
| Qwen3 Coder 30B A3B Instruct | 22.90 | 12.70 | 27.30 | 1680.00 |
| Qwen3 30B A3B Instruct 2507 | 20.20 | 18.20 | 21.60 | 1690.00 |
| GLM 5.2 | 19.90 | 17.30 | 21.90 | 2410.00 |
| Qwen3 Coder 30B A3B Instruct | 19.50 | 6.90 | 31.40 | 2160.00 |
| Nex-N2-Pro | 18.40 | 15.80 | 22.30 | 3110.00 |
| DeepSeek V3 0324 | 17.30 | 6.51 | 23.20 | 1260.00 |
| Qwen3 30B A3B Instruct 2507 | 16.50 | 12.30 | 22.80 | 1500.00 |
| DeepSeek V3 0324 | 16.30 | 12.80 | 17.90 | 1530.00 |
| GLM 4.5 Air | 15.90 | 8.15 | 31.00 | 4930.00 |
| DeepSeek V3.2 | 15.30 | 13.70 | 16.80 | 1830.00 |
| DeepSeek V3.2 Exp | 14.80 | 13.00 | 16.40 | 2000.00 |
| Step 3.5 Flash | 14.70 | 14.70 | 14.70 | 3450.00 |
| DeepSeek V3.1 Terminus | 14.10 | 13.30 | 14.70 | 1280.00 |
| DeepSeek V4 Flash 0731 | 13.80 | 4.47 | 18.30 | 5830.00 |
| DeepSeek V3.1 Terminus | 13.40 | 7.35 | 15.60 | 1720.00 |
| Step 3.5 Flash | 12.30 | 10.70 | 14.50 | 4480.00 |
| DeepSeek V3.1 | 12.10 | 4.24 | 16.20 | 4000.00 |
| Qwen3 VL 30B A3B Instruct | 11.40 | 2.14 | 23.80 | 10100.00 |
| Qwen3 VL 30B A3B Thinking | 11.00 | 10.70 | 11.30 | 5330.00 |
| Kimi K2.7 Code | 10.40 | 7.64 | 13.50 | 5670.00 |
| Qwen3.5-122B-A10B | 7.87 | 3.85 | 17.20 | 13600.00 |
| DeepSeek V4 Flash 0423 | 7.25 | 3.80 | 13.20 | 9830.00 |
| gpt-oss-120b | 7.11 | 3.12 | 14.10 | 10100.00 |
| DeepSeek V4 Pro 0423 | 5.20 | 1.60 | 8.78 | 16800.00 |
| Qwen3.5-122B-A10B | 5.01 | 3.23 | 6.97 | 12800.00 |
| Kimi K2.5 | 4.85 | 2.60 | 8.91 | 16100.00 |
| Kimi K2.6 | 4.79 | 3.08 | 6.12 | 13300.00 |
| gpt-oss-20b | 4.58 | 1.59 | 7.96 | 18900.00 |
| MiniMax M2.5 | 4.09 | 3.18 | 5.28 | 15000.00 |
| Kimi K2.5 | 4.05 | 1.51 | 7.45 | 17500.00 |
| Qwen3.8 2.4T A95B | 3.88 | 3.88 | 3.88 | 14800.00 |
| Qwen3 32B | 3.43 | 1.26 | 6.81 | 25700.00 |
| Qwen3.6 35B A3B | 3.39 | 1.23 | 4.80 | 26100.00 |
| Qwen3.5-9B | 2.97 | 1.51 | 4.75 | 24600.00 |
| GLM 5 | 2.80 | 2.68 | 2.95 | 22000.00 |
| R1 0528 | 2.62 | 1.35 | 7.83 | 22100.00 |
| Qwen3.5-35B-A3B | 2.57 | 2.12 | 3.01 | 24600.00 |
| GLM 5.1 | 2.51 | 2.11 | 3.15 | 25200.00 |
| R1 0528 | 2.37 | 1.88 | 2.59 | 24100.00 |
| Qwen3.5-9B | 2.23 | 1.90 | 2.66 | 28200.00 |
| Qwen3.5-27B | 1.59 | 1.34 | 2.34 | 40200.00 |
| Qwen3.6 27B | 1.57 | 1.40 | 1.74 | 39900.00 |
| Qwen3.5-27B | 1.56 | 1.37 | 1.85 | 39800.00 |
Frequently Asked Questions
Which siliconflow model is fastest?
Based on recent tests, Gemma 4 26B A4B shows the highest average throughput among tracked siliconflow models.
How many recent measurements feed this dashboard?
This provider summary aggregates 7968 individual prompts measured across 5977 monitoring runs over the past month.