Measured over the last 30 days
- Models Tracked
- 13
- Avg Tokens / Second
- 19.00
- Avg Time to First Token (ms)
- 5947.69
- Last Updated
- Sep 23, 2026
All together models
13 live · one row per model · fastest first · 2 folded behind their current lane| Model | Avg Toks/Sec | Min | Max | Avg TTF (ms) |
|---|---|---|---|---|
| GLM 5.2 | 56.90 | 49.50 | 64.30 | 740.00 |
| Llama 3.3 70B Instruct | 34.40 | 9.68 | 55.90 | 1310.00 |
| gpt-oss-120b | 33.70 | 20.10 | 47.10 | 1590.00 |
| MiniMax M3 | 18.00 | 3.69 | 37.60 | 4200.00 |
| Gemma 4 31B | 16.80 | 1.54 | 41.60 | 6390.00 |
| Qwen3.8 2.4T A95B | 15.60 | 15.60 | 15.60 | 3840.00 |
| Inkling Small | 15.30 | 15.30 | 15.30 | 3860.00 |
| DeepSeek V4 Flash 0731 | 14.10 | 8.99 | 19.10 | 4140.00 |
| Muse Glimmer 30B | 13.00 | 10.30 | 15.00 | 4430.00 |
| Inkling | 11.40 | 10.20 | 12.50 | 5370.00 |
| gpt-oss-20b | 8.84 | 8.84 | 8.84 | 6850.00 |
| Qwen3.5-9B | 4.84 | 4.84 | 4.84 | 12800.00 |
| DeepSeek V4 Pro 0813 | 4.15 | 1.38 | 8.57 | 21800.00 |
Frequently Asked Questions
Which together model is fastest?
Based on recent tests, GLM 5.2 shows the highest average throughput among tracked together models.
How many recent measurements feed this dashboard?
This provider summary aggregates 2334 individual prompts measured across 2064 monitoring runs over the past month.