Measured over the last 30 days
- Models Tracked
- 8
- Avg Tokens / Second
- 17.67
- Avg Time to First Token (ms)
- 9710.17
- Last Updated
- Sep 2, 2026
All baseten models
8 tracked · fastest first| Model | Avg Toks/Sec | Min | Max | Avg TTF (ms) |
|---|---|---|---|---|
| gpt-oss-120b | 51.90 | 39.70 | 67.00 | 1000.00 |
| GLM 5.3 | 32.40 | 27.00 | 37.40 | 1460.00 |
| GLM 5.2 | 29.30 | 22.20 | 35.50 | 1710.00 |
| Nemotron 3 Ultra | 23.60 | 5.25 | 50.50 | 3030.00 |
| Nemotron 3 Ultra | 21.90 | 9.83 | 29.10 | 2740.00 |
| Inkling | 13.80 | 13.80 | 13.80 | 4380.00 |
| Inkling | 9.80 | 8.81 | 10.90 | 6050.00 |
| DeepSeek V4 Flash 0731 | 7.07 | 3.26 | 10.70 | 10800.00 |
Frequently Asked Questions
Which baseten model is fastest?
Based on recent tests, gpt-oss-120b shows the highest average throughput among tracked baseten models.
How many recent measurements feed this dashboard?
This provider summary aggregates 707 individual prompts measured across 620 monitoring runs over the past month.