llm-benchmarks
githubdrose.io

baseten Provider Benchmarks


Measured over the last 30 days

Models Tracked
8
Avg Tokens / Second
17.67
Avg Time to First Token (ms)
9710.17
Last Updated
Sep 2, 2026

All baseten models

8 tracked · fastest first
ModelAvg Toks/SecMinMaxAvg TTF (ms)
gpt-oss-120b51.9039.7067.001000.00
GLM 5.332.4027.0037.401460.00
GLM 5.229.3022.2035.501710.00
Nemotron 3 Ultra23.605.2550.503030.00
Nemotron 3 Ultra21.909.8329.102740.00
Inkling13.8013.8013.804380.00
Inkling9.808.8110.906050.00
DeepSeek V4 Flash 07317.073.2610.7010800.00

Frequently Asked Questions

Which baseten model is fastest?
Based on recent tests, gpt-oss-120b shows the highest average throughput among tracked baseten models.
How many recent measurements feed this dashboard?
This provider summary aggregates 707 individual prompts measured across 620 monitoring runs over the past month.