llm-benchmarks
githubdrose.io

open-inference Provider Benchmarks


Measured over the last 30 days

Models Tracked
2
Avg Tokens / Second
17.58
Avg Time to First Token (ms)
12570.32
Last Updated
Sep 1, 2026

All open-inference models

2 tracked · fastest first
ModelAvg Toks/SecMinMaxAvg TTF (ms)
Gemma 4 31B24.605.5038.202580.00
DeepSeek V4 Flash 07311.591.591.59980.00

Frequently Asked Questions

Which open-inference model is fastest?
Based on recent tests, Gemma 4 31B shows the highest average throughput among tracked open-inference models.
How many recent measurements feed this dashboard?
This provider summary aggregates 425 individual prompts measured across 371 monitoring runs over the past month.