llm-benchmarks
githubdrose.io

groq Provider Benchmarks


Measured over the last 30 days

Models Tracked
7
Avg Tokens / Second
74.14
Avg Time to First Token (ms)
1008.57
Last Updated
Sep 23, 2026

All groq models

7 live · one row per model · fastest first · 1 folded behind their current lane
groq site
ModelAvg Toks/SecMinMaxAvg TTF (ms)
Llama 3.1 8B Instruct110.0099.20121.00470.00
Llama 3.3 70B Instruct86.8060.40103.00530.00
Llama 4 Scout85.3085.3085.30620.00
gpt-oss-20b73.5052.50110.00850.00
gpt-oss-120b72.7040.1099.50780.00
gpt-oss-safeguard-20b61.6034.30110.001030.00
MiniMax M2.729.1010.6045.602780.00

Frequently Asked Questions

Which groq model is fastest?
Based on recent tests, Llama 3.1 8B Instruct shows the highest average throughput among tracked groq models.
How many recent measurements feed this dashboard?
This provider summary aggregates 748 individual prompts measured across 714 monitoring runs over the past month.