llm-benchmarks
githubdrose.io

Llama 4 Maverick Benchmarks


Measured on deepinfra

2 runs
deepinfra site
Avg Tokens / Second
44.70
Decode Tokens / Second
Avg Time to First Token (ms)
610.00
Runs Analysed
2
Last Updated
Sep 15, 2026, 12:02 PM

Throughput distribution

Throughput over time

Llama 4 Maverickdeepinfra45tok/s

Generated throughput, including reasoning tokens — so a thinking model's line does not collapse; the table ranks on visible tokens · one cell per model, showing its best-covered provider · models with at least 4 samples in the window first, fastest of those at the top · shared vertical scale, 0 to 53 tok/s (99th percentile) · dashed rule is the model's own mean · 1 cell carry fewer than 4 samples

Every lane measuring this model

2 measured · live first · a lane is a provider, a catalogue id and a transport
ProviderLaneAvg Toks/SecMinMaxAvg TTF (ms)n
deepinfravia OpenRouter44.7038.6050.90610.002
deepinfradirect41.7026.9056.50700.003

Frequently Asked Questions

How fast is Llama 4 Maverick?
The latest rolling average throughput is 44.70 tokens per second with an average time to first token of 610.00 ms across 2 recent runs.
How often are these benchmarks updated?
Benchmarks refresh automatically whenever the monitoring cron runs. The most recent run completed on Sep 15, 2026, 12:02 PM.