llm-benchmarks
githubdrose.io

deepinfra Provider Benchmarks


Measured over the last 30 days

Models Tracked
68
Avg Tokens / Second
17.78
Avg Time to First Token (ms)
9621.32
Last Updated
Sep 24, 2026

All deepinfra models

68 live · one row per model · fastest first · 30 folded behind their current lane
deepinfra site
ModelAvg Toks/SecMinMaxAvg TTF (ms)
Nemotron 3.5 Lightning96.4096.4096.40510.00
Qwen3 Next 80B A3B Instruct55.7055.7055.70720.00
Llama 3 8B Lunaris54.4048.3062.00440.00
Llama 4 Maverick44.7038.6050.90610.00
Gemma 3 12B41.4026.8056.00560.00
Mistral Small 340.9039.3042.40470.00
Phi 440.2031.2049.10550.00
Qwen3 Coder 480B A35B37.8029.6046.00510.00
Gemma 4 31B37.305.3773.70930.00
MythoMax 13B36.9035.5038.40480.00
Llama 3.1 Euryale 70B v2.234.0033.7034.30470.00
Mistral Small 3.2 24B33.0015.3050.60550.00
Hermes 3 70B Instruct30.9030.7031.20520.00
Llama 3.1 70B Instruct30.4027.3033.50720.00
Nemotron 3 Nano 30B A3B29.4029.4029.401400.00
gpt-oss-120b28.006.7149.202830.00
Llama 4 Scout27.4027.4027.40850.00
Granite 4.2 8B25.901.7344.503240.00
DeepSeek V4 Flash 042325.1023.7027.00770.00
Mistral Nemo24.2024.0024.40560.00
Llama 3.1 8B Instruct24.0016.7031.30750.00
Gemma 4 26B A4B20.8011.7033.90870.00
GLM 5.219.307.6930.904460.00
DeepSeek V3 032419.1014.8023.401340.00
Llama 3.3 70B Instruct18.8013.5027.30660.00
MiMo-V2.5-Pro18.0018.0018.002300.00
MiMo-V2.517.9014.7021.202270.00
Hermes 3 405B Instruct17.3011.2023.501370.00
Gemma 3 4B17.0010.0023.901630.00
Gemma 3 27B15.805.5326.206180.00
Muse Glimmer 30B15.7012.2018.103600.00
Qwen3 VL 30B A3B Instruct15.204.6525.805270.00
Qwen3 235B A22B Instruct 250713.605.8221.201070.00
DeepSeek V3.213.204.6424.101530.00
Qwen3 14B11.709.4414.104950.00
Qwen3.8 2.4T A95B11.5011.5011.504630.00
Qwen3 VL 235B A22B Instruct11.3011.3011.301340.00
gpt-oss-20b11.208.0916.405530.00
DeepSeek V310.504.1626.605690.00
Ling 3.0 Flash10.3010.3010.305340.00
MiniMax M2.79.841.5724.9012100.00
Qwen2.5 72B Instruct8.626.8210.403900.00
DeepSeek V3.18.447.988.90920.00
Kimi K2.57.877.877.877390.00
Nemotron 3 Super7.707.138.287560.00
Qwen3 30B A3B7.207.207.208850.00
Qwen3.5-122B-A10B7.074.659.509860.00
MiniMax M35.692.4910.0010700.00
DeepSeek V4 Flash 07315.661.1711.7024300.00
Llama Guard 4 12B5.234.595.57520.00
Kimi K2.7 Code4.742.376.8913500.00
Qwen3.5-35B-A3B4.342.915.4214900.00
Inkling Small4.283.215.5515700.00
GLM 4.7 Flash4.254.254.2514500.00
Qwen3.6 35B A3B3.120.596.1044400.00
Qwen3.5-27B2.942.563.3121100.00
Qwen3.8 27B2.900.8111.9046200.00
GLM 5.12.670.714.3836800.00
R1 05282.601.693.3323500.00
Qwen3.5-9B2.542.542.5424400.00
DeepSeek V4.1 Flash2.392.392.3925900.00
Qwen3.5 397B A17B2.231.632.9528300.00
Kimi K2.62.101.022.9333800.00
GLM 52.061.542.5031100.00
GLM 4.71.831.582.0733100.00
Qwen3 32B1.761.661.8534100.00
GLM 4.61.310.951.67580.00
Hy31.301.301.3047800.00

Frequently Asked Questions

Which deepinfra model is fastest?
Based on recent tests, Nemotron 3.5 Lightning shows the highest average throughput among tracked deepinfra models.
How many recent measurements feed this dashboard?
This provider summary aggregates 4178 individual prompts measured across 3555 monitoring runs over the past month.