llm-benchmarks
githubdrose.io

deepinfra Provider Benchmarks


Measured over the last 30 days

Models Tracked
123
Avg Tokens / Second
28.80
Avg Time to First Token (ms)
0.00
Last Updated
Aug 9, 2026

All deepinfra models

123 tracked · fastest first
deepinfra site
ModelAvg Toks/SecMinMaxAvg TTF (ms)
qwen-3.5-35b-a3b110.0010.70156.000.00
GPT-oss-120b-Turbo65.805.07180.000.00
Nemotron-3-Nano-30B-A3B64.502.8793.100.00
Qwen3-30B-A3B64.0026.6098.200.00
NVIDIA-Nemotron-Nano-9B-v262.903.66103.000.00
qwen-3.5-122b-a10b61.9033.5080.000.00
MiMo-V2.5-Pro60.003.96120.000.00
Inkling58.605.36102.000.00
GPT-oss-20b58.1017.10112.000.00
Inkling-Small55.9016.8087.400.00
MiniMax-M2.7-Turbo50.404.0291.200.00
nvidia/Llama-3.1-Nemotron-70B-Instruct50.1045.9054.400.00
Llama-3.3-Nemotron-Super-49B-v1.547.908.1973.800.00
qwen-3.5-27b47.6010.0070.900.00
qwen-3.5-397b-a17b47.2026.2097.000.00
GLM-4.7-Flash46.602.23129.000.00
Qwen3.6-27B45.702.6574.100.00
Qwen3-Next-80B-A3B-Instruct45.3010.0077.700.00
phi-445.203.9253.000.00
nemotron-3-ultra-550b-a55b44.409.1371.400.00
Nemotron-3-Nano-Omni-30B-A3B-Reasoning44.006.4679.900.00
llama-3.2-90b43.1019.0057.600.00
Qwen 2.5 Coder 32B42.7019.9062.600.00
gemma-4-E4B-it42.4012.6066.200.00
llama-4-maverick42.005.0864.200.00
Qwen3-32B41.0013.4063.600.00
NVIDIA-Nemotron-Nano-12B-v2-VL40.402.3178.300.00
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF1639.901.6167.400.00
NVIDIA-Nemotron-3-Super-120B-A12B39.0011.9064.700.00
llama-3-8b36.4017.4069.400.00
mixtral-8x7b36.2017.0052.000.00
DeepSeek-V3.1-Terminus36.105.0071.700.00
qwen-3.5-2b36.005.7689.000.00
GLM-5.236.0011.3078.000.00
Mistral-Small-24B-Instruct-250135.9017.3056.300.00
gemma-4-31B-it-turbo35.506.6559.600.00
qwen-3.5-4b35.408.4086.900.00
Kimi-K2.7-Code34.504.4987.900.00
gemini-2.5-flash34.3014.5042.000.00
GLM-4.633.8011.7065.000.00
Qwen3.6-35B-A3B33.501.2475.000.00
gemini-3.1-flash-lite33.3011.7043.100.00
gemma-3-12b-it33.0010.7045.300.00
MiniMax-M2.132.7023.4040.500.00
Hy332.700.9962.500.00
claude-4-opus32.0025.4037.300.00
Ornith-1.0-35B32.003.4389.400.00
claude-opus-4-731.1023.9036.100.00
claude-sonnet-530.6019.0036.600.00
L3.3-70B-Euryale-v2.330.4017.9034.600.00
Qwen3-Coder-480B-A35B-Instruct30.302.2454.600.00
llama-3.1-70b30.002.7745.000.00
claude-haiku-4.529.7012.9037.300.00
claude-opus-529.7028.9030.500.00
llama-3-70b29.702.8244.700.00
Qwen3-Coder-480B-A35B-Instruct-Turbo29.601.7657.400.00
qwen-3.5-0.8b29.203.4896.000.00
MiMo-V2.528.607.4255.600.00
llama-2-70b28.402.0444.100.00
gemini-3.5-flash27.9010.2034.800.00
claude-opus-4-827.6016.1033.000.00
Qwen/Qwen2.5-VL-32B-Instruct27.4023.8030.900.00
GLM-5.127.4010.3050.400.00
GLM-4.726.908.2555.300.00
MiniMax-M225.2022.0028.300.00
gemma-3-4b-it25.2010.1051.300.00
DeepSeek-V4-Flash24.705.3738.300.00
DeepSeek-V4-Flash-073124.3023.4025.200.00
Qwen3-Max23.4017.0029.400.00
Qwen3-VL-30B-A3B-Instruct23.403.3038.500.00
MiniMax-M2.522.7019.9027.300.00
mistral-7b22.301.3853.300.00
devstral-small22.102.9351.500.00
gemini-3.1-pro21.8017.6026.800.00
Mistral-Small-3.2-24B-Instruct-250621.703.3051.900.00
GPT-oss-120b21.207.6946.400.00
Mistral-Nemo-Instruct-240721.003.3669.000.00
gemma-3-27b-it20.704.2831.000.00
gemma-4-26B-A4B-it20.705.2434.000.00
Kimi-K2.520.101.1578.000.00
GLM-519.604.1547.400.00
qwen-3-235b19.102.4134.600.00
Qwen3-Max-Thinking18.9012.9024.800.00
GLM-4.6V18.701.3738.400.00
qwen-3-14b18.002.8154.700.00
claude-fable-517.9011.3023.600.00
DeepSeek-V4-Pro17.601.5241.000.00
qwen-2.5-72b17.302.3232.000.00
deepseek-ai/DeepSeek-R1-0528-Turbo17.309.4929.600.00
claude-4-sonnet17.2010.0023.700.00
DeepSeek-R1-052817.209.9430.800.00
claude-sonnet-4.616.907.2321.300.00
DeepSeek-V316.801.7644.800.00
claude-3-7-sonnet-latest16.708.6919.800.00
llama-3.1-8b16.708.7128.200.00
llama-3.1-405b16.603.0124.000.00
Kimi-K2.5-Turbo16.302.2464.700.00
deepseek-ai/DeepSeek-R1-Distill-Llama-70B16.1011.8020.400.00
Kimi-K2-Instruct-090516.102.7462.900.00
Olmo-3.1-32B-Instruct16.001.4762.900.00
MiniMax-M2.714.807.7533.400.00
Kimi-K2-Thinking14.202.0762.000.00
Kimi-K2.614.002.1934.700.00
llama-3.3-70b13.203.6827.600.00
Llama-3.3-70B-Instruct-Turbo13.201.2024.900.00
Qwen3-VL-235B-A22B-Instruct13.101.0540.200.00
gemma-4-31B-it12.801.0651.200.00
llama-3.2-11b12.601.1856.700.00
llama-3.2-1b12.401.2050.100.00
microsoft/WizardLM-2-8x22B12.405.8418.900.00
deepseek-v3.212.101.0232.300.00
llama-3.2-3b11.601.1340.200.00
MiniMax-M310.500.8761.100.00
DeepSeek-V3.18.802.2817.600.00
Qwen3-235B-A22B-Thinking-25077.871.3828.300.00
deepseek-ai/DeepSeek-OCR7.154.2910.000.00
PaddlePaddle/PaddleOCR-VL-0.9B5.904.267.530.00
allenai/olmOCR-2-7B-10255.703.088.320.00
Qwen3.8-Max4.722.666.780.00
Llama-Guard-4-12B2.760.824.250.00
Seed-2.0-pro2.311.543.060.00
Qwen3.7-Max2.292.042.630.00
Seed-2.0-code0.640.530.810.00

Frequently Asked Questions

Which deepinfra model is fastest?
Based on recent tests, qwen-3.5-35b-a3b shows the highest average throughput among tracked deepinfra models.
How many recent measurements feed this dashboard?
This provider summary aggregates 4284 individual prompts measured across 2726 monitoring runs over the past month.