llm-benchmarks
githubdrose.io

nebius Provider Benchmarks


Measured over the last 30 days

Models Tracked
21
Avg Tokens / Second
21.98
Avg Time to First Token (ms)
9078.65
Last Updated
Sep 1, 2026

All nebius models

21 tracked · fastest first
ModelAvg Toks/SecMinMaxAvg TTF (ms)
Nemotron 3 Nano 30B A3B65.9065.9065.90780.00
Hermes 4 70B55.0050.9057.80570.00
Qwen3 30B A3B Instruct 250741.3037.3046.20510.00
Hermes 4 70B36.908.4461.001240.00
gpt-oss-120b35.8025.1042.101570.00
Qwen3 235B A22B Instruct 250730.405.7551.503680.00
Hermes 4 405B29.2024.1036.00530.00
Hermes 4 405B26.406.7233.40760.00
Qwen2.5 VL 72B Instruct25.006.4135.201020.00
Qwen2.5 VL 72B Instruct24.509.1632.10600.00
Nemotron 3 Super21.9020.0024.402800.00
Qwen3 30B A3B Instruct 250719.609.4729.002480.00
Gemma 3 27B18.607.6940.102560.00
Gemma 3 27B18.105.4748.701360.00
Nemotron 3 Super13.704.8922.407780.00
Qwen3 Next 80B A3B Thinking7.826.808.837660.00
Llama 3.3 70B Instruct6.596.596.595740.00
MiniMax M2.56.273.577.7510900.00
Qwen3 32B3.982.096.6418200.00
Llama 3.3 70B Instruct1.461.461.4611700.00
GLM 5.11.451.161.6643400.00

Frequently Asked Questions

Which nebius model is fastest?
Based on recent tests, Nemotron 3 Nano 30B A3B shows the highest average throughput among tracked nebius models.
How many recent measurements feed this dashboard?
This provider summary aggregates 2706 individual prompts measured across 2602 monitoring runs over the past month.