llm-benchmarks
githubdrose.io

novita Provider Benchmarks


Measured over the last 30 days

Models Tracked
100
Avg Tokens / Second
17.70
Avg Time to First Token (ms)
10249.29
Last Updated
Sep 1, 2026

All novita models

100 tracked · fastest first
ModelAvg Toks/SecMinMaxAvg TTF (ms)
Llama 3.1 8B Instruct53.6041.5061.20830.00
Llama 3 8B Lunaris48.9044.3051.30640.00
Llama 4 Maverick48.8038.2058.90660.00
Nemotron 3 Nano 30B A3B47.1032.9066.201220.00
Ling-3.0-flash45.1041.0047.001280.00
Qwen3 Next 80B A3B Instruct44.8024.6055.601010.00
Qwen3 Coder 30B A3B Instruct43.6025.4063.801100.00
Qwen3 Coder 30B A3B Instruct42.0041.2042.801040.00
Qwen3 VL 30B A3B Instruct40.2032.3051.401160.00
Ling-2.6-flash40.1021.2057.50910.00
Ling-3.0-flash39.8035.3043.301460.00
Ling-2.6-flash39.0034.7042.20900.00
Ling-2.6-1T36.0034.1037.301090.00
Qwen3 Coder Next33.3026.7041.801620.00
Llama 4 Maverick32.4021.0038.801350.00
Kimi K2 071131.7031.7031.70720.00
Llama 3.1 Euryale 70B v2.231.0019.6038.20650.00
Qwen3 Coder 480B A35B30.3030.3030.301200.00
Llama 4 Scout29.4022.5043.10780.00
Ring-2.6-1T29.2029.2029.201490.00
Ling-2.6-1T27.905.9242.001580.00
Llama 3.1 8B Instruct27.8027.8027.802010.00
MiniMax M327.3022.0032.402120.00
Kimi K2 090527.1027.1027.10800.00
Gemma 3 27B25.0023.6026.50930.00
Kimi K2 090525.0022.5027.501090.00
Qwen3 235B A22B Instruct 250724.7020.6029.90990.00
Llama 3.3 70B Instruct23.6014.9032.50920.00
Kimi K2 071123.6022.5025.801220.00
Ring-2.6-1T23.3020.2026.002250.00
DeepSeek V321.5020.8022.501320.00
Gemma 3 27B21.4013.4028.901130.00
Qwen3 VL 235B A22B Instruct21.3012.9035.00780.00
DeepSeek V3.221.2012.5026.101560.00
Qwen3 VL 235B A22B Instruct21.008.7333.301010.00
DeepSeek V3 032420.9017.9022.701560.00
gpt-oss-120b19.6014.3027.902740.00
DeepSeek V3 032419.4016.9023.001620.00
ERNIE 4.5 VL 424B A47B18.8015.0021.801730.00
DeepSeek V318.4015.0022.801810.00
Gemma 4 26B A4B18.104.3027.701160.00
DeepSeek V3.2 Exp17.4016.2018.501510.00
DeepSeek V3.1 Terminus17.1015.2018.601770.00
DeepSeek V3.2 Exp17.1014.0021.801700.00
Qwen3 235B A22B Instruct 250717.008.9225.102800.00
DeepSeek V3.116.2015.0018.301700.00
DeepSeek V3.1 Terminus16.2015.0017.801960.00
ERNIE 4.5 VL 424B A47B15.606.2821.101930.00
DeepSeek V3.115.5012.8017.002010.00
GLM 5.214.909.3120.803380.00
GLM 5.214.4011.7016.803300.00
Auto Router (Beta)14.3014.3014.303990.00
gpt-oss-20b14.0012.1015.903940.00
Llama 3.3 70B Instruct13.205.9421.303210.00
DeepSeek V4 Flash 073113.108.0717.104610.00
MiMo-V2.5-Pro12.607.8515.504350.00
MiniMax M112.5010.3014.406440.00
GLM 4.5 Air11.609.6913.504910.00
Gemma 4 31B11.1011.1011.101400.00
WizardLM-2 8x22B10.7010.4011.001180.00
WizardLM-2 8x22B10.405.5719.302810.00
DeepSeek V4 Flash 04239.702.5914.609630.00
Qwen3.8 2.4T A95B9.169.169.165520.00
GLM 5.18.561.4221.9024400.00
GLM 4.5 Air8.527.908.946500.00
MiMo-V2.58.243.6013.309410.00
MiniMax M2.18.036.659.688690.00
GLM 57.971.9724.6020300.00
Mistral Nemo7.867.867.865120.00
MiniMax M27.155.097.959240.00
MiniMax M2.16.023.4911.7014600.00
Kimi K2.7 Code5.841.729.9815800.00
Kimi K2.55.765.765.769610.00
GLM 4.5V5.494.036.4411500.00
MiniMax M2.75.402.2122.5014600.00
MiniMax M2.55.374.016.8113600.00
Kimi K2 Thinking5.315.315.3110500.00
Qwen3 VL 235B A22B Thinking5.294.107.0111400.00
Kimi K2.55.224.755.9310700.00
Kimi K2 Thinking4.734.005.2912200.00
Qwen3 235B A22B Thinking 25074.662.728.3516200.00
GLM 4.5V4.653.475.2513600.00
Qwen3 VL 235B A22B Thinking4.412.037.2414700.00
Qwen3.5-122B-A10B3.902.145.3518500.00
Qwen3.5-27B3.273.053.5318900.00
Qwen3.5 397B A17B3.222.503.9219700.00
R1 Distill Llama 70B3.111.514.8022000.00
R1 Distill Llama 70B3.072.893.2518000.00
GLM 4.63.052.363.7620700.00
Kimi K2.62.901.824.5224100.00
R12.841.244.6425100.00
GLM 4.7 Flash2.732.732.7322400.00
R12.702.143.2621300.00
Hy32.662.283.3124200.00
GLM 4.72.551.643.3026500.00
MiniMax M2.72.321.314.1037500.00
GLM 4.6V1.490.852.6646100.00
Qwen3.8 27B1.150.651.7565900.00
GLM 5.31.140.102.1823800.00
GLM 4.6V0.950.950.9565400.00

Frequently Asked Questions

Which novita model is fastest?
Based on recent tests, Llama 3.1 8B Instruct shows the highest average throughput among tracked novita models.
How many recent measurements feed this dashboard?
This provider summary aggregates 7980 individual prompts measured across 6262 monitoring runs over the past month.