llm-benchmarks
githubdrose.io

coreweave Provider Benchmarks


Measured over the last 30 days

Models Tracked
32
Avg Tokens / Second
19.19
Avg Time to First Token (ms)
8827.61
Last Updated
Sep 1, 2026

All coreweave models

32 tracked · fastest first
ModelAvg Toks/SecMinMaxAvg TTF (ms)
Llama 3.1 8B Instruct57.1044.6078.00700.00
Granite 4.1 8B55.5033.1074.50500.00
Granite 4.1 8B50.6019.00110.00870.00
Llama 3.3 70B Instruct50.4044.8055.60560.00
Qwen3 30B A3B Instruct 250749.6035.9064.40480.00
Llama 3.1 70B Instruct47.0038.7051.20580.00
DeepSeek V3.138.4031.9045.30630.00
Kimi K2.7 Code35.8019.1047.301820.00
MiniMax M328.309.8561.102140.00
DeepSeek V4 Flash 042325.0019.6037.60920.00
Gemma 4 31B24.909.5232.10660.00
Gemma 4 31B24.6013.9030.20620.00
Nemotron 3.5 Lightning22.4016.4026.702720.00
MiniMax M318.406.9324.603920.00
Kimi K2.616.106.7727.304970.00
gpt-oss-120b13.604.2022.902630.00
gpt-oss-120b13.504.7726.003170.00
Granite 4.2 8B13.4013.4013.403760.00
gpt-oss-20b12.401.8934.805440.00
GLM 5.211.805.8019.106390.00
gpt-oss-20b11.106.3813.805830.00
Qwen3.5-35B-A3B11.109.8312.105580.00
gpt-oss-120b9.834.7216.705300.00
Auto Router (Beta)9.199.199.196660.00
DeepSeek V4 Flash 07314.741.837.5319300.00
GLM 5.24.534.534.5310700.00
Qwen3.6 35B A3B4.364.094.8014200.00
Nemotron 3.5 Lightning4.004.004.0015600.00
GLM 5.23.083.083.081690.00
DeepSeek V4 Pro 08132.602.602.6017300.00
Kimi K2.61.711.711.7135400.00
DeepSeek V4 Flash 07311.571.571.5742500.00

Frequently Asked Questions

Which coreweave model is fastest?
Based on recent tests, Llama 3.1 8B Instruct shows the highest average throughput among tracked coreweave models.
How many recent measurements feed this dashboard?
This provider summary aggregates 3553 individual prompts measured across 3095 monitoring runs over the past month.