llm-benchmarks
githubdrose.io

together Provider Benchmarks


Measured over the last 30 days

Models Tracked
13
Avg Tokens / Second
19.00
Avg Time to First Token (ms)
5947.69
Last Updated
Sep 23, 2026

All together models

13 live · one row per model · fastest first · 2 folded behind their current lane
together site
ModelAvg Toks/SecMinMaxAvg TTF (ms)
GLM 5.256.9049.5064.30740.00
Llama 3.3 70B Instruct34.409.6855.901310.00
gpt-oss-120b33.7020.1047.101590.00
MiniMax M318.003.6937.604200.00
Gemma 4 31B16.801.5441.606390.00
Qwen3.8 2.4T A95B15.6015.6015.603840.00
Inkling Small15.3015.3015.303860.00
DeepSeek V4 Flash 073114.108.9919.104140.00
Muse Glimmer 30B13.0010.3015.004430.00
Inkling11.4010.2012.505370.00
gpt-oss-20b8.848.848.846850.00
Qwen3.5-9B4.844.844.8412800.00
DeepSeek V4 Pro 08134.151.388.5721800.00

Frequently Asked Questions

Which together model is fastest?
Based on recent tests, GLM 5.2 shows the highest average throughput among tracked together models.
How many recent measurements feed this dashboard?
This provider summary aggregates 2334 individual prompts measured across 2064 monitoring runs over the past month.