llm-benchmarks
githubdrose.io

together Provider Benchmarks


Measured over the last 30 days

Models Tracked
20
Avg Tokens / Second
53.87
Avg Time to First Token (ms)
0.00
Last Updated
Aug 9, 2026

All together models

20 tracked · fastest first
together site
ModelAvg Toks/SecMinMaxAvg TTF (ms)
LFM2.5-8B-A1B109.0017.40141.000.00
qwen-2-1.5b-instruct102.003.13147.000.00
Kimi-K2.7-Code77.902.27199.000.00
qwen-2.5-7b70.9040.2092.900.00
GLM-5.270.002.84193.000.00
Qwen2.5-7B-Instruct-Turbo68.1048.3085.000.00
qwen-3.5-9b63.808.10109.000.00
Kimi-K2.661.902.93138.000.00
nemotron-3-ultra-550b-a55b56.307.96118.000.00
GPT-oss-120b54.003.70135.000.00
gemma-4-31B-it53.702.34104.000.00
GPT-oss-20b50.904.01167.000.00
DeepSeek-V4-Flash-073145.701.66109.000.00
llama-3.3-70b44.703.9697.300.00
Kimi-K338.402.7872.700.00
gemma-3n-e4b-it31.9013.0055.000.00
MiniMax-M325.902.8768.700.00
DeepSeek-V4-Pro24.705.0880.000.00
Inkling24.4010.1047.500.00
Llama-Guard-4-12B3.140.734.040.00

Frequently Asked Questions

Which together model is fastest?
Based on recent tests, LFM2.5-8B-A1B shows the highest average throughput among tracked together models.
How many recent measurements feed this dashboard?
This provider summary aggregates 2530 individual prompts measured across 2467 monitoring runs over the past month.