Provider Overview (by Benchmark Speed)
Sorted by best measured throughput. Click any provider for full model listings and detailed benchmark history.
| Provider | Models Benchmarked | Best Tok/s | Median Tok/s | Best Model | Strength |
|---|---|---|---|---|---|
| cerebras | 2 | 129 | 105 | Gemma 4 31B | General purpose |
| bedrock | 32 | 115 | 54 | Nova Micro 1.0 | Enterprise compliance |
| openai | 30 | 111 | 27 | o3 Mini | Frontier quality, ecosystem |
| groq | 6 | 93 | 65 | Llama 3.3 70B Instruct | Speed, low latency |
| mara | 3 | 86 | 72 | gpt-oss-120b | General purpose |
| friendli | 5 | 83 | 40 | Gemma 4 31B | General purpose |
| modelrun | 3 | 78 | 24 | Gemma 4 31B | General purpose |
| parasail | 21 | 76 | 24 | Qwen3 Next 80B A3B Instruct | General purpose |
| coreweave | 15 | 75 | 21 | Llama 3.1 8B Instruct | General purpose |
| 3 | 75 | 2 | Gemini 2.5 Flash Lite | General purpose | |
| sambanovaturbo | 1 | 73 | 73 | Llama 3.3 70B Instruct | General purpose |
| baseten | 6 | 68 | 32 | Kimi K2.7 Code | General purpose |
| mistral | 2 | 65 | 49 | GLM 5.2 | European, GDPR |
| nebius | 5 | 54 | 14 | gpt-oss-120b | General purpose |
| crusoe | 6 | 53 | 29 | Llama 3.3 70B Instruct | General purpose |
Provider Profiles
OpenAI
The market leader for frontier models. GPT-5, GPT-5 Nano, o3, and o1 series lead reasoning, coding, and general capability benchmarks. Best ecosystem support and widest third-party integrations.
Best for: Frontier model quality, reasoning tasks, coding assistants, and applications where API reliability and ecosystem maturity matter most.
Watch out for: Direct OpenAI API throughput is lower than inference-optimized providers. Premium pricing on frontier models.
Anthropic
Claude 3 and Claude 4 models excel at instruction-following, long-context analysis, and safe output. Strong performer for document processing, multi-turn dialogue, and tasks requiring precise adherence to complex instructions.
Best for: Long-context analysis, instruction-following, regulated industries where output safety matters, and applications needing reliable structured output.
Watch out for: Smaller model catalog than OpenAI. API throughput is moderate.
Groq
Purpose-built inference hardware (LPUs) delivers category-leading throughput on open-weight models. Llama 3.3 70B and Qwen 3-32B regularly top speed benchmarks at 150+ tok/s with near-zero TTFT.
Best for: Real-time applications, voice interfaces, gaming, and any workload where sub-second streaming response start matters more than frontier model quality.
Watch out for: Limited to open-weight models. No GPT-4 or Claude. Rate limits can be tight on free tier.
AWS Bedrock
Aggregates models from Anthropic, Meta, Mistral, Amazon Nova, and others under one AWS-native API. Nova Micro (~118 tok/s) is the fastest Bedrock model. Strong compliance story: SOC2, HIPAA, FedRAMP.
Best for: Enterprise teams already on AWS. VPC-native deployments. HIPAA/FedRAMP compliance requirements. Data residency control.
Watch out for: API overhead adds latency vs direct provider calls. More complex IAM setup.
DeepInfra
One of the fastest open-weight inference providers, particularly on smaller models. Qwen 3.5-2B at ~203 tok/s and multiple 100+ tok/s options. OpenAI-compatible API.
Best for: High-volume, cost-sensitive workloads on open-weight models. Speed-critical applications where proprietary models are not required.
Watch out for: Less brand recognition than major providers. SLA / support coverage less mature than Groq or Fireworks.
Key Insights
OpenAI and Anthropic lead on model quality and capability breadth; inference-optimized providers (Groq, DeepInfra, Fireworks) lead on raw speed.
AWS Bedrock and Azure OpenAI suit enterprise teams with existing cloud agreements, compliance requirements, or VPC-native data residency needs.
Google Vertex AI integrates natively with Gemini models and GCP infrastructure — the strongest choice if you're already on Google Cloud.
Together AI and Fireworks offer the widest selection of open-weight models (Llama, Mistral, Qwen, DeepSeek) at competitive prices.
No single provider wins across all dimensions: speed, quality, price, compliance, and model variety all trade off differently.