High-performance GPU cloud delivering serverless and dedicated endpoints for Llama 4, DeepSeek V4/R1, Qwen 2.5, and open-weight models with fine-tuning.
3 models tracked · cheapest input: Together AI — GPT-OSS-20B at $0.05/M · official pricing page
meta-llama/Llama-4-Maverick-Together
Input /M
$0.35
Output /M
$0.90
Context
1M
Serverless optimized Llama 4 Maverick endpoint with automated failover and OpenAI-compatible API at $0.35/$0.90.
deepseek-ai/DeepSeek-V4-Together
Input /M
$0.50
Output /M
$1.50
Context
1M
High-concurrency DeepSeek V4 deployment hosted in North American Tier-4 data centers at $0.50/$1.50 per 1M tokens.
together/gpt-oss-20b
Input /M
$0.05
Output /M
$0.20
Context
32.8K
High-volume open baseline at $0.05 input and $0.20 output per 1M tokens.
Other providers
All Together AI prices verified Aug 19, 2026. Prices change frequently — each model page links to the authoritative source.