Extreme wafer-scale engine inference: up to 3,000 tokens/second on GPT-OSS-120B, Gemma 4, Qwen 3, and GLM 4.7 with generous daily free developer tiers.
3 models tracked · cheapest input: Cerebras — GPT OSS 120B at $0.35/M · official pricing page
gpt-oss-120b
Input /M
$0.35
Output /M
$0.75
Context
128K
Wafer-Scale Engine inference running 120B parameters at world-record 3,000 tokens/second for $0.35/$0.75 per 1M tokens.
gemma-4-31b-cerebras
Input /M
$0.99
Output /M
$1.49
Context
128K
Google Gemma 4 running on CS-3 wafer systems at 1,850 tokens/second for real-time document analysis at $0.99/$1.49.
qwen-3-32b-cerebras
Input /M
$0.40
Output /M
$0.80
Context
128K
Qwen 3 running on Cerebras hardware delivering 2,000+ tokens/second at $0.40/$0.80 per 1M tokens.
Other providers
All Cerebras CS-3 prices verified Aug 19, 2026. Prices change frequently — each model page links to the authoritative source.