Speed-optimized FireOptimizer serverless inference platform serving Llama 4 Maverick/Scout, DeepSeek V4/R1, and custom fine-tuned LoRA models.
2 models tracked · cheapest input: Fireworks AI — Llama 4 Scout at $0.12/M · official pricing page
accounts/fireworks/models/llama-4-scout
Input /M
$0.12
Output /M
$0.36
Context
10M
FireOptimizer-accelerated Llama 4 Scout endpoint delivering sub-200ms latency on 10M context requests at $0.12/$0.36.
accounts/fireworks/models/deepseek-r1
Input /M
$0.55
Output /M
$2.19
Context
128K
Optimized serverless DeepSeek R1 reasoning deployment with structured function calling support at $0.55/$2.19.
Other providers
All Fireworks AI prices verified Aug 19, 2026. Prices change frequently — each model page links to the authoritative source.