FireOptimizer vs custom LPU: Fireworks AI ($0.12/$0.36) vs Groq ($0.11/$0.34).
accounts/fireworks/models/llama-4-scout · fireworks
llama-4-scout-lpu · groq
Benchmark
| Workload Scenario | Fireworks AI — Llama 4 Scout | Groq LPU — Llama 4 Scout | Price Delta |
|---|---|---|---|
| 1M input tokens (raw text) | $0.12 | $0.11 | +$0.01 |
| 1M output tokens (generation) | $0.36 | $0.34 | +$0.02 |
| 1M tokens · 70% input / 30% output mix | $0.192 | $0.179 | +$0.013 |
| Standard chat turn (4K in / 800 out, 50% cached) | $0.000588 | $0.000602 | −$0.000014 |
| Monthly scale (10K requests / day) | $176.40 | $180.60 | −$4.20 |
Negative difference = Fireworks AI — Llama 4 Scout is cheaper. Positive = Groq LPU — Llama 4 Scout is cheaper.
Scaling Curve
Total cost of a token volume at 70/30 input/output split (uncached). Log-log scale.
Verdict
Groq LPU — Llama 4 Scout is cheaper on both input and output rates, so it costs less at every input/output mix. Price alone still isn't the whole decision: capability, latency and context limits (Fireworks AI — Llama 4 Scout: 10M, Groq LPU — Llama 4 Scout: 512K) may justify the premium for your task.
FAQ
On input, Groq LPU — Llama 4 Scout is cheaper ($0.11/M vs $0.12/M). On output, Groq LPU — Llama 4 Scout is cheaper ($0.34/M vs $0.36/M). The same model is cheaper on both sides, so it wins at every mix.
A chat-style request (4,000 input + 800 output tokens, 50% cached) costs $0.000588 on Fireworks AI — Llama 4 Scout and $0.000602 on Groq LPU — Llama 4 Scout — Fireworks AI — Llama 4 Scout is 1.0× cheaper for that workload.
Fireworks AI — Llama 4 Scout supports 10,000,000 tokens (16.4K max output); Groq LPU — Llama 4 Scout supports 512,000 (16.4K max output). Fireworks AI — Llama 4 Scout fits 19.5× more context, which matters for long documents and agents.