Head-to-head showdown: Fireworks AI — DeepSeek R1 ($0.55 in / $2.19 out per 1M) vs Devstral 2 (2512) (OpenRouter) ($0.44 in / $2.20 out per 1M). Devstral 2 (2512) (OpenRouter) is 1.0× cheaper across standard token mixes, with 128K vs 262.1K context windows.
accounts/fireworks/models/deepseek-r1 · fireworks
devstral-2512 · mistral
Benchmark
| Workload Scenario | Fireworks AI — DeepSeek R1 | Devstral 2 (2512) (OpenRouter) | Price Delta |
|---|---|---|---|
| 1M input tokens (raw text) | $0.55 | $0.44 | +$0.11 |
| 1M output tokens (generation) | $2.19 | $2.20 | −$0.01 |
| 1M tokens · 70% input / 30% output mix | $1.042 | $0.968 | +$0.074 |
| Standard chat turn (4K in / 800 out, 50% cached) | $0.003952 | $0.002728 | +$0.001224 |
| Monthly scale (10K requests / day) | $1,185.60 | $818.40 | +$367.20 |
Negative difference = Fireworks AI — DeepSeek R1 is cheaper. Positive = Devstral 2 (2512) (OpenRouter) is cheaper.
Scaling Curve
Total cost of a token volume at 70/30 input/output split (uncached). Log-log scale.
Verdict
Devstral 2 (2512) (OpenRouter) has the cheaper input rate, while Fireworks AI — DeepSeek R1 has the cheaper output rate. The crossover happens when output makes up about 92% of your total tokens.
Below that share (retrieval, summarization, extraction — lots of context in, little text out) Devstral 2 (2512) (OpenRouter) is cheaper. Above it (generation, translation, coding — long completions) Fireworks AI — DeepSeek R1 wins.
FAQ
On input, Devstral 2 (2512) (OpenRouter) is cheaper ($0.44/M vs $0.55/M). On output, Fireworks AI — DeepSeek R1 is cheaper ($2.19/M vs $2.20/M). For workloads where more than 92% of tokens are output, the output-cheaper model wins overall.
A chat-style request (4,000 input + 800 output tokens, 50% cached) costs $0.003952 on Fireworks AI — DeepSeek R1 and $0.002728 on Devstral 2 (2512) (OpenRouter) — Devstral 2 (2512) (OpenRouter) is 1.4× more expensive for that workload.
Fireworks AI — DeepSeek R1 supports 128,000 tokens (16.4K max output); Devstral 2 (2512) (OpenRouter) supports 262,144 (209.7K max output). Devstral 2 (2512) (OpenRouter) fits 2.0× more context, which matters for long documents and agents.