The two cheapest usable models in the catalog, head to head.
gemini-2.5-flash-lite · google
ministral-3-latest-8b · mistral
Benchmark
| Workload Scenario | Gemini 2.5 Flash-Lite | Ministral 3 (8B) | Price Delta |
|---|---|---|---|
| 1M input tokens (raw text) | $0.10 | $0.15 | −$0.05 |
| 1M output tokens (generation) | $0.40 | $0.15 | +$0.25 |
| 1M tokens · 70% input / 30% output mix | $0.19 | $0.15 | +$0.04 |
| Standard chat turn (4K in / 800 out, 50% cached) | $0.00054 | $0.00045 | +$0.000090 |
| Monthly scale (10K requests / day) | $162.00 | $135.00 | +$27.00 |
Negative difference = Gemini 2.5 Flash-Lite is cheaper. Positive = Ministral 3 (8B) is cheaper.
Scaling Curve
Total cost of a token volume at 70/30 input/output split (uncached). Log-log scale.
Verdict
Gemini 2.5 Flash-Lite has the cheaper input rate, while Ministral 3 (8B) has the cheaper output rate. The crossover happens when output makes up about 17% of your total tokens.
Below that share (retrieval, summarization, extraction — lots of context in, little text out) Gemini 2.5 Flash-Lite is cheaper. Above it (generation, translation, coding — long completions) Ministral 3 (8B) wins.
FAQ
On input, Gemini 2.5 Flash-Lite is cheaper ($0.10/M vs $0.15/M). On output, Ministral 3 (8B) is cheaper ($0.15/M vs $0.40/M). For workloads where more than 17% of tokens are output, the output-cheaper model wins overall.
A chat-style request (4,000 input + 800 output tokens, 50% cached) costs $0.00054 on Gemini 2.5 Flash-Lite and $0.00045 on Ministral 3 (8B) — Ministral 3 (8B) is 1.2× more expensive for that workload.
Gemini 2.5 Flash-Lite supports 1,048,576 tokens (65.5K max output); Ministral 3 (8B) supports 131,072 (16.4K max output). Gemini 2.5 Flash-Lite fits 8.0× more context, which matters for long documents and agents.