Flagship, balanced, and budget — the three-tier model landscape and how to pick the right one for each workload without over-spending or under-delivering.
| Tier | Example Models | Input $/1M | Output $/1M | Context | Best For |
|---|---|---|---|---|---|
| Flagship | GPT-5.6 Sol, Claude Opus 5, o3 | $5–$10 | $15–$40 | 200K–1M | Complex reasoning, nuanced writing, hard code gen |
| Balanced | GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.1 Pro | $1.50–$3 | $7–$15 | 200K–1M | Most production chatbots, RAG, document analysis |
| Budget | GPT-5.6 Luna, Gemini Flash, Claude Haiku | $0.06–$0.50 | $0.45–$2.50 | 200K–1M | Extraction, classification, high-volume simple tasks |
Use when:
Avoid for:
Use when:
Avoid for:
Use when:
Avoid for:
Does this task require multi-step reasoning or expert-level judgment?
Is this customer-facing with visible quality bar?
Is request volume > 500K/month?
Did quality evaluation pass at current tier?
Not for every task. Flagship models win on complex multi-step reasoning, subtle creative writing, and hard code generation. But for extraction, classification, FAQ answering, and simple summarization, a balanced or budget model gets 95%+ of the quality at 10–50% of the price. Most teams should be using a mix of tiers.
Start with the balanced tier and evaluate output quality. If quality is insufficient, upgrade to flagship. If quality exceeds your bar (i.e., you're over-paying), step down to budget. Use the same 10–20 representative test examples consistently across evaluations.
Most modern budget models (GPT-5.6 Luna, Gemini Flash Lite, Claude Haiku 4.5) support structured output including JSON mode and function/tool calling. Check our model catalog for feature flags — older budget models may lack reliable structured output support.
At current rates (August 2026), the spread is roughly 10–100× depending on provider. GPT-5.6 Luna ($0.20 input / $1.20 output per 1M tokens) vs o3 ($10 / $40) = 50× price difference on input, 33× on output. Even balanced vs flagship is a 5–10× difference.