The Families, Side by Side
Representative models per family, verified 2026-08-28. Per-request cost = 4K in / 1K out, uncached:
| Family | Model | Rate ($/1M in, out) | Int / Cod / Agt | Per request | Monthly @1M |
|---|---|---|---|---|---|
| OpenAI — frontier | GPT-5.6 Sol | $5 / $30 | 60.9 / 77.4 / 57.8 | $0.0500 | $50,000 |
| OpenAI — balanced | GPT-5.6 Terra | $2 / $12 | 56.6 / 76.7 / 50.2 | $0.0200 | $20,000 |
| Anthropic — frontier | Claude Opus 5 | $5 / $25 | 63.1 / 78.0 / 59.2 | $0.0450 | $45,000 |
| Anthropic — frontier | Claude Fable 5 | $10 / $50 | 62.1 / 76.5 / 56.6 | $0.0900 | $90,000 |
| Anthropic — balanced | Claude Sonnet 5 | $2 / $10 | 55.3 / 71.5 / 49.7 | $0.0180 | $18,000 |
| Google — frontier | Gemini 3.1 Pro | $2 / $12 | 47.7 / 68.8 / 23.0 | $0.0200 | $20,000 |
| Google — balanced | Gemini 3.7 Flash | $1.50 / $6 | 56.0 / 76.1 / 45.1 | $0.0120 | $12,000 |
The balanced tier is where the fight lives: Flash is 33-40% cheaper per request than Sonnet or Terra while scoring 76.1 coding — effectively tied with Terra's 76.7. The frontier tier is closer than expected: Opus 5 beats Sol on both price and every composite axis. And Google's 3.1 Pro is the outlier — frontier-priced at balanced-quality on agentic.
Caching and Batch: The Second Price War
All three families discount cached input ~90% and batch ~50% — but the mechanics differ:
OpenAI
90% cached input (Sol $5→$0.50), 50% batch, cache_control markers, 2× long-context surcharge past 272K tokens on Sol.
Anthropic
90% cached input (Sonnet $2→$0.20), 50% batch, cache breakpoints, 5-minute cache TTL — the friendliest for chat history prefixes.
Up to 90% context caching, automatic on prefix match, batch discounts up to 75% on Flash during promos, 2M-token context — the cheapest effective input at scale.
With 90% cached input, the per-request costs compress: Terra $0.0135, Sonnet $0.0115, Flash $0.0071. Google's automatic caching plus the cheapest output rate makes it the volume winner; Anthropic's breakpoints make it the chat-history winner.
Per-Workload Verdicts
| Workload | Winner | Runner-up | Why |
|---|---|---|---|
| Production chatbot | Gemini 3.7 Flash | Claude Sonnet 5 | Cheapest output + automatic caching; Sonnet for tone-sensitive copy |
| Coding assistant | GPT-5.6 Terra | Gemini 3.7 Flash | 76.7 coding with Sol-class reasoning; Flash within a point at 60% of the price |
| Long documents / agents | Claude Opus 5 | GPT-5.6 Sol | Best agentic (59.2) + instruction following for multi-step work |
| High-volume extraction | Gemini 3.7 Flash | — | At $6 output, the same pipeline costs 40-50% less than the other two |
| Hard reasoning | Claude Opus 5 | GPT-5.6 Sol | 63.1 vs 60.9 intelligence, and Opus is cheaper on output |
Every head-to-head above is interactive on the comparison arena with the per-use-case fit math. And the honest footnote: our five-model build study shows DeepSeek V4 Pro and GLM 5.3 delivering 90% of this quality at 40-60% of the price — the flagship families are the premium option, not the default.
Frequently Asked Questions
Which flagship family is cheapest in 2026?
At the balanced tier with a 4K-in/1K-out request: Gemini 3.7 Flash costs $0.012 (cheapest), then Claude Sonnet 5 at $0.018, then GPT-5.6 Terra at $0.020 — Terra's $12 output rate makes it the priciest despite matching input rates. At the frontier tier, Claude Opus 5 ($0.045) undercuts GPT-5.6 Sol ($0.05) on the same request.
Which family has the best benchmark scores?
By third-party composites (verified 2026-08-28): Anthropic leads frontier coding (Opus 5 at 78.0, Fable 5 at 76.5) and intelligence (Opus 5 at 63.1); Google's Gemini 3.7 Flash (76.1 coding) and OpenAI's GPT-5.6 Sol (77.4 coding, 60.9 intelligence) are within a few points. The families are closer than their marketing.
How do caching discounts compare?
All three offer ~90% cached-input discounts: Sol $5→$0.50, Sonnet 5 $2→$0.20, Gemini 3.1 Pro $2→$0.20. Google's caching is automatic on prefix match; OpenAI and Anthropic need explicit cache markers.
Which family should my team standardize on?
For most production workloads, Google wins on price (Flash output $6 vs Sonnet $10 vs Terra $12), Anthropic wins on writing and instruction-following, and OpenAI wins on ecosystem breadth. Our fit verdicts below give the per-use-case answer.
Are there cheaper alternatives to all three?
Yes — DeepSeek V4 Pro and GLM 5.3 deliver 90%+ of the balanced-tier quality at 40-60% of the price, and budget tiers (Luna, V4 Flash) serve most volume traffic. The flagship families justify themselves on the hard 10% of requests.