The Same Request Has 8 Price Points
Take one production request — 4,000 input tokens, 1,000 output tokens — and price it across the current catalog (rates verified 2026-08-19; DeepSeek off-peak is 50% of peak):
| Model | Tier | Rate ($/1M in, out) | Cost / request | vs Sonnet 5 |
|---|---|---|---|---|
| GPT-5.6 Terra | Balanced | $2 / $12 | $0.0200 | +11% |
| Claude Sonnet 5 | Balanced | $2 / $10 | $0.0180 | baseline |
| Grok 4.5 | Balanced | $2 / $6 | $0.0140 | −22% |
| Gemini 3.7 Flash | Balanced | $1.50 / $6 | $0.0120 | −33% |
| DeepSeek V4 Pro | Balanced | $1.32 / $3.96 | $0.0092 | −49% |
| Gemini 3.5 Flash-Lite | Budget | $0.30 / $2.50 | $0.0037 | −79% |
| DeepSeek V4 Flash | Budget | $0.44 / $1.32 | $0.0031 | −83% |
| GPT-5.6 Luna | Budget | $0.20 / $1.20 | $0.0020 | −89% |
Note what the table hides: output price drives the spread. GPT-5.6 Terra charges more for the same request than Claude Sonnet 5 purely because its output rate is $12 vs $10 per 1M. And GPT-5.6 Luna costs 1/9th of Claude Sonnet 5 on the identical workload. Routing is the cheapest optimization you haven't done — it needs no prompt changes and no new code paths.
Which Tasks Can Safely Downgrade?
The decision is task-based, not model-based. Use this matrix as your starting point — it matches today's tier labels (flagship / balanced / budget) to common workloads:
- Complex multi-step reasoning
- Large code generation & refactors
- Nuanced long-form writing
- Legal / financial / medical analysis
- Most production chatbots
- RAG & document analysis
- Mid-length code with tests
- Extraction with judgment calls
- FAQ answers & simple Q&A
- Classification & tagging
- Short summaries & rewriting
- Structured extraction from clean input
- The router itself
The budget column is bigger than most teams assume. In typical SaaS traffic, 40–60% of requests are lookups, classifications, extractions, and summaries — tasks where a budget model's output is indistinguishable from a flagship's. The reason the money is left on the table is almost always inertia: one model wired into every endpoint.
Six Tasks, Routed: The Before/After Table
Here is the matrix applied to six common SaaS tasks. Baseline is Claude Sonnet 5 everywhere; costs are per request at the stated token mix (rates verified 2026-08-19; DeepSeek off-peak is 50% of peak):
| Task | Token mix (in/out) | Sonnet 5 cost | Routed to | Routed cost | Saved |
|---|---|---|---|---|---|
| FAQ bot | 4K / 1K | $0.0180 | GPT-5.6 Luna | $0.0020 | −89% |
| Email triage | 4K / 1K | $0.0180 | Gemini 3.5 Flash-Lite | $0.0037 | −79% |
| Structured extraction | 4K / 1K | $0.0180 | DeepSeek V4 Flash (off-peak) | $0.0015 | −91% |
| Summarization | 6K / 1K | $0.0220 | GPT-5.6 Luna | $0.0024 | −89% |
| Code assistant | 2K / 3K | $0.0340 | Gemini 3.7 Flash | $0.0210 | −38% |
| RAG chatbot | 6K / 800 | $0.0200 | DeepSeek V4 Pro | $0.0111 | −45% |
The pattern holds everywhere: lookup and extraction tasks save 79–91%; interactive and reasoning-adjacent tasks save 38–45% — still worth doing, but they stay on models with stronger reasoning. The routing tier for code and chat is "balanced," not "budget."
At 1M requests a month with a 70/20/10 mix of these tasks, the six rows alone move the bill from ~$18,000 to ~$5,600 — the same worked example below, but built bottom-up from tasks instead of top-down from a ratio.
The Routing Framework: Rules, Classifier, Fallback
Production routing has three layers, in increasing order of sophistication. Start at layer 1 — it captures most of the savings:
Layer 1: Rule-based routing
Route by task endpoint, input length, intent keywords, and prior failures. A 30-line function beats a flagship model on 60% of traffic from day one.
Layer 2: A cheap model as the router
Use GPT-5.6 Luna or Gemini 3.5 Flash-Lite to classify each request into a tier. The router costs fractions of a cent and replaces brittle keyword lists.
Layer 3: Confidence fallback (route up, not down)
Start cheap; if the budget model signals low confidence, errors, or a schema violation, escalate to the balanced or flagship tier. This is what protects quality while routing.
The one rule that makes routing safe: every escalation is logged. If 15% of your budget-tier traffic escalates, your routing rule is wrong for 15% of traffic — tune the rules weekly until escalations stay under 3%.
Worked Example: 1M Requests, 4K In / 1K Out
A typical SaaS with 1M requests per month currently runs everything on Claude Sonnet 5 at $18,000/month. Here's the routed version:
Push the mix to 80% budget and the saving lands near 71% ($5,200/month); swap Luna for DeepSeek V4 Flash off-peak as the budget tier and it reaches ~73%. The exact ceiling depends on your task mix — but for most products, 50–80% of traffic can downgrade safely.
Routing and caching stack: caching already takes 90% off your repeated input on most providers, so the combination is ~90% off input and a lower rate on everything. Model your own mix with the savings calculator before touching code.
Off-Peak Routing: The Cheapest Trick Nobody Uses
DeepSeek prices V4 models at 50% off during off-peak hours — everything outside 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday. That is most of the day for most time zones, and the discount applies to input, output, and cache-hit rates. The saving needs no model change, no quality change, and no code beyond a time check:
It also stacks with the cache-hit rate ($0.014 → $0.007 off-peak), making off-peak DeepSeek the cheapest serious production rate on the market today. Batch APIs are the same idea from the other direction: OpenAI, Anthropic, and Google batch endpoints discount ~50% (Google's Gemini 3.7 Flash batch was listed at 75% off on OpenRouter in August 2026) for work that can wait up to 24 hours. Nightly pipelines, evals, and backfills are the obvious candidates.
Practical setup: a cron that flips your extraction and summarization queues to DeepSeek V4 Flash off-peak after 22:00 UTC, and a batch endpoint for anything with a >1-hour SLA. Both changes are invisible to users.
When NOT to Route
Routing is a framework, not a mandate. Keep traffic on the higher tier when any of these hold:
- Quality-critical user-facing writing. Brand voice, marketing copy, and client deliverables — the nuance floor is real and visible.
- Latency-critical interactive features. Budget routes and free tiers can queue at peak hours; keep a paid balanced model for the request path.
- Compliance or contractual constraints. Data residency, model audit, and vendor-approval requirements override price.
- Tiny volume (<10K requests/month). The savings (~$150/month at best) don't justify router complexity and a second integration.
- Long-context tasks on surcharged models. GPT-5.6 bills 2× past 272K tokens; routing a 500K-token job to a cheaper base rate can backfire. Price the actual context first.
- Brand-new features. Let a week of logged escalations prove the tier before you route it. Re-routing after launch is embarrassing; logging after launch is impossible.
What You Actually Give Up by Downgrading
Be honest about the tradeoffs so routing doesn't backfire:
- Nuance ceiling: On ambiguous or creative work, budget models flatten tone, miss subtext, and hedge more. Keep flagship routing for those tasks.
- Reasoning depth: Multi-step deduction degrades on budget tiers. Long agent loops or math-heavy analysis should stay balanced or above.
- Context limits: Budget tiers often have smaller context windows or long-context surcharges (e.g., GPT-5.6 bills 2× above 272K tokens). Long documents change the calculus.
- Latency under load: Budget tiers are fast, but free/cheap routes (OpenRouter free models, free tiers) can queue at peak hours — keep a paid fallback for user-facing latency.
The framework handles all four: tradeoffs are acceptable for the tasks you route, and the fallback catches the rest. If a task type escalates more than 3% of the time, it belongs on a higher tier — the logging tells you, not the marketing page.
The Routing Decision Tree
When a new request type arrives, run it through five questions in order — the first "yes" wins:
1. Is it user-facing writing?
Brand copy, emails to customers, client deliverables → flagship. The nuance floor is visible and embarrassing.
2. Does it need multi-step reasoning or large refactors?
Deep deduction, architecture decisions, 60+ line code generation → balanced tier at minimum.
3. Is it lookup, classification, extraction, or a short summary?
→ budget tier. These four task types are the 79–91% savings rows from the table above.
4. Is the context huge or compliance-bound?
Past ~272K tokens on surcharged models, or subject to data-residency rules → price the actual context before choosing; the cheapest base rate can lose.
5. Unsure?
Default to balanced, log for a week, and let the escalation rate decide. The logging is the tiebreaker — not the marketing page.
The 3-2-1 Rollout Plan
Routing ships like a feature, not a flag-flip. The 3-2-1 plan keeps quality risk near zero:
Week 1: route 3 low-risk tasks
FAQ answering, classification, structured extraction — the rows with 79–91% savings. Wire logging on every routing decision from day one.
Week 2: add 2 more tasks
Summarization and email triage join the budget tier. Check the escalation rate daily; anything above 3% goes back up a tier immediately.
Every week after: 1 review
Rebalance the mix, move offenders up, move proven tasks down. Re-run the pricing model monthly — model prices still drop ~40% a year, so last quarter's budget tier may now be mid-tier.
The plan front-loads logging because the log is the asset that makes every later decision safe — escalation rates, task mix, and price changes all get decided from the same sheet.
Getting Started in a Weekend
Export a month of usage and group requests by task type — you need cost per task, not cost per endpoint.
Pick one high-volume, low-complexity task (FAQ, classification, extraction) and point it at a budget model.
Run side-by-side for a week on a sample; log escalations and manual review flags.
Roll out the router for that task, then repeat for the next task type. Rinse quarterly — model prices keep falling.
Frequently Asked Questions
What is the cheapest AI model for simple tasks?
At verified 2026-08-19 rates, GPT-5.6 Luna ($0.20/$1.20 per 1M in/out tokens), Gemini 3.5 Flash-Lite ($0.30/$2.50), and DeepSeek V4 Flash off-peak ($0.22/$0.66) are the strongest budget options. For FAQ lookup, classification, extraction, and short summaries, they deliver 80-90% of a flagship's output at 1/25th the price.
How much can model routing save?
In our worked example, moving 70% of a 1M-request/month workload off Claude Sonnet 5 onto GPT-5.6 Luna and Gemini 3.7 Flash cut the bill from $18,000 to $5,600 per month — a 69% saving. Most teams can route 50-80% of their traffic without measurable quality loss.
Which tasks should stay on flagship models?
Complex multi-step reasoning, large code generation and refactoring, nuanced long-form writing, and high-stakes analysis (legal, financial, medical) still justify flagship pricing. Everything else — routing, classification, extraction, summaries, FAQ answers — has a cheaper tier that is good enough.
Does routing hurt output quality?
Task-appropriate routing does not. The failure pattern is routing to the cheapest model for a task it cannot do (e.g., a 60-line refactor on a budget tier). The fix is a confidence fallback: start cheap, and escalate to the flagship when the cheap model signals low confidence or fails.
How do I implement model routing in a weekend?
Start with rules: route by input length, intent keywords, and task type; log every routing decision and its outcome for a week. Then add a cheap classifier to do the routing itself, and wire a failure fallback that escalates to a better model. The logging layer is the part most teams skip — without it you cannot measure quality loss.
Does off-peak pricing really save 50%?
Yes, for DeepSeek V4: everything outside 01:00-04:00 and 06:00-10:00 UTC Monday to Friday is billed at half the peak rate, on input, output, and cache-hit tokens. Move batch and nightly work into those hours and the discount applies automatically. Provider batch APIs are a similar ~50% (sometimes more) for jobs that can wait.