Enable Prompt Caching
Prompt caching is the highest-ROI optimization for any workload that repeats the same system prompt or context. Providers match a prefix of your input against their KV cache — if it matches, you pay the cached rate instead of the full input rate.
Route by Complexity
Not every request needs a flagship model. Build a routing layer that classifies request complexity and sends simple queries to budget models — reserving flagship calls for tasks that genuinely require it. A rule-based classifier (keyword matching, input length, prior failure) can reduce spend by 40–70% without measurable quality loss.
Compress System Prompts
System prompts are sent on every request. Audit yours for redundant instructions, repetitive examples, and verbose formatting rules. Every 1,000 tokens removed saves $2–$20 per 1M requests depending on your model. Use a prompt linter or LLM-assisted compression pass to trim 20–40% without losing quality.
Control Output Length Explicitly
Output tokens cost 3–5× more than input. Without an explicit max_tokens cap, models often generate far more than necessary. Measure your median real output length, set max_tokens at P90 + 20% buffer, and log truncations. This alone saves 30–60% on output-heavy workloads like content generation and code completion.
Use Batch APIs for Async Workloads
OpenAI, Anthropic, and Google all offer async batch APIs with ~50% price discounts. If your workload doesn't need a real-time response (nightly data processing, bulk embeddings, content moderation queues, backfill jobs), batch processing is free money. The only cost is a 24-hour response SLA.
Cache at the Application Layer
Many requests are semantically identical. Build an exact-match or embedding-based cache in Redis or PostgreSQL that intercepts duplicate queries before they reach the API. A 20% cache hit rate on 1M requests/month at $0.01/request = $200K/year in avoided spend.
Measure Real Costs in Your Provider Dashboard
Estimates (including ours) are approximations. Actual billing uses each provider's tokenizer, which can differ 5–20% from cl100k_base estimates for non-English text. Use your provider's usage dashboard to measure actual token counts, then feed that back into your budget planning.
Prune Retrieved Context in RAG
RAG input is dominated by retrieved chunks. Reducing top-k from 10 to 5, re-ranking and dropping marginal chunks, or summarizing retrieved context before passing to the model can cut input tokens by 30–60%. Since input tokens are the majority of RAG cost, this has outsized impact.
Right-Size Your Model for the Task
Audit your model choices by task type every quarter. AI pricing drops ~40% per year on average — a model that was the right budget choice 6 months ago may now have cheaper, better alternatives. Set a calendar reminder to re-evaluate your stack against updated pricing.
Set Budget Alerts and Per-Feature Caps
Runaway agent loops, unexpected traffic spikes, and misconfigured tools are the top causes of surprise AI bills. Set provider-level spend alerts at 50% and 90% of monthly budget. Instrument cost per feature at the code level so you know exactly which features are expensive.
Quick Reference: Tactics by Savings Potential
Highest impact (60–90% savings)
- Prompt caching
- Model routing
- Batch APIs
High impact (30–60% savings)
- Output length caps
- Context pruning
- App-layer caching
Medium impact (10–30% savings)
- Prompt compression
- Quarterly model audits
- Usage monitoring
Frequently Asked Questions
What is the single biggest way to reduce AI API costs?
Prompt caching is typically the highest-leverage optimization — it can reduce input token costs by 60-90% for workloads where the same prefix (system prompt, retrieved documents) is repeated across requests. Most major providers offer cached-input discounts of 50-90% off the standard input rate.
Does switching models actually save significant money?
Yes — the price difference between tiers is enormous. A balanced-tier model like Claude Sonnet 5 ($2/$10 per 1M in/out tokens) costs about 5-10x less than a flagship like o3 ($10/$40). For most production chatbots, RAG pipelines, and extraction workloads, the balanced tier delivers equivalent quality at a fraction of the price.
How much does controlling output length actually save?
Output tokens typically cost 3-5x more than input tokens. Reducing output from 2000 to 800 tokens per response cuts output costs by 60%. At 1M monthly requests, this alone can save thousands of dollars per month on high-output workloads like content generation or code completion.
When does the batch API not make sense?
Batch APIs have a 24-hour SLA, so they're wrong for any user-facing feature that needs a real-time response. They're ideal for nightly data processing, bulk embeddings, dataset analysis, content moderation queues, and other async pipelines where latency doesn't matter.
How do I know which optimization will save me the most?
Start by identifying your dominant cost driver: if input tokens are > 70% of your cost, focus on caching and context pruning. If output tokens dominate, focus on max_tokens limits and switching to cheaper output models. Use our Cost Savings Calculator to model the impact of each optimization before implementing it.