What Prompt Caching Is (No Jargon)
When your product calls an AI model, it doesn't just send the user's question. It sends a package: your instruction sheet (the persona, tone rules, formatting), the conversation history, and often documents. That package is billed by the token — every word counts.
Here's the thing nobody tells you at signup: the first part of that package is identical on every single call. User #1,247 asks "refund policy?" and user #1,248 asks "shipping time?" — but both times your app sends the exact same instruction sheet. The providers noticed, and built prompt caching: they remember the repeated part and charge a fraction of the normal price for it.
That's it. Same text, sent again and again, bought at a bulk discount. If you've ever seen "cached input" on an invoice and skipped past it — that line is your best friend.
Why Your App Pays Full Price for the Same Words
Three things get re-sent on nearly every request:
- The instruction sheet: Your system prompt — brand voice, rules, tool definitions. It is probably longer than you think (1,000–10,000+ words).
- Conversation history: Every turn of a chat re-sends all previous turns, so 10-turn conversations bill their history 9 extra times.
- Shared documents: Policy PDFs, product specs, help-center articles — attached to every relevant request.
Without caching, you pay the premium rate for all of it, every time. With caching, the repeated part drops to 3–10% of its normal price. The difference is usually the single biggest line item on an AI bill — see the system prompt tax in our hidden-cost guide for the full math.
The Cached-Token Price List (Verified 2026-08-19)
Per 1 million input tokens, standard vs cached, for a representative model from each provider:
| Provider | Example model | Standard input | Cached input | Discount |
|---|---|---|---|---|
| OpenAI | GPT-5.6 Sol | $5.00 | $0.50 | 90% off |
| Anthropic | Claude Sonnet 5 | $2.00 | $0.20 | 90% off |
| Gemini 3.1 Pro | $2.00 | $0.20 | 90% off | |
| DeepSeek | DeepSeek V4 Flash | $0.44 | $0.014 | 97% off |
DeepSeek's 97% discount is the most aggressive right now — one reason its V4 Flash is the cheapest serious option for high-volume products. Cached rates also stack with batch processing (~50% more off) where you don't need instant answers.
How Much You Can Cache Depends on Your Product
Different products repeat different amounts of text. These are typical cacheable shares of input for common product types, with the monthly saving at 100,000 requests on Claude Sonnet 5 with a 4,000-token stable prefix (rounded):
| Product type | Cacheable share | What gets cached | Saving / month |
|---|---|---|---|
| Agent loop | 70–90% | System prompt + tool schemas | ~$550 |
| Chatbot | 60–80% | System prompt + stable history prefix | ~$500 |
| Batch / nightly jobs | 50–70% | Parent task prefix + batch discount | ~$430 |
| RAG search | 40–60% | System prompt + repeated chunks | ~$350 |
| Extraction pipeline | 30–50% | System prompt + output schema | ~$280 |
Two takeaways for a founder reading this: agents and chatbots are the biggest prizes because their repeated prefix is huge and stable; and if your product is a search or pipeline, expect a smaller — but still very real — share. Scale the table linearly: at 1M requests, multiply every row by ten.
The Cache-Breaking Mistakes Founders Make
Caching fails silently: the bill just doesn't move. The usual culprits — each one looks innocent in a code review:
Dynamic content at the start of the prompt
A timestamp, request ID, or date in the first 100 tokens changes the prefix on every call — no match, no discount. Rule: stable front, dynamic tail.
History re-serialized differently each call
Same data, different order or formatting per request defeats the byte-for-byte match. Freeze the serialization format.
Per-user template prefixes
1,000 users with slightly different instruction variants produce 1,000 distinct prefixes — none repeated enough to cache. Share one template, personalize in the tail.
Cosmetic edits between deploys
Trailing whitespace, reordered bullets, 'minor tweaks' to the system prompt invalidate every cached copy until traffic rebuilds it. Batch prompt changes.
Missing cache markers on OpenAI and Anthropic
Google and DeepSeek cache automatically; OpenAI and Anthropic need explicit markers in code. If nothing was marked, nothing is cached.
The diagnostic is the dashboard: if the cached-input line is near zero two days after enabling caching, one of these five is active.
Three Budgets, One Page: What Caching Is Worth to You
The quick estimate: if input is ~40% of your bill and ~60% of that input is cacheable at 90% off, caching recovers roughly 21% of your total bill. Applied to three budgets:
| Monthly AI bill | Input share (40%) | Cacheable (60%) | Recovered / month |
|---|---|---|---|
| $500 / month | $200 | $120 | ~$110 |
| $5,000 / month | $2,000 | $1,200 | ~$1,080 |
| $50,000 / month | $20,000 | $12,000 | ~$10,800 |
Your mix will differ — output-heavy products recover less (caching only touches input), and chat-heavy products recover more. But the order of magnitude is right, and the effort is the same for every budget: one sprint item called "enable caching."
The $800 → $80 Example
Suppose your app sends a 4,000-token instruction sheet (roughly a 3,000-word document — normal for a chatbot with brand rules and tool definitions) on 100,000 requests per month, on Claude Sonnet 5:
Double the instruction sheet to 8,000 tokens and it's $1,440/month recovered. At 1M requests — the scale where bills get scary — it's $7,200/month. Caching is not a micro-optimization; it is often the difference between profitable and not.
One nuance for high-volume products: cached prefixes expire after minutes to hours, and right after expiry the next batch of traffic pays full price until the cache rebuilds. On steady traffic the rebuild takes seconds and the discount is effectively continuous — but if you see the cached-input share dip periodically, check expiry timing against your traffic pattern before assuming caching is broken.
What Caching Does NOT Fix
Caching is the highest-leverage lever, but it has hard limits. Three things it won't touch:
Output tokens
The text the model writes is never discounted. Output-heavy features (long articles, code generation) need model choice, not caching — that's the model routing playbook.
Cold-start traffic
After a deploy, a prompt rewrite, or a cache expiry, the first requests rebuild the cache at full price. Batch your prompt changes so you pay that tax once, not weekly.
Per-user unique prefixes
If every user gets a slightly different instruction template, no prefix is repeated enough to matter. One shared template with personalization in the tail is the only way.
The 30-Second Caching Audit for Founders
Five yes/no questions — thirty seconds in your provider dashboard, no code required:
- 1Does your dashboard show a cached-input (or cache-read) line?
- 2Is that line more than ~20% of your input tokens?
- 3Does the instruction sheet start with stable text (no timestamps or IDs)?
- 4Did your developer add cache markers for OpenAI and Anthropic?
- 5Is dynamic content (user data, dates) in the tail of the prompt, not the front?
A "no" on any of the five is a one-sprint item with a triple-digit monthly payoff. A "yes" on all five and the cached-input line is still tiny? That's the audit finding to paste into your developer's ticket — it means the code serializes the prompt differently on every call.
Two Questions to Ask Your Developer This Week
Is prompt caching enabled on every model we use?
For Google and DeepSeek it may already be automatic. For OpenAI and Anthropic it needs markers in the code — often a 5-minute change your developer can ship this sprint.
Can you keep the instruction sheet stable?
Caching only works when the repeated text is byte-identical. If your code adds a timestamp or random formatting to the prompt, the cache misses and you pay full price. One rule: keep the front of the prompt stable, put dynamic content at the end.
Then verify in your provider dashboard: the cached input (or "cache read") line should be a large share of your input tokens within a few days. If it's near zero, caching is off — or your prompt isn't stable.
What Not to Put in the Cacheable Zone
Cached prefixes are held briefly by the provider, then expire — but "briefly" is long enough to matter for sensitive data. Keep API keys, passwords, personal data, and unpublished content out of the stable front of the prompt, and check each provider's retention and training policy. For free or anonymous model routes (like OpenRouter's stealth previews), assume the strictest policy — our privacy tradeoff guide lists what to verify before sending anything sensitive.
What a Healthy Cached-Input Line Looks Like
A week after enabling caching, check the ratio of cached to total input tokens. Healthy ranges by product type:
- Chat and agents: 50–80% of input cached. The instruction sheet plus stable history is most of the context — this is the range that says caching is working.
- RAG: 30–60%. Retrieval brings fresh chunks every call, which lowers the ceiling. Below 30% is suspicious.
- Extraction and batch: 30–50%. The schema and instructions repeat; the data doesn't. Batch adds the ~50% async discount on top.
Below those ranges, the fix is almost always one of the cache-breaking mistakes above — dynamic prefixes being the most common. Above them, you're leaving the discount on the table in the other direction: move the tail (user data) out of the cached zone so more of each request hits the cheaper rate.
Jargon Buster: The Words on Your Bill
| Term | What it means |
|---|---|
| Token | The billing unit — roughly 3/4 of an English word |
| Input tokens | Text sent to the model (your package of instructions and history) |
| Output tokens | Text the model writes back — priced 3–5x higher than input |
| System prompt | Your app's instruction sheet, re-sent on every request |
| Cache hit | Repeated prefix recognized — billed at the discounted cached rate |
| Cache miss | Prefix not recognized — billed at the full rate |
| Cache breakpoint | A code marker telling the provider where caching starts (OpenAI, Anthropic) |
| Batch API | Discounted asynchronous processing with up to 24-hour turnaround |
If a vendor's invoice or docs use terms not on this table, that's a signal to ask for a plain-English explanation before you scale — surprising terminology on an AI bill is almost always surprising cost.
Frequently Asked Questions
What is prompt caching in plain English?
When your app calls an AI model, it sends a big text package: instructions, conversation history, maybe documents. Providers noticed the first part of that package is identical on every call, so they now remember it and charge far less for the repeated portion — usually 90% less, and up to 97% less on DeepSeek.
How much can prompt caching save me?
In our worked example, a 4,000-token instruction sheet sent 100,000 times per month costs $800/month on Claude Sonnet 5 at the normal rate and only $80/month when cached — $720/month, over $8,600 per year. Most apps can cache 60-90% of their repeated input.
Do I need to be technical to benefit from prompt caching?
For Google and DeepSeek, caching is automatic — you benefit the moment your developer ships it. For OpenAI and Anthropic, your developer needs to add small markers in the code (a cache breakpoint). Neither requires a redesign; ask for it in your next sprint.
Is my data safe if it's cached?
Cached prefixes are held transiently by the provider and typically expire within minutes to hours. Still, avoid putting secrets, personal data, or unpublished material in the stable part of your prompt, and check each provider's retention policy — our free-model privacy guide covers what to verify.
What is the difference between prompt caching and the batch API?
Prompt caching discounts repeated input tokens on any request. The batch API is a separate 50% discount for jobs you don't need back immediately (nightly processing, bulk analysis). They stack: cached tokens in a batch job get both discounts.
What share of my bill should be cached input?
For chat and agent products, 50-80% of input tokens should be billed at the cached rate within a few days of enabling caching. If the cached-input line stays below about 20% of input, caching is off or the prompt prefix isn't byte-stable — check the cache markers and how history is serialized.
Does prompt caching affect response quality?
No. The model still reads the full context on every request — caching only changes how the repeated prefix is billed. Cached and uncached responses are identical. The only noticeable side effect is occasional cold-start latency right after the cache expires, which is seconds on steady traffic.