The Setup: Three Tools and One Rule
Budgeting AI spend only needs three things, and two of them are free:
- Monthly usage exports: Every provider (OpenAI, Anthropic, Google, DeepSeek) lets you download a usage CSV. If a vendor doesn't, that's your first red flag.
- A statement generator: I merged my CSVs into Bill Doctor, which maps rows to verified model rates, estimates unknown models, and flags spikes, cache gaps, and concentration risks.
- A feature tag on every row: Which feature generated this call — chat, RAG, extraction? You can't judge a cost until you know what bought it.
The one rule: before paying anything next month, I had to explain every line on this statement. Not "it's AI stuff" — the feature, the model, the reason. That rule is what surfaces the waste.
Five Signs You Need This System
Not sure whether budget-style tracking is worth an hour of your month? If any of these describe you, it is:
- 1You can't name your top-spend model without opening a dashboard
- 2Your bill moved more than 20% in a month and nobody noticed until the invoice
- 3You've never once looked at the cached-input line on your provider invoice
- 4You're on three or more providers and manually adding up the invoices
- 5Your agents retry silently — you don't know your failure rate or retry cost
Every one of these was true for me in week zero. Thirty days later, none of them were — which is the whole pitch: the system works even when the starting point is a mess.
The Statement: One Month, Five Categories
My "AI bank statement" — five views of the same $5,000 month:
By feature
By provider
By model (top 3)
Waste lines
The bottom line: $5,000 spent, ~$2,050 explainable as waste or misallocation — 41% of the month, spent on autopilot. Not fraud, not a bug: just untracked defaults, the AI equivalent of subscriptions you forgot to cancel.
Reading It Like a Personal Budget
The three checks that turned the statement into actions:
Concentration check: one model = 68%
Sonnet 5 was serving everything — including extraction, which a budget tier does equally well. Budgeters call this 'putting everything on one card'; it hid the fact that 25% of spend was on the wrong tier.
Cache gap: 40% cached, ~90% cacheable
The chat assistant re-sends the same instruction sheet 100,000 times a month. With caching markers enabled, that line drops by ~$720. The money wasn't missing — it was never turned on.
Waste check: 23% in retries and failed calls
Retries after rate limits and agent tool failures re-bill full context. Exponential backoff plus a fallback model cuts this in half — about $575/month of the $1,150.
Week by Week: What Each Week Taught Me
The thirty days weren't one big audit — they were four small ones, one per week, each building on the last:
Week 1 — The baseline shock
Assembled the first statement and realized I couldn't explain roughly 30% of the lines. Installed the tagging rule (every row = feature + model + reason) before spending another dollar.
Week 2 — The cache gap
The dashboard showed cached input at ~0% of input tokens. Enabled caching markers on the chat assistant; the first ~$500 came back within days. No code redesign, one sprint item.
Week 3 — The retry tax
Logs showed a 2.4% failure rate with three blind retries each. Added exponential backoff and a Claude Haiku 4.5 fallback; roughly $300 recovered and error rates dropped with it.
Week 4 — The routing decision
Two weeks of escalation logs proved extraction could downgrade safely. Routed it to GPT-5.6 Luna, re-ran the statement: $5,000 → $3,350. The month closed 33% down with zero feature changes.
The pattern is the point: every week's number came from a different screen — the dashboard, the logs, the logs again, then a routing change. None of it required rewriting the product.
What I Changed (In Order of Impact)
Routed the extraction pipeline to GPT-5.6 Luna
$1,250/mo feature → ~$350/mo. Quality held; we kept Sonnet 5 on the chat assistant where tone matters.
Enabled prompt caching on the chat assistant
One developer sprint, one cache marker. Recovered ~$720/mo of the cache gap.
Added retry backoff + fallback routing
Rate-limit retries now wait and escalate to a cheaper model instead of re-hammering Sonnet 5. Cut ~$575/mo of the waste line.
Set per-feature budgets and alerts
Extraction now has a hard monthly ceiling and a 50%/90% alert. The invoice can no longer surprise me.
Month two: $3,350 — down 33%, with the same features and no user-facing quality change. The routing details live in the downgrade playbook; the retry math is in the hidden-tax analysis.
The 50/30/20 Rule for AI Budgets
Once you can see the categories, give them a target. The budgeter's 50/30/20 translates to AI spend cleanly:
Chat, RAG, extraction — whatever the product sells. This bucket gets the best models and the strictest SLAs.
New features, evals, prototype loops. Cap it monthly so experiments can't become a second production bill.
Spikes, retries, subscriptions, and the month-end audit itself. The buffer is what absorbs surprises without tripping alerts.
The rule's job is allocation, not permission: experiments don't get cancelled, they get budgeted — which is the whole difference between AI spend as a cost center and AI spend as a line item you actually manage.
Month Two: When the System Starts Catching Leaks
The second statement came out at $3,350, and the shape had changed, not just the total:
- Cache share: ~0% → 78% of input: Caching markers plus a frozen prompt prefix turned the chat assistant's instruction sheet into a 90%-discounted line item.
- Waste line: $1,150 → ~$430: Backoff, argument validation, and the Haiku fallback cut retries and failed calls by two-thirds without touching the product.
- Concentration: 68% → 41%: Extraction and classification moved to GPT-5.6 Luna; Sonnet 5 now serves what it's actually good at.
The month also proved the system's second job: catching leaks, not just finding waste. A test account hit the production model through a leaked key mid-month, and the 50% budget alert fired within a day. We rotated the key and stopped what would have been a mystery $400 spike on the next invoice — the kind of line that, pre-system, got blamed on "AI being expensive."
The ritual took 40 minutes instead of 60 — most of the categories were already tagged. That is the real metric: the first month found the money, and every month after it is insurance.
The 12-Category Chart of Accounts for AI Spend
Every budgeting system needs a chart of accounts, and AI spend is no different. If a category doesn't exist in your tracking yet, it still exists in your bill:
| Category | What counts | Budget rule of thumb |
|---|---|---|
| Chat assistant | Conversational UI, support bot | < 40% of bill |
| RAG search | Retrieval + answer synthesis | < 25% of bill |
| Extraction | Structured data pipelines | < 15% of bill |
| Agents | Multi-turn loops, tool calls | Cap at 25 turns, track per session |
| Batch | Async and nightly jobs | Should be at least 50% off |
| Embeddings | Vectorization, retrieval index | < 5% of bill |
| Vision & image | Image/video input and generation | Track separately |
| Evals & testing | CI runs, regression suites | Flat monthly cap |
| Retries & errors | Failed calls re-billed | < 10% of bill |
| Cache gap | Input that should have been cached | < 10% of bill |
| Tier premium | Flagship where budget would do | < 30% of bill |
| Fixed subscriptions | Cursor, tools, other AI vendors | Audit quarterly |
Two rows deserve special attention. Cache gap is the friendliest category in the ledger — the money is already spent, but the fix is a developer sprint, not a strategy change. Tier premium is where the biggest dollar amounts hide, and it is covered end to end in the model routing playbook.
The Monthly Ritual: 45 Minutes, 3 Screens
Once the system exists, it's a repeating ritual — about 45 minutes on the first of the month:
Screen 1 — Dashboards (15 min)
Open each provider's usage page. Check spend against the 50%/90% budget alerts and flag anything that moved more than 20% month over month.
Screen 2 — Bill Doctor (15 min)
Upload the merged CSVs. Read the findings: spikes, cache gaps, model concentration, unknown model IDs. Pick the top three to act on.
Screen 3 — Targets + Wrapped card (15 min)
Set three next-month targets (one per feature). Generate the AI Cost Wrapped recap card and share it — public accountability keeps the ritual honest.
Forty-five minutes a month is the cheapest optimization in the AI stack. It found $2,050 in month one and has kept the bill trending down ever since — because the number that gets reviewed monthly is a number that gets managed.
Alert configuration is part of the ritual, not an add-on: set provider-level alerts at 50% and 90% of the monthly budget, and a per-feature cap on anything experimental. The 50% alert gives you a mid-month nudge while there is still time to act; the 90% alert is the emergency brake. In month two, the 50% alert is what caught the leaked-key incident before it hit the invoice.
Run It Yourself: The 5-Step Template
Download last month's usage CSV from every provider you use.
Upload to Bill Doctor (or merge into a sheet) — it maps models to verified rates and flags anomalies.
Tag every row with the feature that generated it: chat, RAG, extraction, agent, batch.
Run the three checks: concentration (one model >60%?), cache gap (cached vs cacheable), waste (retries + errors).
Pick the three biggest findings and set one action and one budget per feature for next month. Re-run monthly.
Done tracking a month? Generate your own shareable recap card — spend, token mix, cache savings, and the cheapest equivalent workload — and hold yourself accountable to it publicly. That's what this month's card is for.
Frequently Asked Questions
How do I track AI API costs?
Export the usage CSV from each provider (OpenAI, Anthropic, Google, DeepSeek all support it), merge the rows into one sheet, and categorize by feature and model. Bill Doctor does the heavy lifting automatically — it maps rows to verified rates, estimates unknown models, and flags spikes and cache gaps.
What is a healthy model concentration?
If one model is more than ~60% of your spend, you are almost certainly overpaying on some tasks. In our tracked month, Claude Sonnet 5 was 68% of the bill — routing just the extraction pipeline to a budget tier fixed the concentration and cut the bill by 18%.
How much should a startup spend on AI APIs per month?
There is no universal number — the right frame is unit economics: cost per request or per feature vs the value that feature creates. Start by tracking, then set a per-feature budget. Most teams we see find 20-30% of the bill is avoidable waste within the first month of tracking.
What is the 'cache gap'?
The difference between the input tokens that could be cached and the input tokens actually billed at the cached rate. In our tracked month, only 40% of cacheable input was cached — a $900 gap. Fixing it is usually developer-side: enable caching and keep the prompt prefix stable.
How do I do AI cost tracking as a non-technical founder?
You don't need to read code — you need two exports: a usage CSV from your provider and a calculator that understands it. Bill Doctor turns the CSV into findings, and AI Cost Wrapped gives you a shareable monthly recap card. The whole loop takes about an hour a month.
How often should I check AI costs?
A weekly glance at the provider dashboards plus a 45-minute monthly ritual (review, Bill Doctor upload, three targets, Wrapped card). Set provider-level spend alerts at 50% and 90% of budget so a runaway month alerts you instead of surprising you.