Why Loops Bill Differently Than Chats
A chat request is one transaction: your prompt, your history, one response. An agent is a loop — the model calls a tool, receives the result, plans, calls again, and only finishes when it decides it's done. Each iteration is a full API request that re-sends everything before it:
- The system prompt — again, at every turn
- The entire conversation so far — growing every turn
- Tool call outputs, appended as new input tokens
- New output tokens for every planning step, even invisible ones
The arithmetic consequence: cost scales with turns × tokens × price, and most agent frameworks make all three larger than they need to be. Two identical features can differ 30× in running cost purely because of loop design and model choice.
The Formula: Calculate It Before You Ship
One session's cost is the sum over turns, with tool results priced as input and retries priced as extra turns:
Simplified: N turns × (I in + O out) × per-token price. N is the variable you control least — and the one that moves the bill most. Once you can estimate N, I, and O for a typical session, plug the numbers into the AI cost calculator and multiply by your projected sessions per month.
Anatomy of One Turn: Where the Tokens Go
To see why real bills exceed the simple estimate, open one turn of a running agent and count what gets sent. A typical turn in a production loop — system prompt, growing history, one tool result, one new tool call — looks like this:
| Turn component | Tokens | Billed as |
|---|---|---|
| System prompt (persona + tool schemas) | 2,000 | Input |
| Conversation history (prior turns) | 3,200 | Input |
| Last assistant message | 400 | Input |
| Tool result | 800 | Input |
| New instruction / tool-call JSON | 500 | Input |
| Completion (incl. reasoning) | 500 | Output |
Note the ratio: only ~7% of the input is the user's actual new instruction. Everything else is context being re-billed. That is exactly why the worked examples below call the simple model an estimate — measure your real context per turn before projecting a monthly budget, or double your projection to be safe.
Worked Example: The Same 20-Turn Loop on Three Models
Take a realistic agent session: 20 turns, 2,000 input and 500 output tokens per turn. Tool call results are priced separately below. Prices verified 2026-08-19 (DeepSeek off-peak rates are 50% of peak):
| Model | Rate ($/1M in, out) | Per session | 10K sessions/mo |
|---|---|---|---|
| GPT-5.6 Sol | $5 / $30 | $0.50 | $5,000 |
| GPT-5.6 Terra | $2 / $12 | $0.20 | $2,000 |
| Claude Sonnet 5 | $2 / $10 | $0.18 | $1,800 |
| DeepSeek V4 Flash (peak) | $0.44 / $1.32 | $0.031 | $310 |
| DeepSeek V4 Flash (off-peak) | $0.22 / $0.66 | $0.015 | $154 |
The same loop costs over 30× more on the flagship than on the budget model at off-peak hours. Neither choice is objectively wrong — quality and latency differ — but the decision should be made with this table visible. For many agentic workloads (research, data extraction, code with tests), the cheaper tier is indistinguishable in outcome.
The Runaway Loop: Where Budgets Die
Agents don't always finish in 20 turns. A missing tool, an ambiguous goal, or a validation error can send the model back around: turn 30, turn 40, turn 60. Because every turn re-bills the full context, the cost scales linearly with turn count — and nothing stops it.
The failure mode isn't the model — it's the harness. Teams rarely ship a loop with a hard iteration cap, a per-session token budget, or circuit breakers on repeated failures. A 20-turn default with a hard cap of 25 converts a $15,800 surprise into a $5,300 line item that fails loudly instead of silently draining the account.
Related reading: the same failure resends that inflate agent loops are one of the five hidden waste sinks in our wasted-token analysis.
The 4 Loop Failure Patterns That Inflate Turn Count
Runaway loops are rarely random — they follow four recurring patterns. Each is cheap to fix and each is worth turns-per-session on every run:
1. Tool schema ambiguity
Vague parameter descriptions make the model guess argument shapes, producing malformed calls that fail and re-enter the loop. Fix: strict schemas, example values, and enum constraints on every tool.
2. Missing termination criteria
Without an explicit 'you are done when X' condition, a capable model keeps proposing next steps forever. Fix: write the done-condition into the system prompt and enforce a max-turn cap regardless.
3. Error cascades
When a tool is down, the recovery turn also fails, then the retry fails — each iteration re-billing full context. Fix: a circuit breaker that switches to a fallback model or asks the user after two consecutive failures.
4. Context truncation
When history exceeds the window, frameworks silently drop the oldest turns, the model loses the thread, and starts asking clarifying questions that add turns. Fix: rolling summaries and pruning irrelevant tool results instead of blind truncation.
These four patterns account for most 40+ turn sessions we see in the wild. Fixing just pattern 1 and 2 typically halves the average turn count — which, at $1.05 per 40-turn session on Sol, is the difference between $10,500 and ~$5,250/month at 10K sessions.
The Tool Call Tax: Results and Failures
Tool calls have a double cost. The happy path: each result is appended as input and re-sent on every later turn. The failure path: a failed call re-enters the loop with the full context plus the error text plus a new recovery response. In our example, four tool calls with 800-token results append ~3,200 tokens per session — and because results are re-sent on each later turn, their real footprint is roughly 8× that: about $0.13 of the $0.50 Sol session (~$1,300/month at 10K sessions), and only ~$56/month on DeepSeek V4 Flash off-peak.
Failed calls are the expensive variant. At a 25% tool failure rate with one full-context resend per failure (10,000-token context), on Sol:
Mitigation is cheap: validate arguments before the model fires the call, cache tool results for repeat calls, and format errors so the model can recover in one turn instead of three.
Loop Economics at a Glance
The same 20-turn loop (2K in / 500 out per turn, 10,000 sessions a month) at three design points — the whole article in one table:
| Design point | Setup | Per session | Per month |
|---|---|---|---|
| Budget loop | DeepSeek V4 Flash off-peak + cached prefix | ~$0.007 | ~$70 |
| Balanced loop | GPT-5.6 Terra + cached prefix | ~$0.13 | ~$1,280 |
| Flagship loop | GPT-5.6 Sol, no caching | $0.50 | $5,000 |
The spread between design points is ~70× — the same feature, the same harness, a different budget line. Nobody chooses the expensive loop deliberately; it is always the accumulation of defaults: no cache markers, no turn cap, no fallback routing. Each one is a one-line fix.
Caching Changes the Loop Economics
The system prompt and tool schemas are identical at every turn — exactly what prompt caching is designed for. With a 90% cached-input discount (OpenAI, Anthropic, Google), the 2,000 input tokens per turn drop from $5/1M to $0.50/1M on Sol:
Output is untouched by caching, which is why output-heavy loops still reward a cheaper model. The combined play — cache the prefix and run on a budget tier — is what turns $5,000/month into ~$160/month for the same loop. DeepSeek's cache-hit rate ($0.014 vs $0.44 input) is the most aggressive on the market right now.
Why 1M-Token Context Doesn't Fix Cost
Bigger context windows are a quality feature, not a cost feature — the loop re-bills the full context on every turn, so a retrieval-heavy agent with 80,000 tokens per turn scales the bill linearly. Compare it with the short-context loop from earlier, on GPT-5.6 Terra ($2/$12):
One common misconception: a model with a 1M-token window (like the free Ox Alpha preview or Gemini 3.1 Pro) makes long-horizon agents possible, but it does not make them cheap — every turn still re-bills whatever context you feed it. The cost lever is context discipline: cache the stable prefix, prune retrieved chunks between turns, and summarize history instead of carrying it.
The Pre-Ship Cost Checklist
Run this before your agent reaches production — every item is minutes of work and months of avoided spend. A quick ceiling table helps: pick your worst-case turn count, multiply by the per-turn price, and make that number your per-session alert threshold (2K in / 500 out per turn, no caching):
| Max turns | GPT-5.6 Terra | GPT-5.6 Sol | DeepSeek V4 Flash (off-peak) |
|---|---|---|---|
| 10 | $0.10 | $0.25 | $0.008 |
| 25 | $0.25 | $0.63 | $0.019 |
| 50 | $0.50 | $1.25 | $0.039 |
| 100 | $1.00 | $2.50 | $0.077 |
Hard iteration cap
Max turns per task (e.g., 25). Test where your quality curve flattens; cap slightly above it.
Per-session token budget
Abort or downgrade when a session exceeds N tokens. One metric, enforced in the harness.
Prompt caching on the stable prefix
System prompt + tool schemas must be byte-stable and cached. 90% off on most providers.
Validate tool arguments
Catch malformed calls before the model fires them; format errors to recover in one turn.
Fallback routing on failure
After 1 retry, escalate to a cheaper model instead of re-hammering the same endpoint.
Log turns and tokens per session
Export to a CSV your cost tooling understands — Bill Doctor reads provider exports directly.
Frequently Asked Questions
Why is an AI agent more expensive than a chatbot?
A chatbot bills one request per user message. An agent loops: every turn re-sends the full conversation plus tool results, and each tool call or failure can trigger another turn. A 20-turn agent session bills roughly 20 requests worth of context — before counting tool outputs and retries.
How much does one AI agent session cost?
For a 20-turn session with 2,000 input and 500 output tokens per turn: about $0.50 on GPT-5.6 Sol, $0.20 on GPT-5.6 Terra or Claude Sonnet 5, and $0.015 on DeepSeek V4 Flash at off-peak rates. At 10,000 sessions per month that is $5,000, $2,000, or $154.
Do tool call results cost tokens?
Yes. Tool outputs are appended to the conversation and billed as input tokens on every subsequent turn. A session with four tool calls and 800-token results adds roughly 3,200 input tokens per session — and failed calls add a full resend plus new output.
How do I stop a runaway agent loop?
Set a hard iteration cap per task, add a token budget check between turns, and route recovery to a cheap fallback model instead of retrying the same model. Log turn count per session from day one so a regression shows up in your metrics, not your invoice.
What is the cheapest way to run agents?
The combination of a budget model (DeepSeek V4 Flash off-peak, GPT-5.6 Luna, Gemini 3.5 Flash-Lite), prompt caching on the stable prefix, strict turn caps, and validated tool arguments — in that order. We measured the same 20-turn loop at $0.50 vs $0.015 per session depending on those choices.
How do long-context agents change the cost math?
Every turn re-bills the full context, so a retrieval-heavy agent with 80,000 tokens per turn costs about $9.20 per session on GPT-5.6 Terra — roughly 46x a short-context loop. Prompt caching is the critical lever: with 90% cached input the same session drops to about $2.00. Long context windows help quality, never cost.