An agent plans, calls tools and iterates: every step re-sends the growing conversation, so input tokens compound across steps before the final answer.
Input / request
40,000
Output / request
1,500
Cacheable input
65%
Requests / user / day
8
Cost by model
| Model | Per request | Per user / month | 500 users / month | vs cheapest |
|---|---|---|---|---|
| bestGemini 2.5 Flash-Lite | $0.00226 | $0.5424 | $271.20 | — |
| Ministral 3 (8B) | $0.002715 | $0.6516 | $325.80 | 1.2× |
| Mistral Small 4 | $0.00339 | $0.8136 | $406.80 | 1.5× |
| GPT-5.6 Luna | $0.00512 | $1.229 | $614.40 | 2.3× |
| GPT-5.4 nano | $0.005195 | $1.247 | $623.40 | 2.3× |
| Codestral | $0.00633 | $1.519 | $759.60 | 2.8× |
| DeepSeek V4 Flash | $0.008504 | $2.041 | $1k | 3.8× |
| Gemini 3.5 Flash-Lite | $0.00873 | $2.095 | $1k | 3.9× |
| Gemini 2.5 Flash | $0.00873 | $2.095 | $1k | 3.9× |
| Mistral Large 3 | $0.0106 | $2.532 | $1.3k | 4.7× |
| GPT-5.4 mini | $0.0192 | $4.608 | $2.3k | 8.5× |
| Claude Haiku 4.5 | $0.0241 | $5.784 | $2.9k | 10.7× |
| DeepSeek V4 Pro | $0.0256 | $6.135 | $3.1k | 11.3× |
| Grok 4.3 | $0.0265 | $6.348 | $3.2k | 11.7× |
| o4-mini | $0.0292 | $6.996 | $3.5k | 12.9× |
| GLM 5.2 | $0.0298 | $7.162 | $3.6k | 13.2× |
| GPT-5.1 | $0.0358 | $8.58 | $4.3k | 15.8× |
| Gemini 2.5 Pro | $0.0358 | $8.58 | $4.3k | 15.8× |
| Gemini 3.6 Flash | $0.0362 | $8.676 | $4.3k | 16.0× |
| Mistral Medium 3.5 | $0.0362 | $8.676 | $4.3k | 16.0× |
| Gemini 3.5 Flash | $0.0384 | $9.216 | $4.6k | 17.0× |
| Grok 4.5 | $0.0448 | $10.752 | $5.4k | 19.8× |
| Claude Sonnet 5 | $0.0482 | $11.568 | $5.8k | 21.3× |
| Grok 4.6 | $0.05 | $12.00 | $6k | 22.1× |
| GPT-5.2 | $0.0501 | $12.012 | $6k | 22.1× |
| GPT-5.6 Terra | $0.0512 | $12.288 | $6.1k | 22.7× |
| Gemini 3.1 Pro | $0.0512 | $12.288 | $6.1k | 22.7× |
| o3 | $0.053 | $12.72 | $6.4k | 23.5× |
| GPT-4.1 | $0.053 | $12.72 | $6.4k | 23.5× |
| Muse Spark | $0.0564 | $13.53 | $6.8k | 24.9× |
| GPT-5.4 | $0.064 | $15.36 | $7.7k | 28.3× |
| Claude Sonnet 4.5 | $0.0723 | $17.352 | $8.7k | 32.0× |
| Kimi K3 | $0.0723 | $17.352 | $8.7k | 32.0× |
| Claude Opus 5 | $0.1205 | $28.92 | $14.5k | 53.3× |
| Claude Opus 4.5 | $0.1205 | $28.92 | $14.5k | 53.3× |
| GPT-5.6 Sol | $0.128 | $30.72 | $15.4k | 56.6× |
| GPT-5.5 | $0.128 | $30.72 | $15.4k | 56.6× |
| Claude Fable 5 | $0.241 | $57.84 | $28.9k | 106.6× |
Assumes the typical cacheable share (65% of input at cached rates where published). 8 requests/user/day. Adjust everything in the monthly calculator.
How to spend less
Related workflow costs
FAQ
Using a typical profile of 40,000 input and 1,500 output tokens with 65% cacheable input: from $0.00226 on Gemini 2.5 Flash-Lite up to $0.241 on Claude Fable 5. See the table for every model.
At 8 requests per user per day and the typical token profile, budget from $0.5424 per active user per month on Gemini 2.5 Flash-Lite. A team of 500 active users lands around $271.20/month at that tier. Model your exact numbers in the monthly cost calculator.
Prompt caching is the single biggest lever — each step re-reads prior context. Summarize or prune tool outputs before appending them to the transcript. Cap the step budget; runaway loops are the #1 surprise on agent invoices.