Retrieval-augmented generation embeds a query, retrieves top-k chunks, and asks the model to answer grounded in those chunks. Input is dominated by retrieved context.
Input / request
8,000
Output / request
500
Cacheable input
50%
Requests / user / day
10
Cost by model
| Model | Per request | Per user / month | 500 users / month | vs cheapest |
|---|---|---|---|---|
| bestGemini 2.5 Flash-Lite | $0.00064 | $0.192 | $96.00 | — |
| Ministral 3 (8B) | $0.000735 | $0.2205 | $110.25 | 1.1× |
| Mistral Small 4 | $0.00096 | $0.288 | $144.00 | 1.5× |
| GPT-5.6 Luna | $0.00148 | $0.444 | $222.00 | 2.3× |
| GPT-5.4 nano | $0.001505 | $0.4515 | $225.75 | 2.4× |
| Codestral | $0.00177 | $0.531 | $265.50 | 2.8× |
| DeepSeek V4 Flash | $0.002476 | $0.7428 | $371.40 | 3.9× |
| Gemini 3.5 Flash-Lite | $0.00257 | $0.771 | $385.50 | 4.0× |
| Gemini 2.5 Flash | $0.00257 | $0.771 | $385.50 | 4.0× |
| Mistral Large 3 | $0.00295 | $0.885 | $442.50 | 4.6× |
| GPT-5.4 mini | $0.00555 | $1.665 | $832.50 | 8.7× |
| Claude Haiku 4.5 | $0.0069 | $2.07 | $1k | 10.8× |
| Grok 4.3 | $0.00705 | $2.115 | $1.1k | 11.0× |
| DeepSeek V4 Pro | $0.007436 | $2.231 | $1.1k | 11.6× |
| o4-mini | $0.0077 | $2.31 | $1.2k | 12.0× |
| GLM 5.2 | $0.00836 | $2.508 | $1.3k | 13.1× |
| Gemini 3.6 Flash | $0.0104 | $3.105 | $1.6k | 16.2× |
| Mistral Medium 3.5 | $0.0104 | $3.105 | $1.6k | 16.2× |
| GPT-5.1 | $0.0105 | $3.15 | $1.6k | 16.4× |
| Gemini 2.5 Pro | $0.0105 | $3.15 | $1.6k | 16.4× |
| Gemini 3.5 Flash | $0.0111 | $3.33 | $1.7k | 17.3× |
| Muse Spark | $0.0121 | $3.638 | $1.8k | 18.9× |
| Grok 4.5 | $0.0122 | $3.66 | $1.8k | 19.1× |
| Grok 4.6 | $0.013 | $3.90 | $2k | 20.3× |
| Claude Sonnet 5 | $0.0138 | $4.14 | $2.1k | 21.6× |
| o3 | $0.014 | $4.20 | $2.1k | 21.9× |
| GPT-4.1 | $0.014 | $4.20 | $2.1k | 21.9× |
| GPT-5.2 | $0.0147 | $4.41 | $2.2k | 23.0× |
| GPT-5.6 Terra | $0.0148 | $4.44 | $2.2k | 23.1× |
| Gemini 3.1 Pro | $0.0148 | $4.44 | $2.2k | 23.1× |
| GPT-5.4 | $0.0185 | $5.55 | $2.8k | 28.9× |
| Claude Sonnet 4.5 | $0.0207 | $6.21 | $3.1k | 32.3× |
| Kimi K3 | $0.0207 | $6.21 | $3.1k | 32.3× |
| Claude Opus 5 | $0.0345 | $10.35 | $5.2k | 53.9× |
| Claude Opus 4.5 | $0.0345 | $10.35 | $5.2k | 53.9× |
| GPT-5.6 Sol | $0.037 | $11.10 | $5.5k | 57.8× |
| GPT-5.5 | $0.037 | $11.10 | $5.5k | 57.8× |
| Claude Fable 5 | $0.069 | $20.70 | $10.4k | 107.8× |
Assumes the typical cacheable share (50% of input at cached rates where published). 10 requests/user/day. Adjust everything in the monthly calculator.
How to spend less
Related workflow costs
FAQ
Using a typical profile of 8,000 input and 500 output tokens with 50% cacheable input: from $0.00064 on Gemini 2.5 Flash-Lite up to $0.069 on Claude Fable 5. See the table for every model.
At 10 requests per user per day and the typical token profile, budget from $0.192 per active user per month on Gemini 2.5 Flash-Lite. A team of 500 active users lands around $96.00/month at that tier. Model your exact numbers in the monthly cost calculator.
retrieved-context dominates cost — tune top-k and chunk size before switching models. Re-rank and drop marginal chunks; halving context roughly halves input cost. Deduplicate repeated chunks across queries to raise the cache hit rate.