An IDE or code-review assistant sends the open file plus relevant repository context and expects long, precise code output — output tokens dominate the bill.
Input / request
12,000
Output / request
2,000
Cacheable input
60%
Requests / user / day
40
Cost by model
| Model | Per request | Per user / month | 500 users / month | vs cheapest |
|---|---|---|---|---|
| bestMinistral 3 (8B) | $0.001128 | $1.354 | $676.80 | — |
| Gemini 2.5 Flash-Lite | $0.001352 | $1.622 | $811.20 | 1.2× |
| Mistral Small 4 | $0.002028 | $2.434 | $1.2k | 1.8× |
| Codestral | $0.003456 | $4.147 | $2.1k | 3.1× |
| GPT-5.6 Luna | $0.003504 | $4.205 | $2.1k | 3.1× |
| GPT-5.4 nano | $0.003604 | $4.325 | $2.2k | 3.2× |
| DeepSeek V4 Flash | $0.004853 | $5.823 | $2.9k | 4.3× |
| Mistral Large 3 | $0.00576 | $6.912 | $3.5k | 5.1× |
| Gemini 3.5 Flash-Lite | $0.006656 | $7.987 | $4k | 5.9× |
| Gemini 2.5 Flash | $0.006656 | $7.987 | $4k | 5.9× |
| Grok 4.3 | $0.0124 | $14.928 | $7.5k | 11.0× |
| GPT-5.4 mini | $0.0131 | $15.768 | $7.9k | 11.6× |
| DeepSeek V4 Pro | $0.0146 | $17.487 | $8.7k | 12.9× |
| Claude Haiku 4.5 | $0.0155 | $18.624 | $9.3k | 13.8× |
| o4-mini | $0.0161 | $19.272 | $9.6k | 14.2× |
| GLM 5.2 | $0.0165 | $19.834 | $9.9k | 14.7× |
| Gemini 3.6 Flash | $0.0233 | $27.936 | $14k | 20.6× |
| Mistral Medium 3.5 | $0.0233 | $27.936 | $14k | 20.6× |
| Muse Spark | $0.0235 | $28.20 | $14.1k | 20.8× |
| Grok 4.5 | $0.0238 | $28.512 | $14.3k | 21.1× |
| Grok 4.6 | $0.0252 | $30.24 | $15.1k | 22.3× |
| Gemini 3.5 Flash | $0.0263 | $31.536 | $15.8k | 23.3× |
| GPT-5.1 | $0.0269 | $32.28 | $16.1k | 23.8× |
| Gemini 2.5 Pro | $0.0269 | $32.28 | $16.1k | 23.8× |
| o3 | $0.0292 | $35.04 | $17.5k | 25.9× |
| GPT-4.1 | $0.0292 | $35.04 | $17.5k | 25.9× |
| Claude Sonnet 5 | $0.031 | $37.248 | $18.6k | 27.5× |
| GPT-5.6 Terra | $0.035 | $42.048 | $21k | 31.1× |
| Gemini 3.1 Pro | $0.035 | $42.048 | $21k | 31.1× |
| GPT-5.2 | $0.0377 | $45.192 | $22.6k | 33.4× |
| GPT-5.4 | $0.0438 | $52.56 | $26.3k | 38.8× |
| Claude Sonnet 4.5 | $0.0466 | $55.872 | $27.9k | 41.3× |
| Kimi K3 | $0.0466 | $55.872 | $27.9k | 41.3× |
| Claude Opus 5 | $0.0776 | $93.12 | $46.6k | 68.8× |
| Claude Opus 4.5 | $0.0776 | $93.12 | $46.6k | 68.8× |
| GPT-5.6 Sol | $0.0876 | $105.12 | $52.6k | 77.7× |
| GPT-5.5 | $0.0876 | $105.12 | $52.6k | 77.7× |
| Claude Fable 5 | $0.1552 | $186.24 | $93.1k | 137.6× |
Assumes the typical cacheable share (60% of input at cached rates where published). 40 requests/user/day. Adjust everything in the monthly calculator.
How to spend less
Related workflow costs
FAQ
Using a typical profile of 12,000 input and 2,000 output tokens with 60% cacheable input: from $0.001128 on Ministral 3 (8B) up to $0.1552 on Claude Fable 5. See the table for every model.
At 40 requests per user per day and the typical token profile, budget from $1.354 per active user per month on Ministral 3 (8B). A team of 500 active users lands around $676.80/month at that tier. Model your exact numbers in the monthly cost calculator.
Output is the expensive side — prefer models with cheap output for autocomplete-style calls. Cache repository context between keystrokes; diffs change far less than the full file. Measure acceptance rate: paying for output users delete is pure waste.