A support bot answers customer questions using product documentation. Each turn sends a system prompt plus retrieved docs as input and generates a short helpful reply.
Input / request
3,500
Output / request
350
Cacheable input
70%
Requests / user / day
6
Cost by model
| Model | Per request | Per user / month | 500 users / month | vs cheapest |
|---|---|---|---|---|
| bestMinistral 3 (8B) | $0.000247 | $0.0444 | $22.207 | — |
| Gemini 2.5 Flash-Lite | $0.00027 | $0.0485 | $24.255 | 1.1× |
| Mistral Small 4 | $0.000404 | $0.0728 | $36.383 | 1.6× |
| GPT-5.6 Luna | $0.000679 | $0.1222 | $61.11 | 2.8× |
| GPT-5.4 nano | $0.000697 | $0.1254 | $62.685 | 2.8× |
| Codestral | $0.000703 | $0.1266 | $63.315 | 2.9× |
| DeepSeek V4 Flash | $0.000958 | $0.1725 | $86.247 | 3.9× |
| Mistral Large 3 | $0.001173 | $0.2111 | $105.525 | 4.8× |
| Gemini 3.5 Flash-Lite | $0.001263 | $0.2274 | $113.715 | 5.1× |
| Gemini 2.5 Flash | $0.001263 | $0.2274 | $113.715 | 5.1× |
| GPT-5.4 mini | $0.002546 | $0.4583 | $229.163 | 10.3× |
| Grok 4.3 | $0.002677 | $0.4819 | $240.975 | 10.9× |
| DeepSeek V4 Pro | $0.00288 | $0.5184 | $259.182 | 11.7× |
| Claude Haiku 4.5 | $0.003045 | $0.5481 | $274.05 | 12.3× |
| GLM 5.2 | $0.003353 | $0.6035 | $301.77 | 13.6× |
| o4-mini | $0.003369 | $0.6064 | $303.188 | 13.7× |
| Gemini 3.6 Flash | $0.004568 | $0.8222 | $411.075 | 18.5× |
| Mistral Medium 3.5 | $0.004568 | $0.8222 | $411.075 | 18.5× |
| Grok 4.5 | $0.004935 | $0.8883 | $444.15 | 20.0× |
| Gemini 3.5 Flash | $0.005093 | $0.9167 | $458.325 | 20.6× |
| GPT-5.1 | $0.005119 | $0.9214 | $460.688 | 20.7× |
| Gemini 2.5 Pro | $0.005119 | $0.9214 | $460.688 | 20.7× |
| Grok 4.6 | $0.005425 | $0.9765 | $488.25 | 22.0× |
| Muse Spark | $0.005863 | $1.055 | $527.625 | 23.8× |
| Claude Sonnet 5 | $0.00609 | $1.096 | $548.10 | 24.7× |
| o3 | $0.006125 | $1.103 | $551.25 | 24.8× |
| GPT-4.1 | $0.006125 | $1.103 | $551.25 | 24.8× |
| GPT-5.6 Terra | $0.00679 | $1.222 | $611.10 | 27.5× |
| Gemini 3.1 Pro | $0.00679 | $1.222 | $611.10 | 27.5× |
| GPT-5.2 | $0.007166 | $1.29 | $644.963 | 29.0× |
| GPT-5.4 | $0.008488 | $1.528 | $763.875 | 34.4× |
| Claude Sonnet 4.5 | $0.009135 | $1.644 | $822.15 | 37.0× |
| Kimi K3 | $0.009135 | $1.644 | $822.15 | 37.0× |
| Claude Opus 5 | $0.0152 | $2.741 | $1.4k | 61.7× |
| Claude Opus 4.5 | $0.0152 | $2.741 | $1.4k | 61.7× |
| GPT-5.6 Sol | $0.017 | $3.056 | $1.5k | 68.8× |
| GPT-5.5 | $0.017 | $3.056 | $1.5k | 68.8× |
| Claude Fable 5 | $0.0304 | $5.481 | $2.7k | 123.4× |
Assumes the typical cacheable share (70% of input at cached rates where published). 6 requests/user/day. Adjust everything in the monthly calculator.
How to spend less
Related workflow costs
FAQ
Using a typical profile of 3,500 input and 350 output tokens with 70% cacheable input: from $0.000247 on Ministral 3 (8B) up to $0.0304 on Claude Fable 5. See the table for every model.
At 6 requests per user per day and the typical token profile, budget from $0.0444 per active user per month on Ministral 3 (8B). A team of 500 active users lands around $22.207/month at that tier. Model your exact numbers in the monthly cost calculator.
Cache the system prompt and static documentation chunks — cached input is often 4–10× cheaper. Route simple FAQ turns to a nano-tier model and escalate only complex tickets. Cap max_output per reply; support answers rarely need more than a few hundred tokens.