Summarizing contracts, reports, tickets or transcripts: the entire document is input, and the summary is a small fraction of its length.
Input / request
25,000
Output / request
600
Cacheable input
10%
Requests / user / day
4
Cost by model
| Model | Per request | Per user / month | 500 users / month | vs cheapest |
|---|---|---|---|---|
| bestGemini 2.5 Flash-Lite | $0.002515 | $0.3018 | $150.90 | — |
| Ministral 3 (8B) | $0.003503 | $0.4203 | $210.15 | 1.4× |
| Mistral Small 4 | $0.003773 | $0.4527 | $226.35 | 1.5× |
| GPT-5.6 Luna | $0.00527 | $0.6324 | $316.20 | 2.1× |
| GPT-5.4 nano | $0.0053 | $0.636 | $318.00 | 2.1× |
| Codestral | $0.007365 | $0.8838 | $441.90 | 2.9× |
| Gemini 3.5 Flash-Lite | $0.008325 | $0.999 | $499.50 | 3.3× |
| Gemini 2.5 Flash | $0.008325 | $0.999 | $499.50 | 3.3× |
| DeepSeek V4 Flash | $0.0107 | $1.287 | $643.62 | 4.3× |
| Mistral Large 3 | $0.0123 | $1.473 | $736.50 | 4.9× |
| GPT-5.4 mini | $0.0198 | $2.372 | $1.2k | 7.9× |
| Claude Haiku 4.5 | $0.0258 | $3.09 | $1.5k | 10.2× |
| o4-mini | $0.0281 | $3.369 | $1.7k | 11.2× |
| Grok 4.3 | $0.0301 | $3.615 | $1.8k | 12.0× |
| DeepSeek V4 Pro | $0.0322 | $3.862 | $1.9k | 12.8× |
| Muse Spark | $0.0338 | $4.056 | $2k | 13.4× |
| GPT-5.1 | $0.0344 | $4.132 | $2.1k | 13.7× |
| Gemini 2.5 Pro | $0.0344 | $4.132 | $2.1k | 13.7× |
| GLM 5.2 | $0.0345 | $4.139 | $2.1k | 13.7× |
| Gemini 3.6 Flash | $0.0386 | $4.635 | $2.3k | 15.4× |
| Mistral Medium 3.5 | $0.0386 | $4.635 | $2.3k | 15.4× |
| Gemini 3.5 Flash | $0.0395 | $4.743 | $2.4k | 15.7× |
| GPT-5.2 | $0.0482 | $5.786 | $2.9k | 19.2× |
| Grok 4.5 | $0.0494 | $5.922 | $3k | 19.6× |
| Grok 4.6 | $0.0499 | $5.982 | $3k | 19.8× |
| o3 | $0.0511 | $6.126 | $3.1k | 20.3× |
| GPT-4.1 | $0.0511 | $6.126 | $3.1k | 20.3× |
| Claude Sonnet 5 | $0.0515 | $6.18 | $3.1k | 20.5× |
| GPT-5.6 Terra | $0.0527 | $6.324 | $3.2k | 21.0× |
| Gemini 3.1 Pro | $0.0527 | $6.324 | $3.2k | 21.0× |
| GPT-5.4 | $0.0659 | $7.905 | $4k | 26.2× |
| Claude Sonnet 4.5 | $0.0773 | $9.27 | $4.6k | 30.7× |
| Kimi K3 | $0.0773 | $9.27 | $4.6k | 30.7× |
| Claude Opus 5 | $0.1287 | $15.45 | $7.7k | 51.2× |
| Claude Opus 4.5 | $0.1287 | $15.45 | $7.7k | 51.2× |
| GPT-5.6 Sol | $0.1317 | $15.81 | $7.9k | 52.4× |
| GPT-5.5 | $0.1317 | $15.81 | $7.9k | 52.4× |
| Claude Fable 5 | $0.2575 | $30.90 | $15.4k | 102.4× |
Assumes the typical cacheable share (10% of input at cached rates where published). 4 requests/user/day. Adjust everything in the monthly calculator.
How to spend less
Related workflow costs
FAQ
Using a typical profile of 25,000 input and 600 output tokens with 10% cacheable input: from $0.002515 on Gemini 2.5 Flash-Lite up to $0.2575 on Claude Fable 5. See the table for every model.
At 4 requests per user per day and the typical token profile, budget from $0.3018 per active user per month on Gemini 2.5 Flash-Lite. A team of 500 active users lands around $150.90/month at that tier. Model your exact numbers in the monthly cost calculator.
Long-context models pay off here — compare price per 1M tokens at your true document size. Summarize once, store the result; don't re-summarize unchanged documents. For batch backfills, nightly jobs can use cache-friendly request ordering.