Turning unstructured text into JSON: invoices, tickets, leads, moderation labels. Structured input in, compact structured output.
Input / request
2,000
Output / request
250
Cacheable input
20%
Requests / user / day
200
Cost by model
| Model | Per request | Per user / month | 500 users / month | vs cheapest |
|---|---|---|---|---|
| bestGemini 2.5 Flash-Lite | $0.000264 | $1.584 | $792.00 | — |
| Ministral 3 (8B) | $0.000284 | $1.701 | $850.50 | 1.1× |
| Mistral Small 4 | $0.000396 | $2.376 | $1.2k | 1.5× |
| GPT-5.6 Luna | $0.000628 | $3.768 | $1.9k | 2.4× |
| GPT-5.4 nano | $0.000641 | $3.843 | $1.9k | 2.4× |
| Codestral | $0.000717 | $4.302 | $2.2k | 2.7× |
| DeepSeek V4 Flash | $0.00104 | $6.238 | $3.1k | 3.9× |
| Gemini 3.5 Flash-Lite | $0.001117 | $6.702 | $3.4k | 4.2× |
| Gemini 2.5 Flash | $0.001117 | $6.702 | $3.4k | 4.2× |
| Mistral Large 3 | $0.001195 | $7.17 | $3.6k | 4.5× |
| GPT-5.4 mini | $0.002355 | $14.13 | $7.1k | 8.9× |
| Grok 4.3 | $0.002705 | $16.23 | $8.1k | 10.2× |
| Claude Haiku 4.5 | $0.00289 | $17.34 | $8.7k | 10.9× |
| o4-mini | $0.00297 | $17.82 | $8.9k | 11.3× |
| DeepSeek V4 Pro | $0.00312 | $18.718 | $9.4k | 11.8× |
| GLM 5.2 | $0.003396 | $20.376 | $10.2k | 12.9× |
| Muse Spark | $0.003563 | $21.375 | $10.7k | 13.5× |
| Gemini 3.6 Flash | $0.004335 | $26.01 | $13k | 16.4× |
| Mistral Medium 3.5 | $0.004335 | $26.01 | $13k | 16.4× |
| GPT-5.1 | $0.00455 | $27.30 | $13.7k | 17.2× |
| Gemini 2.5 Pro | $0.00455 | $27.30 | $13.7k | 17.2× |
| Gemini 3.5 Flash | $0.00471 | $28.26 | $14.1k | 17.8× |
| Grok 4.5 | $0.00482 | $28.92 | $14.5k | 18.3× |
| Grok 4.6 | $0.0049 | $29.40 | $14.7k | 18.6× |
| o3 | $0.0054 | $32.40 | $16.2k | 20.5× |
| GPT-4.1 | $0.0054 | $32.40 | $16.2k | 20.5× |
| Claude Sonnet 5 | $0.00578 | $34.68 | $17.3k | 21.9× |
| GPT-5.6 Terra | $0.00628 | $37.68 | $18.8k | 23.8× |
| Gemini 3.1 Pro | $0.00628 | $37.68 | $18.8k | 23.8× |
| GPT-5.2 | $0.00637 | $38.22 | $19.1k | 24.1× |
| GPT-5.4 | $0.00785 | $47.10 | $23.5k | 29.7× |
| Claude Sonnet 4.5 | $0.00867 | $52.02 | $26k | 32.8× |
| Kimi K3 | $0.00867 | $52.02 | $26k | 32.8× |
| Claude Opus 5 | $0.0145 | $86.70 | $43.4k | 54.7× |
| Claude Opus 4.5 | $0.0145 | $86.70 | $43.4k | 54.7× |
| GPT-5.6 Sol | $0.0157 | $94.20 | $47.1k | 59.5× |
| GPT-5.5 | $0.0157 | $94.20 | $47.1k | 59.5× |
| Claude Fable 5 | $0.0289 | $173.40 | $86.7k | 109.5× |
Assumes the typical cacheable share (20% of input at cached rates where published). 200 requests/user/day. Adjust everything in the monthly calculator.
How to spend less
Related workflow costs
FAQ
Using a typical profile of 2,000 input and 250 output tokens with 20% cacheable input: from $0.000264 on Gemini 2.5 Flash-Lite up to $0.0289 on Claude Fable 5. See the table for every model.
At 200 requests per user per day and the typical token profile, budget from $1.584 per active user per month on Gemini 2.5 Flash-Lite. A team of 500 active users lands around $792.00/month at that tier. Model your exact numbers in the monthly cost calculator.
Volume makes nano-tier models attractive — extraction rarely needs frontier reasoning. Constrain output with JSON schema modes to avoid retry loops on malformed responses. Cache shared schema instructions and few-shot examples across calls.