An agent plans, calls tools and iterates: every step re-sends the growing conversation, so input tokens compound across steps before the final answer.
Input / request
40,000
Output / request
1,500
Cacheable input
65%
Requests / user / day
8
Cost by model
| Model | Per request | Per user / month | 500 users / month | vs cheapest |
|---|---|---|---|---|
| bestGLM-4.7-Flash (Z.ai) | Free | Free | Free | — |
| GLM-4.5-Flash (Z.ai) | Free | Free | Free | Not comparable |
| GLM-4.6V-Flash (Z.ai) | Free | Free | Free | Not comparable |
| OpenRouter Free Tier (Llama 3.3) | Free | Free | Free | Not comparable |
| Ox Alpha (Z.ai GLM preview) | Free | Free | Free | Not comparable |
| Text Embedding 3 (Small) | $0.0008 | $0.192 | $96.00 | Not comparable |
| Text Embedding 004 | $0.0008 | $0.192 | $96.00 | Not comparable |
| Amazon Nova Micro | $0.000928 | $0.2226 | $111.30 | Not comparable |
| Qwen 2.5 Turbo | $0.00113 | $0.2712 | $135.60 | Not comparable |
| GLM-4.6V-FlashX (Z.ai) | $0.001264 | $0.3034 | $151.68 | Not comparable |
| Llama 3.2 3B Instruct | $0.001275 | $0.306 | $153.00 | Not comparable |
| Llama 3.2 1B Instruct (OpenRouter) | $0.001382 | $0.3316 | $165.78 | Not comparable |
| Amazon Nova Lite | $0.00159 | $0.3816 | $190.80 | Not comparable |
| GLM-5.3-Flash (Z.ai) | $0.001815 | $0.4356 | $217.80 | Not comparable |
| GLM-4.7-FlashX (Z.ai) | $0.00184 | $0.4416 | $220.80 | Not comparable |
| Gemini 2.0 Flash-Lite | $0.001987 | $0.477 | $238.50 | Not comparable |
| Gemini 1.5 Flash | $0.001987 | $0.477 | $238.50 | Not comparable |
| Llama 3.1 8B Instruct | $0.00212 | $0.5088 | $254.40 | Not comparable |
| Groq LPU — Llama 3.1 8B Instant | $0.00212 | $0.5088 | $254.40 | Not comparable |
| Cohere Command R7B | $0.002225 | $0.534 | $267.00 | Not comparable |
| Together AI — GPT-OSS-20B | $0.0023 | $0.552 | $276.00 | Not comparable |
| Gemini 2.5 Flash-Lite | $0.00265 | $0.636 | $318.00 | Not comparable |
| Gemini 2.0 Flash | $0.00265 | $0.636 | $318.00 | Not comparable |
| Ministral 3 8B | $0.002715 | $0.6516 | $325.80 | Not comparable |
| DeepSeek V3 (Chat) | $0.002744 | $0.6586 | $329.28 | Not comparable |
| DeepSeek Coder V2.5 | $0.002744 | $0.6586 | $329.28 | Not comparable |
| Fireworks AI — Llama 4 Scout | $0.003 | $0.72 | $360.00 | Not comparable |
| Mistral Small 3.2 24B (OpenRouter) | $0.0033 | $0.792 | $396.00 | Not comparable |
| Mistral Small 4 | $0.00339 | $0.8136 | $406.80 | Not comparable |
| Hy-MT2-30B-A3B (OpenRouter) | $0.003403 | $0.8166 | $408.30 | Not comparable |
| Groq LPU — Llama 4 Scout | $0.00348 | $0.8352 | $417.60 | Not comparable |
| DeepInfra — Qwen 2.5 Coder 32B | $0.00356 | $0.8544 | $427.20 | Not comparable |
| Ministral 3 14B | $0.00362 | $0.8688 | $434.40 | Not comparable |
| GLM-4-32B-0414-128K (Z.ai) | $0.00415 | $0.996 | $498.00 | Not comparable |
| Qwen 2.5 Coder 32B | $0.00422 | $1.013 | $506.40 | Not comparable |
| Muse Spark 1.2 Contributor (OpenRouter) | $0.0043 | $1.032 | $516.00 | Not comparable |
| Microsoft Phi-4 (14B) | $0.00445 | $1.068 | $534.00 | Not comparable |
| DeepInfra — Llama 3.3 70B | $0.00445 | $1.068 | $534.00 | Not comparable |
| Grok 4.1 Fast | $0.00459 | $1.102 | $550.80 | Not comparable |
| Qwen3 Coder Next (OpenRouter) | $0.0047 | $1.128 | $564.00 | Not comparable |
| GPT-4o mini | $0.00495 | $1.188 | $594.00 | Not comparable |
| GPT-5.6 Luna | $0.00512 | $1.229 | $614.40 | Not comparable |
| GPT-5.6 Luna Pro (OpenRouter) | $0.00512 | $1.229 | $614.40 | Not comparable |
| GPT-5.4 nano | $0.005195 | $1.247 | $623.40 | Not comparable |
| Text Embedding 3 (Large) | $0.0052 | $1.248 | $624.00 | Not comparable |
| Qwen3 Coder Flash (OpenRouter) | $0.005207 | $1.25 | $624.78 | Not comparable |
| GLM-4.5-Air (Z.ai) | $0.00523 | $1.255 | $627.60 | Not comparable |
| MiniMax-01 (4M Context) | $0.00549 | $1.318 | $658.80 | Not comparable |
| Yi-Lightning (01.AI) | $0.00581 | $1.394 | $697.20 | Not comparable |
| Claude 3 Haiku | $0.006025 | $1.446 | $723.00 | Not comparable |
| Groq LPU — Llama 4 Maverick | $0.0063 | $1.512 | $756.00 | Not comparable |
| Codestral 2501 | $0.00633 | $1.519 | $759.60 | Not comparable |
| Codestral 2508 (OpenRouter) | $0.00633 | $1.519 | $759.60 | Not comparable |
| Gemini 3.1 Flash Lite (OpenRouter) | $0.0064 | $1.536 | $768.00 | Not comparable |
| Llama 3.2 11B Vision | $0.00664 | $1.594 | $796.80 | Not comparable |
| Llama 4 Scout (109B MoE) | $0.006675 | $1.602 | $801.00 | Not comparable |
| GLM-4.6V (Z.ai) | $0.00685 | $1.644 | $822.00 | Not comparable |
| Qwen 2.5 72B Instruct | $0.00686 | $1.646 | $823.20 | Not comparable |
| Cohere Command R | $0.0069 | $1.656 | $828.00 | Not comparable |
| Microsoft Phi-3.5 MoE | $0.0069 | $1.656 | $828.00 | Not comparable |
| Qwen3.8-Flash (QwenCloud preview) | $0.007105 | $1.705 | $852.60 | Not comparable |
| MiniMax M2.7 (OpenRouter) | $0.00756 | $1.814 | $907.20 | Not comparable |
| MiniMax M3 (OpenRouter) | $0.00756 | $1.814 | $907.20 | Not comparable |
| DeepSeek V3.2 (OpenRouter) | $0.00759 | $1.822 | $910.80 | Not comparable |
| Qwen 3 32B Instruct | $0.00784 | $1.882 | $940.80 | Not comparable |
| Llama 3.3 70B Instruct | $0.008085 | $1.94 | $970.20 | Not comparable |
| Gemma 2 9B Instruct | $0.0083 | $1.992 | $996.00 | Not comparable |
| QwQ 32B (Reasoner) | $0.00844 | $2.026 | $1,012.80 | Not comparable |
| DeepSeek V4 Flash | $0.008504 | $2.041 | $1,020.48 | Not comparable |
| DeepSeek V4 Flash Vision Exp | $0.008504 | $2.041 | $1,020.48 | Not comparable |
| AI21 Jamba 1.5 Mini | $0.0086 | $2.064 | $1,032.00 | Not comparable |
| Gemini 3.5 Flash-Lite | $0.00873 | $2.095 | $1,047.60 | Not comparable |
| Gemini 2.5 Flash | $0.0099 | $2.376 | $1,188.00 | Not comparable |
| Mistral Large 3 | $0.0106 | $2.532 | $1,266.00 | Not comparable |
| Devstral 2 (2512) (OpenRouter) | $0.0106 | $2.545 | $1,272.48 | Not comparable |
| Gemma 4 31B Instruct | $0.0111 | $2.67 | $1,335.00 | Not comparable |
| Groq LPU — QwQ 32B Reasoner | $0.0122 | $2.924 | $1,462.20 | Not comparable |
| GLM-4.5V (Z.ai) | $0.014 | $3.35 | $1,675.20 | Not comparable |
| GLM-4.6 (Z.ai) | $0.0146 | $3.494 | $1,747.20 | Not comparable |
| GLM-4.5 (Z.ai) | $0.0146 | $3.494 | $1,747.20 | Not comparable |
| GLM 4.7 (Zhipu) | $0.0146 | $3.494 | $1,747.20 | Not comparable |
| DeepSeek R1 (Reasoner) | $0.0146 | $3.51 | $1,755.00 | Not comparable |
| NVIDIA Llama 3.1 Nemotron 70B | $0.0151 | $3.612 | $1,806.00 | Not comparable |
| Cerebras — GPT OSS 120B | $0.0151 | $3.63 | $1,815.00 | Not comparable |
| Together AI — Llama 4 Maverick | $0.0153 | $3.684 | $1,842.00 | Not comparable |
| Llama 3.1 70B Instruct (OpenRouter) | $0.0166 | $3.984 | $1,992.00 | Not comparable |
| Cerebras — Qwen 3 32B | $0.0172 | $4.128 | $2,064.00 | Not comparable |
| Qwen3.5 397B A17B (OpenRouter) | $0.0191 | $4.586 | $2,293.20 | Not comparable |
| GPT-5.4 mini | $0.0192 | $4.608 | $2,304.00 | Not comparable |
| Claude 3.5 Haiku | $0.0193 | $4.627 | $2,313.60 | Not comparable |
| Kimi K2.7 Code (OpenRouter) | $0.0194 | $4.661 | $2,330.40 | Not comparable |
| Llama 4 Maverick (400B MoE) | $0.0199 | $4.77 | $2,385.00 | Not comparable |
| Amazon Nova Pro | $0.0212 | $5.088 | $2,544.00 | Not comparable |
| Qwen3.8 27B (OpenRouter) | $0.0213 | $5.112 | $2,556.00 | Not comparable |
| Qwen3 VL 235B Thinking (OpenRouter) | $0.022 | $5.28 | $2,640.00 | Not comparable |
| Grok Build 0.1 | $0.0222 | $5.328 | $2,664.00 | Not comparable |
| GPT-3.5 Turbo | $0.0223 | $5.34 | $2,670.00 | Not comparable |
| Together AI — DeepSeek V4 | $0.0223 | $5.34 | $2,670.00 | Not comparable |
| GLM-5 (Z.ai) | $0.024 | $5.76 | $2,880.00 | Not comparable |
| Claude Haiku 4.5 | $0.0241 | $5.784 | $2,892.00 | Not comparable |
| Groq LPU — Llama 3.3 70B | $0.0248 | $5.948 | $2,974.20 | Not comparable |
| Fireworks AI — DeepSeek R1 | $0.0253 | $6.068 | $3,034.20 | Not comparable |
| DeepSeek V4 Pro | $0.0256 | $6.135 | $3,067.68 | Not comparable |
| Grok 4.3 | $0.0265 | $6.348 | $3,174.00 | Not comparable |
| Databricks DBRX Instruct | $0.0267 | $6.408 | $3,204.00 | Not comparable |
| GLM-4.5-AirX (Z.ai) | $0.0279 | $6.689 | $3,344.40 | Not comparable |
| GLM-5-Turbo (Z.ai) | $0.029 | $6.97 | $3,484.80 | Not comparable |
| GLM-5V-Turbo (Z.ai) | $0.029 | $6.97 | $3,484.80 | Not comparable |
| o4-mini | $0.0292 | $6.996 | $3,498.00 | Not comparable |
| GLM-5.3 (Z.ai) | $0.033 | $7.91 | $3,955.20 | Not comparable |
| GLM-5.2 (Z.ai) | $0.033 | $7.91 | $3,955.20 | Not comparable |
| GLM-5.1 (Z.ai) | $0.033 | $7.91 | $3,955.20 | Not comparable |
| Gemini 1.5 Pro | $0.0331 | $7.95 | $3,975.00 | Not comparable |
| Gemini 3.7 Flash | $0.0339 | $8.136 | $4,068.00 | Not comparable |
| Kimi K2.6 | $0.0339 | $8.136 | $4,068.00 | Not comparable |
| Gemini 3.6 Flash | $0.0362 | $8.676 | $4,338.00 | Not comparable |
| Mistral Medium 3.5 | $0.0362 | $8.676 | $4,338.00 | Not comparable |
| Qwen 2.5 Max | $0.0362 | $8.678 | $4,339.20 | Not comparable |
| o3-mini | $0.0363 | $8.712 | $4,356.00 | Not comparable |
| o1-mini | $0.0363 | $8.712 | $4,356.00 | Not comparable |
| Llama 3.2 90B Vision | $0.0374 | $8.964 | $4,482.00 | Not comparable |
| Gemini 3.5 Flash (OpenRouter) | $0.0384 | $9.216 | $4,608.00 | Not comparable |
| Gemini 2.5 Pro | $0.0406 | $9.75 | $4,875.00 | Not comparable |
| Perplexity Sonar | $0.0415 | $9.96 | $4,980.00 | Not comparable |
| Cerebras — Gemma 4 31B | $0.0418 | $10.04 | $5,020.20 | Not comparable |
| Pixtral Large | $0.0422 | $10.128 | $5,064.00 | Not comparable |
| Mistral Large 2 (2411) | $0.0422 | $10.128 | $5,064.00 | Not comparable |
| Qwen 3.8 Max (2.4T MoE) | $0.0422 | $10.128 | $5,064.00 | Not comparable |
| OpenRouter Auto-Best Router | $0.0445 | $10.68 | $5,340.00 | Not comparable |
| Grok 4.5 | $0.0448 | $10.752 | $5,376.00 | Not comparable |
| Amazon Nova Premier | $0.0452 | $10.848 | $5,424.00 | Not comparable |
| Claude Sonnet 5 | $0.0482 | $11.568 | $5,784.00 | Not comparable |
| GPT-5.6 Sol Pro (OpenRouter) | $0.0482 | $11.568 | $5,784.00 | Not comparable |
| Grok 4.6 Flagship | $0.05 | $12.00 | $6,000.00 | Not comparable |
| GPT-5.3 Codex | $0.0501 | $12.012 | $6,006.00 | Not comparable |
| GPT-5.6 Terra | $0.0512 | $12.288 | $6,144.00 | Not comparable |
| Gemini 3.1 Pro | $0.0512 | $12.288 | $6,144.00 | Not comparable |
| GPT-5.6 Terra Pro (OpenRouter) | $0.0512 | $12.288 | $6,144.00 | Not comparable |
| GLM-4.5-X (Z.ai) | $0.0559 | $13.404 | $6,702.00 | Not comparable |
| Cohere Command A+ | $0.0565 | $13.56 | $6,780.00 | Not comparable |
| GPT-5.4 Workhorse | $0.064 | $15.36 | $7,680.00 | Not comparable |
| NVIDIA Nemotron-4 340B | $0.0645 | $15.48 | $7,740.00 | Not comparable |
| Claude 3.7 Sonnet (Hybrid Reasoning) | $0.0723 | $17.352 | $8,676.00 | Not comparable |
| Claude 3.5 Sonnet | $0.0723 | $17.352 | $8,676.00 | Not comparable |
| Kimi K3 (Moonshot) | $0.0723 | $17.352 | $8,676.00 | Not comparable |
| Claude Sonnet 4.6 (OpenRouter) | $0.0723 | $17.352 | $8,676.00 | Not comparable |
| Llama 3.1 405B Instruct | $0.0753 | $18.06 | $9,030.00 | Not comparable |
| GPT-4o (Omni) | $0.0825 | $19.80 | $9,900.00 | Not comparable |
| Perplexity Sonar Reasoning Pro | $0.092 | $22.08 | $11,040.00 | Not comparable |
| Perplexity Sonar Deep Research | $0.092 | $22.08 | $11,040.00 | Not comparable |
| AI21 Jamba 1.5 Large | $0.092 | $22.08 | $11,040.00 | Not comparable |
| Grok 2 | $0.095 | $22.80 | $11,400.00 | Not comparable |
| Grok 2 Vision | $0.095 | $22.80 | $11,400.00 | Not comparable |
| Cohere Command R+ | $0.115 | $27.60 | $13,800.00 | Not comparable |
| Claude Opus 5 | $0.1205 | $28.92 | $14,460.00 | Not comparable |
| Claude Opus 4.8 (OpenRouter) | $0.1205 | $28.92 | $14,460.00 | Not comparable |
| GPT-5.6 Sol | $0.128 | $30.72 | $15,360.00 | Not comparable |
| GPT-5.5 Standard | $0.128 | $30.72 | $15,360.00 | Not comparable |
| Perplexity Sonar Pro | $0.1425 | $34.20 | $17,100.00 | Not comparable |
| o3 (Reasoning Frontier) | $0.226 | $54.24 | $27,120.00 | Not comparable |
| Claude Fable 5 | $0.241 | $57.84 | $28,920.00 | Not comparable |
| GPT-4 Turbo | $0.315 | $75.60 | $37,800.00 | Not comparable |
| GPT-5.6 Cyber | $0.32 | $76.80 | $38,400.00 | Not comparable |
| Claude 3 Opus | $0.3615 | $86.76 | $43,380.00 | Not comparable |
| o3-pro (Frontier Reasoning) | $0.452 | $108.48 | $54,240.00 | Not comparable |
| o1 (Reasoning) | $0.495 | $118.80 | $59,400.00 | Not comparable |
| GPT-5.5 Pro | $0.768 | $184.32 | $92,160.00 | Not comparable |
Assumes the typical cacheable share (65% of input at cached rates where published). 8 requests/user/day. Adjust everything in the monthly calculator.
Best value picks for Agentic workflow
| Model | Fit score | Cost / 1K requests | Value (fit ÷ cost) |
|---|---|---|---|
| best valueGLM-5.3-Flash (Z.ai) | 60.7 | $1.815 | 33.4 |
| GPT-5.6 Luna | 52.9 | $5.12 | 10.3 |
| GPT-5.4 nano | 37.0 | $5.195 | 7.1 |
| DeepSeek V3.2 (OpenRouter) | 44.2 | $7.59 | 5.8 |
| MiniMax M3 (OpenRouter) | 42.5 | $7.56 | 5.6 |
Fit = weighted blend of third-party composite indices (intelligence, coding, agentic) for this use case; value = fit ÷ cost per 1,000 requests. Models without a verified composite are excluded. Methodology · Benchmark hub.
How to spend less
Related workflow costs
FAQ
Using a typical profile of 40,000 input and 1,500 output tokens with 65% cacheable input: from Free on GLM-4.7-Flash (Z.ai) up to $0.768 on GPT-5.5 Pro. See the table for every model.
At 8 requests per user per day and the typical token profile, budget from Free per active user per month on GLM-4.7-Flash (Z.ai). A team of 500 active users lands around Free/month at that tier. Model your exact numbers in the monthly cost calculator.
Prompt caching is the single biggest lever — each step re-reads prior context. Summarize or prune tool outputs before appending them to the transcript. Cap the step budget; runaway loops are the #1 surprise on agent invoices.