Summarizing contracts, reports, tickets or transcripts: the entire document is input, and the summary is a small fraction of its length.
Input / request
25,000
Output / request
600
Cacheable input
10%
Requests / user / day
4
Cost by model
| Model | Per request | Per user / month | 500 users / month | vs cheapest |
|---|---|---|---|---|
| bestGLM-4.7-Flash (Z.ai) | Free | Free | Free | — |
| GLM-4.5-Flash (Z.ai) | Free | Free | Free | Not comparable |
| GLM-4.6V-Flash (Z.ai) | Free | Free | Free | Not comparable |
| OpenRouter Free Tier (Llama 3.3) | Free | Free | Free | Not comparable |
| Ox Alpha (Z.ai GLM preview) | Free | Free | Free | Not comparable |
| Text Embedding 3 (Small) | $0.0005 | $0.06 | $30.00 | Not comparable |
| Text Embedding 004 | $0.0005 | $0.06 | $30.00 | Not comparable |
| Llama 3.2 3B Instruct | $0.00078 | $0.0936 | $46.80 | Not comparable |
| Llama 3.2 1B Instruct (OpenRouter) | $0.000796 | $0.0955 | $47.736 | Not comparable |
| Amazon Nova Micro | $0.000893 | $0.1072 | $53.603 | Not comparable |
| GLM-4.6V-FlashX (Z.ai) | $0.00115 | $0.138 | $69.00 | Not comparable |
| Qwen 2.5 Turbo | $0.001257 | $0.1509 | $75.45 | Not comparable |
| Llama 3.1 8B Instruct | $0.001298 | $0.1558 | $77.88 | Not comparable |
| Groq LPU — Llama 3.1 8B Instant | $0.001298 | $0.1558 | $77.88 | Not comparable |
| Cohere Command R7B | $0.00134 | $0.1608 | $80.40 | Not comparable |
| Together AI — GPT-OSS-20B | $0.00137 | $0.1644 | $82.20 | Not comparable |
| Amazon Nova Lite | $0.001532 | $0.1838 | $91.89 | Not comparable |
| GLM-4.7-FlashX (Z.ai) | $0.00184 | $0.2208 | $110.40 | Not comparable |
| GLM-5.3-Flash (Z.ai) | $0.001875 | $0.225 | $112.50 | Not comparable |
| Gemini 2.0 Flash-Lite | $0.001914 | $0.2297 | $114.863 | Not comparable |
| Gemini 1.5 Flash | $0.001914 | $0.2297 | $114.863 | Not comparable |
| Mistral Small 3.2 24B (OpenRouter) | $0.001995 | $0.2394 | $119.70 | Not comparable |
| Hy-MT2-30B-A3B (OpenRouter) | $0.002027 | $0.2432 | $121.62 | Not comparable |
| DeepInfra — Qwen 2.5 Coder 32B | $0.002144 | $0.2573 | $128.64 | Not comparable |
| Gemini 2.5 Flash-Lite | $0.002553 | $0.3063 | $153.15 | Not comparable |
| Gemini 2.0 Flash | $0.002553 | $0.3063 | $153.15 | Not comparable |
| GLM-4-32B-0414-128K (Z.ai) | $0.00256 | $0.3072 | $153.60 | Not comparable |
| Muse Spark 1.2 Contributor (OpenRouter) | $0.00262 | $0.3144 | $157.20 | Not comparable |
| Microsoft Phi-4 (14B) | $0.00268 | $0.3216 | $160.80 | Not comparable |
| DeepInfra — Llama 3.3 70B | $0.00268 | $0.3216 | $160.80 | Not comparable |
| Groq LPU — Llama 4 Scout | $0.002817 | $0.338 | $168.99 | Not comparable |
| Fireworks AI — Llama 4 Scout | $0.002991 | $0.3589 | $179.46 | Not comparable |
| Text Embedding 3 (Large) | $0.00325 | $0.39 | $195.00 | Not comparable |
| DeepSeek V3 (Chat) | $0.003353 | $0.4024 | $201.18 | Not comparable |
| DeepSeek Coder V2.5 | $0.003353 | $0.4024 | $201.18 | Not comparable |
| Qwen3 Coder Next (OpenRouter) | $0.003355 | $0.4026 | $201.30 | Not comparable |
| Ministral 3 8B | $0.003503 | $0.4203 | $210.15 | Not comparable |
| Yi-Lightning (01.AI) | $0.003584 | $0.4301 | $215.04 | Not comparable |
| Mistral Small 4 | $0.003773 | $0.4527 | $226.35 | Not comparable |
| GPT-4o mini | $0.003923 | $0.4707 | $235.35 | Not comparable |
| Llama 4 Scout (109B MoE) | $0.00402 | $0.4824 | $241.20 | Not comparable |
| Llama 3.2 11B Vision | $0.004096 | $0.4915 | $245.76 | Not comparable |
| Cohere Command R | $0.00411 | $0.4932 | $246.60 | Not comparable |
| Microsoft Phi-3.5 MoE | $0.00411 | $0.4932 | $246.60 | Not comparable |
| Qwen3.8-Flash (QwenCloud preview) | $0.004282 | $0.5138 | $256.92 | Not comparable |
| Ministral 3 14B | $0.00467 | $0.5604 | $280.20 | Not comparable |
| Llama 3.3 70B Instruct | $0.004854 | $0.5825 | $291.24 | Not comparable |
| Grok 4.1 Fast | $0.0049 | $0.588 | $294.00 | Not comparable |
| Qwen 2.5 Coder 32B | $0.00491 | $0.5892 | $294.60 | Not comparable |
| Qwen3 Coder Flash (OpenRouter) | $0.00507 | $0.6084 | $304.20 | Not comparable |
| Groq LPU — Llama 4 Maverick | $0.00511 | $0.6132 | $306.60 | Not comparable |
| Gemma 2 9B Instruct | $0.00512 | $0.6144 | $307.20 | Not comparable |
| GLM-4.5-Air (Z.ai) | $0.005235 | $0.6282 | $314.10 | Not comparable |
| AI21 Jamba 1.5 Mini | $0.00524 | $0.6288 | $314.40 | Not comparable |
| MiniMax-01 (4M Context) | $0.00526 | $0.6312 | $315.60 | Not comparable |
| GPT-5.6 Luna | $0.00527 | $0.6324 | $316.20 | Not comparable |
| GPT-5.6 Luna Pro (OpenRouter) | $0.00527 | $0.6324 | $316.20 | Not comparable |
| GPT-5.4 nano | $0.0053 | $0.636 | $318.00 | Not comparable |
| DeepSeek V3.2 (OpenRouter) | $0.006403 | $0.7684 | $384.18 | Not comparable |
| Claude 3 Haiku | $0.006438 | $0.7725 | $386.25 | Not comparable |
| Gemini 3.1 Flash Lite (OpenRouter) | $0.006588 | $0.7905 | $395.25 | Not comparable |
| Gemma 4 31B Instruct | $0.0067 | $0.804 | $402.00 | Not comparable |
| Codestral 2501 | $0.007365 | $0.8838 | $441.90 | Not comparable |
| Codestral 2508 (OpenRouter) | $0.007365 | $0.8838 | $441.90 | Not comparable |
| GLM-4.6V (Z.ai) | $0.007415 | $0.8898 | $444.90 | Not comparable |
| Groq LPU — QwQ 32B Reasoner | $0.007484 | $0.8981 | $449.04 | Not comparable |
| MiniMax M2.7 (OpenRouter) | $0.00762 | $0.9144 | $457.20 | Not comparable |
| MiniMax M3 (OpenRouter) | $0.00762 | $0.9144 | $457.20 | Not comparable |
| Gemini 3.5 Flash-Lite | $0.008325 | $0.999 | $499.50 | Not comparable |
| Qwen 2.5 72B Instruct | $0.008383 | $1.006 | $502.95 | Not comparable |
| Gemini 2.5 Flash | $0.008438 | $1.013 | $506.25 | Not comparable |
| NVIDIA Llama 3.1 Nemotron 70B | $0.00917 | $1.10 | $550.20 | Not comparable |
| Cerebras — GPT OSS 120B | $0.0092 | $1.104 | $552.00 | Not comparable |
| Together AI — Llama 4 Maverick | $0.00929 | $1.115 | $557.40 | Not comparable |
| Qwen 3 32B Instruct | $0.00958 | $1.15 | $574.80 | Not comparable |
| QwQ 32B (Reasoner) | $0.00982 | $1.178 | $589.20 | Not comparable |
| Llama 3.1 70B Instruct (OpenRouter) | $0.0102 | $1.229 | $614.40 | Not comparable |
| Cerebras — Qwen 3 32B | $0.0105 | $1.258 | $628.80 | Not comparable |
| DeepSeek V4 Flash | $0.0107 | $1.287 | $643.62 | Not comparable |
| DeepSeek V4 Flash Vision Exp | $0.0107 | $1.287 | $643.62 | Not comparable |
| Qwen3.5 397B A17B (OpenRouter) | $0.0112 | $1.338 | $669.24 | Not comparable |
| Devstral 2 (2512) (OpenRouter) | $0.0113 | $1.36 | $679.80 | Not comparable |
| Llama 4 Maverick (400B MoE) | $0.012 | $1.44 | $720.00 | Not comparable |
| Mistral Large 3 | $0.0123 | $1.473 | $736.50 | Not comparable |
| Qwen3.8 27B (OpenRouter) | $0.0123 | $1.476 | $738.00 | Not comparable |
| Qwen3 VL 235B Thinking (OpenRouter) | $0.0124 | $1.488 | $744.00 | Not comparable |
| GPT-3.5 Turbo | $0.0134 | $1.608 | $804.00 | Not comparable |
| Together AI — DeepSeek V4 | $0.0134 | $1.608 | $804.00 | Not comparable |
| DeepSeek R1 (Reasoner) | $0.014 | $1.685 | $842.34 | Not comparable |
| GLM-4.5V (Z.ai) | $0.0149 | $1.783 | $891.30 | Not comparable |
| Fireworks AI — DeepSeek R1 | $0.0151 | $1.808 | $903.84 | Not comparable |
| GLM-4.6 (Z.ai) | $0.0151 | $1.811 | $905.70 | Not comparable |
| GLM-4.5 (Z.ai) | $0.0151 | $1.811 | $905.70 | Not comparable |
| GLM 4.7 (Zhipu) | $0.0151 | $1.811 | $905.70 | Not comparable |
| Groq LPU — Llama 3.3 70B | $0.0152 | $1.827 | $913.44 | Not comparable |
| Databricks DBRX Instruct | $0.0161 | $1.93 | $964.80 | Not comparable |
| Kimi K2.7 Code (OpenRouter) | $0.0176 | $2.111 | $1,055.40 | Not comparable |
| GPT-5.4 mini | $0.0198 | $2.372 | $1,185.75 | Not comparable |
| Amazon Nova Pro | $0.0204 | $2.45 | $1,225.20 | Not comparable |
| Claude 3.5 Haiku | $0.0206 | $2.472 | $1,236.00 | Not comparable |
| Llama 3.2 90B Vision | $0.023 | $2.765 | $1,382.40 | Not comparable |
| Grok Build 0.1 | $0.0242 | $2.904 | $1,452.00 | Not comparable |
| GLM-5 (Z.ai) | $0.0249 | $2.99 | $1,495.20 | Not comparable |
| Perplexity Sonar | $0.0256 | $3.072 | $1,536.00 | Not comparable |
| Cerebras — Gemma 4 31B | $0.0256 | $3.077 | $1,538.64 | Not comparable |
| Claude Haiku 4.5 | $0.0258 | $3.09 | $1,545.00 | Not comparable |
| OpenRouter Auto-Best Router | $0.0268 | $3.216 | $1,608.00 | Not comparable |
| GLM-4.5-AirX (Z.ai) | $0.028 | $3.36 | $1,680.00 | Not comparable |
| o4-mini | $0.0281 | $3.369 | $1,684.65 | Not comparable |
| o3-mini | $0.0288 | $3.452 | $1,725.90 | Not comparable |
| o1-mini | $0.0288 | $3.452 | $1,725.90 | Not comparable |
| GLM-5-Turbo (Z.ai) | $0.03 | $3.60 | $1,800.00 | Not comparable |
| GLM-5V-Turbo (Z.ai) | $0.03 | $3.60 | $1,800.00 | Not comparable |
| Grok 4.3 | $0.0301 | $3.615 | $1,807.50 | Not comparable |
| Gemini 1.5 Pro | $0.0319 | $3.829 | $1,914.375 | Not comparable |
| DeepSeek V4 Pro | $0.0322 | $3.862 | $1,931.16 | Not comparable |
| GLM-5.3 (Z.ai) | $0.0348 | $4.175 | $2,087.40 | Not comparable |
| GLM-5.2 (Z.ai) | $0.0348 | $4.175 | $2,087.40 | Not comparable |
| GLM-5.1 (Z.ai) | $0.0348 | $4.175 | $2,087.40 | Not comparable |
| Gemini 2.5 Pro | $0.0349 | $4.189 | $2,094.375 | Not comparable |
| Gemini 3.7 Flash | $0.0377 | $4.527 | $2,263.50 | Not comparable |
| Kimi K2.6 | $0.0377 | $4.527 | $2,263.50 | Not comparable |
| Gemini 3.6 Flash | $0.0386 | $4.635 | $2,317.50 | Not comparable |
| Mistral Medium 3.5 | $0.0386 | $4.635 | $2,317.50 | Not comparable |
| NVIDIA Nemotron-4 340B | $0.0393 | $4.716 | $2,358.00 | Not comparable |
| Gemini 3.5 Flash (OpenRouter) | $0.0395 | $4.743 | $2,371.50 | Not comparable |
| Qwen 2.5 Max | $0.0402 | $4.829 | $2,414.40 | Not comparable |
| Llama 3.1 405B Instruct | $0.0458 | $5.502 | $2,751.00 | Not comparable |
| GPT-5.3 Codex | $0.0482 | $5.786 | $2,892.75 | Not comparable |
| Pixtral Large | $0.0491 | $5.892 | $2,946.00 | Not comparable |
| Mistral Large 2 (2411) | $0.0491 | $5.892 | $2,946.00 | Not comparable |
| Qwen 3.8 Max (2.4T MoE) | $0.0491 | $5.892 | $2,946.00 | Not comparable |
| Grok 4.5 | $0.0494 | $5.922 | $2,961.00 | Not comparable |
| Grok 4.6 Flagship | $0.0499 | $5.982 | $2,991.00 | Not comparable |
| Amazon Nova Premier | $0.0503 | $6.036 | $3,018.00 | Not comparable |
| Claude Sonnet 5 | $0.0515 | $6.18 | $3,090.00 | Not comparable |
| GPT-5.6 Sol Pro (OpenRouter) | $0.0515 | $6.18 | $3,090.00 | Not comparable |
| GPT-5.6 Terra | $0.0527 | $6.324 | $3,162.00 | Not comparable |
| Gemini 3.1 Pro | $0.0527 | $6.324 | $3,162.00 | Not comparable |
| GPT-5.6 Terra Pro (OpenRouter) | $0.0527 | $6.324 | $3,162.00 | Not comparable |
| Perplexity Sonar Reasoning Pro | $0.0548 | $6.576 | $3,288.00 | Not comparable |
| Perplexity Sonar Deep Research | $0.0548 | $6.576 | $3,288.00 | Not comparable |
| AI21 Jamba 1.5 Large | $0.0548 | $6.576 | $3,288.00 | Not comparable |
| GLM-4.5-X (Z.ai) | $0.056 | $6.716 | $3,357.90 | Not comparable |
| Grok 2 | $0.056 | $6.72 | $3,360.00 | Not comparable |
| Grok 2 Vision | $0.056 | $6.72 | $3,360.00 | Not comparable |
| Cohere Command A+ | $0.0629 | $7.545 | $3,772.50 | Not comparable |
| GPT-4o (Omni) | $0.0654 | $7.845 | $3,922.50 | Not comparable |
| GPT-5.4 Workhorse | $0.0659 | $7.905 | $3,952.50 | Not comparable |
| Cohere Command R+ | $0.0685 | $8.22 | $4,110.00 | Not comparable |
| Claude 3.7 Sonnet (Hybrid Reasoning) | $0.0773 | $9.27 | $4,635.00 | Not comparable |
| Claude 3.5 Sonnet | $0.0773 | $9.27 | $4,635.00 | Not comparable |
| Kimi K3 (Moonshot) | $0.0773 | $9.27 | $4,635.00 | Not comparable |
| Claude Sonnet 4.6 (OpenRouter) | $0.0773 | $9.27 | $4,635.00 | Not comparable |
| Perplexity Sonar Pro | $0.084 | $10.08 | $5,040.00 | Not comparable |
| Claude Opus 5 | $0.1287 | $15.45 | $7,725.00 | Not comparable |
| Claude Opus 4.8 (OpenRouter) | $0.1287 | $15.45 | $7,725.00 | Not comparable |
| GPT-5.6 Sol | $0.1317 | $15.81 | $7,905.00 | Not comparable |
| GPT-5.5 Standard | $0.1317 | $15.81 | $7,905.00 | Not comparable |
| o3 (Reasoning Frontier) | $0.2515 | $30.18 | $15,090.00 | Not comparable |
| GPT-4 Turbo | $0.2555 | $30.66 | $15,330.00 | Not comparable |
| Claude Fable 5 | $0.2575 | $30.90 | $15,450.00 | Not comparable |
| GPT-5.6 Cyber | $0.3294 | $39.525 | $19,762.50 | Not comparable |
| Claude 3 Opus | $0.3862 | $46.35 | $23,175.00 | Not comparable |
| o1 (Reasoning) | $0.3922 | $47.07 | $23,535.00 | Not comparable |
| o3-pro (Frontier Reasoning) | $0.503 | $60.36 | $30,180.00 | Not comparable |
| GPT-5.5 Pro | $0.7905 | $94.86 | $47,430.00 | Not comparable |
Assumes the typical cacheable share (10% of input at cached rates where published). 4 requests/user/day. Adjust everything in the monthly calculator.
Best value picks for Document summarization
| Model | Fit score | Cost / 1K requests | Value (fit ÷ cost) |
|---|---|---|---|
| best valueGLM-5.3-Flash (Z.ai) | 59.1 | $1.875 | 31.5 |
| GPT-5.6 Luna | 52.6 | $5.27 | 10.0 |
| GPT-5.4 nano | 38.3 | $5.30 | 7.2 |
| DeepSeek V3.2 (OpenRouter) | 44.2 | $6.403 | 6.9 |
| MiniMax M3 (OpenRouter) | 43.9 | $7.62 | 5.8 |
Fit = weighted blend of third-party composite indices (intelligence, coding, agentic) for this use case; value = fit ÷ cost per 1,000 requests. Models without a verified composite are excluded. Methodology · Benchmark hub.
How to spend less
Related workflow costs
FAQ
Using a typical profile of 25,000 input and 600 output tokens with 10% cacheable input: from Free on GLM-4.7-Flash (Z.ai) up to $0.7905 on GPT-5.5 Pro. See the table for every model.
At 4 requests per user per day and the typical token profile, budget from Free per active user per month on GLM-4.7-Flash (Z.ai). A team of 500 active users lands around Free/month at that tier. Model your exact numbers in the monthly cost calculator.
Long-context models pay off here — compare price per 1M tokens at your true document size. Summarize once, store the result; don't re-summarize unchanged documents. For batch backfills, nightly jobs can use cache-friendly request ordering.