glm-4.6v-flash · Z.ai (Zhipu)
Free, lightweight Z.ai vision-language model with native function calling.
Last checked Aug 26, 2026. Calculations use the listed base rates; provider-specific tiers, cache writes, batch pricing, and long-context rules may differ. Note: Z.ai lists input, output, and cached input as free. Availability, quotas, and access requirements may vary by route. Confirm the current rate at the official provider source.
Quick answer
At the published base rate, GLM-4.6V-Flash (Z.ai) costs $0.00 per 1M input tokens and $0.00 per 1M output tokens. It supports a 128K context window and is currently marked ga.
Method & trust
This rate card uses the provider's listed base input, output, and cached-input prices. Workload tables below apply those rates to explicit token counts, cache assumptions, and request volumes; verify provider-specific tiers before committing budget.
Input
Free
per 1M tokens
Output
Free
per 1M tokens
Cached input
Free
per 1M tokens
Context window
128K
max output 32.8K
GLM-4.6V-Flash (Z.ai) cost calculator
What GLM-4.6V-Flash (Z.ai) costs per task
| Use case | Input tokens | Output tokens | Cost / request | Monthly @ 1K req/day |
|---|---|---|---|---|
| Customer support chatbot | 3,500 | 350 | $0.00 | $0.00 |
| RAG / search-augmented answers | 8,000 | 500 | $0.00 | $0.00 |
| AI coding assistant | 12,000 | 2,000 | $0.00 | $0.00 |
| Document summarization | 25,000 | 600 | $0.00 | $0.00 |
| Agentic workflow | 40,000 | 1,500 | $0.00 | $0.00 |
| Content generation | 800 | 1,200 | $0.00 | $0.00 |
| Data extraction & tagging | 2,000 | 250 | $0.00 | $0.00 |
| Translation | 5,000 | 5,500 | $0.00 | $0.00 |
Assumes each use case's typical cacheable share of input. See full cost scenarios.
Compare with alternatives
| Model | Input /M | Output /M | Context | Chat request* |
|---|---|---|---|---|
| GLM-4.6V-Flash (Z.ai) | Free | Free | 128K | $0.00 |
| GPT-5.6 Solcompare | $5.00 | $30.00 | 1.1M | $0.035 |
| GPT-5.6 Terracompare | $2.00 | $12.00 | 1.1M | $0.014 |
| GPT-5.6 Lunacompare | $0.20 | $1.20 | 1.1M | $0.0014 |
| GPT-5.6 Cybercompare | $12.50 | $75.00 | 1.1M | $0.0875 |
| GPT-5.5 Standardcompare | $5.00 | $30.00 | 512K | $0.035 |
| GPT-5.5 Procompare | $30.00 | $180.00 | 512K | $0.21 |
| GPT-5.4 Workhorsecompare | $2.50 | $15.00 | 256K | $0.0175 |
| GPT-5.4 minicompare | $0.75 | $4.50 | 256K | $0.00525 |
*4,000 in + 800 out tokens, 50% cached input where available.
Related calculations
FAQ
GLM-4.6V-Flash (Z.ai) costs Free per 1M input tokens and Free per 1M output tokens, with cached input at Free per 1M tokens. Verified Aug 26, 2026 against https://docs.z.ai/guides/overview/pricing.
GLM-4.6V-Flash (Z.ai) supports a 128,000-token context window with up to 32,768 output tokens per request, tokenized with glm_bpe.
At Free/M input and Free/M output, GLM-4.6V-Flash (Z.ai) sits below GPT-5.6 Sol ($5.00/M in, $30.00/M out) — see the comparison table for full-workload differences.