deepseek-v4-flash · DeepSeek
The value monster of 2026: $0.44/$1.32 with a 1M-token context and 384K max output, cache hits at $0.014/M.
Last checked Aug 16, 2026. Calculations use the listed base rates; provider-specific tiers, cache writes, batch pricing, and long-context rules may differ. Note: Peak-hour rates shown; off-peak (most of the day) is 50% off: $0.22/$0.66. Confirm the current rate at the official provider source.
Input
$0.44
per 1M tokens
Output
$1.32
per 1M tokens
Cached input
$0.014
per 1M tokens
Context window
1M
max output 384K
DeepSeek V4 Flash cost calculator
What DeepSeek V4 Flash costs per task
| Use case | Input tokens | Output tokens | Cost / request | Monthly @ 1K req/day |
|---|---|---|---|---|
| Customer support chatbot | 3,500 | 350 | $0.000958 | $28.749 |
| RAG / search-augmented answers | 8,000 | 500 | $0.002476 | $74.28 |
| AI coding assistant | 12,000 | 2,000 | $0.004853 | $145.584 |
| Document summarization | 25,000 | 600 | $0.0107 | $321.81 |
| Agentic workflow | 40,000 | 1,500 | $0.008504 | $255.12 |
| Content generation | 800 | 1,200 | $0.001834 | $55.013 |
| Data extraction & tagging | 2,000 | 250 | $0.00104 | $31.188 |
| Translation | 5,000 | 5,500 | $0.009141 | $274.215 |
Assumes each use case's typical cacheable share of input. See full cost scenarios.
Compare with alternatives
| Model | Input /M | Output /M | Context | Chat request* |
|---|---|---|---|---|
| DeepSeek V4 Flash | $0.44 | $1.32 | 1M | $0.001964 |
| GPT-5.6 Lunacompare | $0.20 | $1.20 | 1.1M | $0.0014 |
| Grok 4.3 | $1.25 | $2.50 | 1M | $0.0049 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | $0.00054 |
*4,000 in + 800 out tokens, 50% cached input where available.
Pricing history
Related calculations
FAQ
DeepSeek V4 Flash costs $0.44 per 1M input tokens and $1.32 per 1M output tokens, with cached input at $0.014 per 1M tokens. Verified Aug 16, 2026 against https://api-docs.deepseek.com/quick_start/pricing.
DeepSeek V4 Flash supports a 1,000,000-token context window with up to 384,000 output tokens per request, tokenized with DeepSeek BPE (V4).
At $0.44/M input and $1.32/M output, DeepSeek V4 Flash sits above GPT-5.6 Luna ($0.20/M in, $1.20/M out) — see the comparison table for full-workload differences.