meta-llama/llama-3.2-1b-instruct · Meta Llama
Compact Meta open-weight text route for low-latency, high-volume inference.
Last checked Aug 26, 2026. Calculations use the listed base rates; provider-specific tiers, cache writes, batch pricing, and long-context rules may differ. Note: Live OpenRouter route snapshot checked 26 Aug 2026. Provider and host pricing may differ. Confirm the current rate at the official provider source.
Quick answer
At the published base rate, Llama 3.2 1B Instruct (OpenRouter) costs $0.027 per 1M input tokens and $0.201 per 1M output tokens. It supports a 60K context window and is currently marked ga.
Method & trust
This rate card uses the provider's listed base input, output, and cached-input prices. Workload tables below apply those rates to explicit token counts, cache assumptions, and request volumes; verify provider-specific tiers before committing budget.
Input
$0.027
per 1M tokens
Output
$0.201
per 1M tokens
Cached input
—
not published
Context window
60K
max output 54K
Llama 3.2 1B Instruct (OpenRouter) cost calculator
What Llama 3.2 1B Instruct (OpenRouter) costs per task
| Use case | Input tokens | Output tokens | Cost / request | Monthly @ 1K req/day |
|---|---|---|---|---|
| Customer support chatbot | 3,500 | 350 | $0.000165 | $4.946 |
| RAG / search-augmented answers | 8,000 | 500 | $0.000317 | $9.495 |
| AI coding assistant | 12,000 | 2,000 | $0.000726 | $21.78 |
| Document summarization | 25,000 | 600 | $0.000796 | $23.868 |
| Agentic workflow | 40,000 | 1,500 | $0.001382 | $41.445 |
| Content generation | 800 | 1,200 | $0.000263 | $7.884 |
| Data extraction & tagging | 2,000 | 250 | $0.000104 | $3.128 |
| Translation | 5,000 | 5,500 | $0.00124 | $37.215 |
Assumes each use case's typical cacheable share of input. See full cost scenarios.
Compare with alternatives
| Model | Input /M | Output /M | Context | Chat request* |
|---|---|---|---|---|
| Llama 3.2 1B Instruct (OpenRouter) | $0.027 | $0.201 | 60K | $0.000269 |
| GPT-5.6 Solcompare | $5.00 | $30.00 | 1.1M | $0.035 |
| GPT-5.6 Terracompare | $2.00 | $12.00 | 1.1M | $0.014 |
| GPT-5.6 Lunacompare | $0.20 | $1.20 | 1.1M | $0.0014 |
| GPT-5.6 Cybercompare | $12.50 | $75.00 | 1.1M | $0.0875 |
| GPT-5.5 Standardcompare | $5.00 | $30.00 | 512K | $0.035 |
| GPT-5.5 Procompare | $30.00 | $180.00 | 512K | $0.21 |
| GPT-5.4 Workhorsecompare | $2.50 | $15.00 | 256K | $0.0175 |
| GPT-5.4 minicompare | $0.75 | $4.50 | 256K | $0.00525 |
*4,000 in + 800 out tokens, 50% cached input where available.
Related calculations
FAQ
Llama 3.2 1B Instruct (OpenRouter) costs $0.027 per 1M input tokens and $0.201 per 1M output tokens. Verified Aug 26, 2026 against https://openrouter.ai/api/v1/models.
Llama 3.2 1B Instruct (OpenRouter) supports a 60,000-token context window with up to 54,000 output tokens per request, tokenized with llama_bpe.
At $0.027/M input and $0.201/M output, Llama 3.2 1B Instruct (OpenRouter) sits below GPT-5.6 Sol ($5.00/M in, $30.00/M out) — see the comparison table for full-workload differences.