meta-llama/llama-3.1-70b-instruct · Meta Llama
Meta open-weight text route for general instruction workloads.
Last checked Aug 26, 2026. Calculations use the listed base rates; provider-specific tiers, cache writes, batch pricing, and long-context rules may differ. Note: Live OpenRouter route snapshot checked 26 Aug 2026. Provider and host pricing may differ. Confirm the current rate at the official provider source.
Quick answer
At the published base rate, Llama 3.1 70B Instruct (OpenRouter) costs $0.40 per 1M input tokens and $0.40 per 1M output tokens. It supports a 131.1K context window and is currently marked ga.
Method & trust
This rate card uses the provider's listed base input, output, and cached-input prices. Workload tables below apply those rates to explicit token counts, cache assumptions, and request volumes; verify provider-specific tiers before committing budget.
Input
$0.40
per 1M tokens
Output
$0.40
per 1M tokens
Cached input
—
not published
Context window
131.1K
max output 16.4K
Llama 3.1 70B Instruct (OpenRouter) cost calculator
What Llama 3.1 70B Instruct (OpenRouter) costs per task
| Use case | Input tokens | Output tokens | Cost / request | Monthly @ 1K req/day |
|---|---|---|---|---|
| Customer support chatbot | 3,500 | 350 | $0.00154 | $46.20 |
| RAG / search-augmented answers | 8,000 | 500 | $0.0034 | $102.00 |
| AI coding assistant | 12,000 | 2,000 | $0.0056 | $168.00 |
| Document summarization | 25,000 | 600 | $0.0102 | $307.20 |
| Agentic workflow | 40,000 | 1,500 | $0.0166 | $498.00 |
| Content generation | 800 | 1,200 | $0.0008 | $24.00 |
| Data extraction & tagging | 2,000 | 250 | $0.0009 | $27.00 |
| Translation | 5,000 | 5,500 | $0.0042 | $126.00 |
Assumes each use case's typical cacheable share of input. See full cost scenarios.
Compare with alternatives
| Model | Input /M | Output /M | Context | Chat request* |
|---|---|---|---|---|
| Llama 3.1 70B Instruct (OpenRouter) | $0.40 | $0.40 | 131.1K | $0.00192 |
| GPT-5.6 Solcompare | $5.00 | $30.00 | 1.1M | $0.035 |
| GPT-5.6 Terracompare | $2.00 | $12.00 | 1.1M | $0.014 |
| GPT-5.6 Lunacompare | $0.20 | $1.20 | 1.1M | $0.0014 |
| GPT-5.6 Cybercompare | $12.50 | $75.00 | 1.1M | $0.0875 |
| GPT-5.5 Standardcompare | $5.00 | $30.00 | 512K | $0.035 |
| GPT-5.5 Procompare | $30.00 | $180.00 | 512K | $0.21 |
| GPT-5.4 Workhorsecompare | $2.50 | $15.00 | 256K | $0.0175 |
| GPT-5.4 minicompare | $0.75 | $4.50 | 256K | $0.00525 |
*4,000 in + 800 out tokens, 50% cached input where available.
Related calculations
FAQ
Llama 3.1 70B Instruct (OpenRouter) costs $0.40 per 1M input tokens and $0.40 per 1M output tokens. Verified Aug 26, 2026 against https://openrouter.ai/api/v1/models.
Llama 3.1 70B Instruct (OpenRouter) supports a 131,072-token context window with up to 16,384 output tokens per request, tokenized with llama_bpe.
At $0.40/M input and $0.40/M output, Llama 3.1 70B Instruct (OpenRouter) sits below GPT-5.6 Sol ($5.00/M in, $30.00/M out) — see the comparison table for full-workload differences.