Llama 4 Maverick (400B MoE), Llama 4 Scout (109B MoE with 10M context), Llama 3.3 70B, Llama 3.1 405B/70B/8B, and Llama 3.2 Vision: open-weight foundation models.
11 models tracked · cheapest input: Llama 3.2 1B Instruct (OpenRouter) at $0.027/M · official pricing page
Meta's open-weight Llama 4 family runs nearly everywhere — first-party, Bedrock, Together, Groq — making it the portability play: same model, many hosts, competitive token prices, and a 10M-token Scout context option.
meta-llama/llama-4-maverick
Input /M
$0.45
Output /M
$1.25
Context
1M
Meta's flagship open-weights MoE model (400B total, 17B active across 128 experts) with 1M context and native multimodal pre-training.
meta-llama/llama-4-scout
Input /M
$0.15
Output /M
$0.45
Context
10M
Meta's ultra-long-context open model: 109B total (17B active across 16 experts) with unprecedented 10 Million token context window.
meta-llama/llama-3.3-70b-instruct
Input /M
$0.18
Output /M
$0.59
Context
128K
Industry-standard open-weights dense 70B model matching 405B capabilities at 70B efficiency ($0.18/$0.59 on leading clouds).
meta-llama/llama-3.1-405b-instruct
Input /M
$1.75
Output /M
$3.50
Context
128K
The world's largest open-weights model: rivaling proprietary frontier models in synthetic data generation and complex reasoning.
meta-llama/llama-3.1-8b-instruct
Input /M
$0.05
Output /M
$0.08
Context
128K
Ultra-fast 8B parameter model ideal for edge devices, local fine-tuning, and ultra-high-throughput classification at $0.05/$0.08.
meta-llama/llama-3.2-90b-vision-instruct
Input /M
$0.90
Output /M
$0.90
Context
128K
Meta's open multimodal model for chart understanding, document visual reasoning, and image QA.
meta-llama/llama-3.2-11b-vision-instruct
Input /M
$0.16
Output /M
$0.16
Context
128K
Lightweight visual reasoning open model capable of running on consumer edge hardware and fast serverless APIs.
meta-llama/llama-3.2-3b-instruct
Input /M
$0.03
Output /M
$0.05
Context
128K
Edge and on-device compact model for mobile deployment and embedded agents at $0.03/$0.05 per 1M tokens.
meta-llama/llama-3.1-70b-instruct
Input /M
$0.40
Output /M
$0.40
Context
131.1K
Meta open-weight text route for general instruction workloads.
meta-llama/llama-3.2-1b-instruct
Input /M
$0.027
Output /M
$0.201
Context
60K
Compact Meta open-weight text route for low-latency, high-volume inference.
meta/muse-spark-1.2-contributor
Input /M
$0.10
Output /M
$0.20
Context
1.1M
Meta's contributor-tier reasoning model: 1M-token context, multimodal input, multi-agent workflows and structured output at a minimal price. Prompts may be used to improve Meta products.
Other providers
All Meta Llama prices verified Aug 28, 2026. Prices change frequently — each model page links to the authoritative source.