Ultra-low latency inference on custom Language Processing Units (LPUs): 300–800+ tokens/sec on Llama 4, Llama 3.3, QwQ, and Mistral models.
5 models tracked · cheapest input: Groq LPU — Llama 3.1 8B Instant at $0.05/M · official pricing page
llama-4-scout-lpu
Input /M
$0.11
Output /M
$0.34
Context
512K
Llama 4 Scout served on Groq LPU hardware delivering 450+ tokens/second with linear per-token pricing ($0.11/$0.34).
llama-4-maverick-lpu
Input /M
$0.20
Output /M
$0.60
Context
512K
400B MoE model running on interconnected LPUs at 250+ tokens/second at $0.20/$0.60 per 1M tokens.
llama-3.3-70b-versatile
Input /M
$0.59
Output /M
$0.79
Context
128K
Industry benchmark for high-speed 70B inference: 300+ tokens/second with sub-150ms TTFT at $0.59/$0.79.
llama-3.1-8b-instant
Input /M
$0.05
Output /M
$0.08
Context
128K
Ultra-fast edge inference delivering 800+ tokens/second at $0.05/$0.08 per 1M tokens.
qwen-qwq-32b-preview
Input /M
$0.29
Output /M
$0.39
Context
128K
Chain-of-thought reasoning served at interactive chat speeds (180+ tokens/sec) on LPUs at $0.29/$0.39.
Other providers
All Groq LPU prices verified Aug 19, 2026. Prices change frequently — each model page links to the authoritative source.