Cost-effective serverless inference provider delivering open models (Llama 3.3, DeepSeek R1, Qwen 2.5) with pure per-token pay-as-you-go pricing.
2 models tracked · cheapest input: DeepInfra — Qwen 2.5 Coder 32B at $0.08/M · official pricing page
meta-llama/Llama-3.3-70B-Instruct
Input /M
$0.10
Output /M
$0.30
Context
128K
Low-cost serverless Llama 3.3 70B inference at $0.10/$0.30 per 1M tokens with pure usage-based billing and zero minimums.
Qwen/Qwen2.5-Coder-32B-Instruct
Input /M
$0.08
Output /M
$0.24
Context
128K
Direct serverless inference of Qwen 2.5 Coder at $0.08/$0.24 per 1M tokens.
Other providers
All DeepInfra prices verified Aug 19, 2026. Prices change frequently — each model page links to the authoritative source.