Explore verified first-party rate cards and model catalogs by provider. Each hub lists the complete model lineup with input, output and prompt-caching rates, context windows, and official documentation source links.
The GPT-5.x family and o-series reasoning models: three tiers per generation (Sol/Terra/Luna naming since 5.6), 1M+ context on latest generations, and 90% cached-input discounts.
Claude models with the widest context in the industry: Sonnet 5 and Opus 5 ship 1M-token windows at standard pricing, with fine-grained prompt caching (cache reads at ~10% of input).
Gemini 3.x and 2.5 families: Pro with tiered long-context pricing, Flash for volume, Flash-Lite for the cheapest usable tier — plus explicit context-caching rates.
Grok 4.x models served from xAI's Colossus clusters, with tiered pricing above 200K-token prompts and up to 1M-token context on grok-4.3.
European provider with aggressive 2026 pricing: Mistral Large 3 at $0.50/$1.50 per 1M, plus Small, Ministral and Codestral tiers — cached input at 90% off.
The Meta Model API serves Muse Spark direct and self-serve. Llama 4 models remain open-weight with no first-party metered price, so they are excluded from this priced catalog.
The V4 generation: 1M-token context, 384K max output, and cache hits at ~3% of fresh input cost — with 50% off-peak discounts.
Kimi K3: a 1M-token-context flagship from Beijing, priced against premium Western tiers with aggressive cache-hit discounts.
GLM 5.x models. First-party Z.ai USD rates are credit-multiplier based; the USD rate below is GLM 5.2 served through Mistral's official API.