Explore verified first-party rate cards and model catalogs by provider. Each hub lists the complete model lineup with input, output and prompt-caching rates, context windows, and official documentation source links.
The GPT-5.6 family (Sol, Terra, Luna), GPT-5.5/5.4 generations, and unified reasoning systems: 1M+ context on frontier tiers, 90% prompt caching discounts, and 50% batch discounts.
Claude 5 generation (Opus 5, Sonnet 5, Fable 5, Haiku 4.5) alongside Claude 3.7 Sonnet hybrid reasoning with industry-leading context windows and 90% prompt caching savings.
Gemini 3.x, 2.5, 2.0, and 1.5 series: 2M token context, ultra-fast Flash tiers, context caching discounts up to 90%, and multimodal native reasoning.
DeepSeek V4 Pro, V4 Flash, R1 Reasoning, and V3 Chat: 1M token context, disruptive low pricing, peak/off-peak tier discounts, and 90% cache-hit savings.
Llama 4 Maverick (400B MoE), Llama 4 Scout (109B MoE with 10M context), Llama 3.3 70B, Llama 3.1 405B/70B/8B, and Llama 3.2 Vision: open-weight foundation models.
European open-weight leader: Mistral Large 3, Medium 3.5, Small 4, Codestral 2501 code engine, Pixtral multimodal, and Ministral edge models.
Grok 4.6 Flagship, Grok 4.5, Grok 4.3, and Grok Build 0.1 models served from Colossus clusters with up to 1M-token context and cheap output token rates.
Qwen3.8-Max and Qwen3.8-Flash, Qwen3.7/3.6/3.5 families, Qwen3-VL and Coder lines, plus Qwen2.5 models offering top-tier price-performance.
Enterprise RAG and agentic powerhouses: Command A+ (218B MoE unified), Command R+, Command R, Embed v3.0, and high-performance Re-ranking models.
Amazon Nova Premier, Nova 2 Pro, Nova 2 Lite, Nova Pro, Nova Lite, and Nova Micro: cost-efficient foundation models natively integrated into AWS Bedrock.
Sonar Reasoning Pro, Sonar Deep Research, Sonar Pro, and Sonar online models with real-time web-grounded search synthesis.
Ultra-low latency inference on custom Language Processing Units (LPUs): 300–800+ tokens/sec on Llama 4, Llama 3.3, QwQ, and Mistral models.
Extreme wafer-scale engine inference: up to 3,000 tokens/second on GPT-OSS-120B, Gemma 4, Qwen 3, and GLM 4.7 with generous daily free developer tiers.
High-performance GPU cloud delivering serverless and dedicated endpoints for Llama 4, DeepSeek V4/R1, Qwen 2.5, and open-weight models with fine-tuning.
Speed-optimized FireOptimizer serverless inference platform serving Llama 4 Maverick/Scout, DeepSeek V4/R1, and custom fine-tuned LoRA models.
Yi-Lightning and Yi-Large models: highly optimized bilingual Chinese/English models with fast throughput.
Kimi K3 and Kimi K2.6: 1M-token-context pioneer from Beijing with fine-grained cache discounts and deep document reasoning.
Z.ai's GLM-5.3/5.3-Flash/5.2/5.1/5 and 4.x text families, plus GLM-V vision models, with 1M-token context on the latest flagship releases.
Jamba 1.5 Large and Mini: hybrid SSM-Transformer (Mamba) architecture providing ultra-efficient 256K long-context throughput.
MiniMax-01 and Text-01: high-performance 4M-token context window models designed for massive document synthesis and multilingual dialog.
Cost-effective serverless inference provider delivering open models (Llama 3.3, DeepSeek R1, Qwen 2.5) with pure per-token pay-as-you-go pricing.
Universal AI gateway routing queries across 200+ models with smart fallback, auto-best routing, and unified billing across providers.
Phi-4 and Phi-3.5 family of highly capable small language models (SLMs) delivering state-of-the-art synthetic reasoning at compact footprints.
DBRX Instruct and Base: enterprise-grade 132B open Mixture-of-Experts (MoE) foundation models built for Lakehouse AI pipelines.
Nemotron-4 340B and Llama-3.1-Nemotron-70B: synthetic data generation, reward modeling, and enterprise alignment engines optimized with TensorRT-LLM.
Hy-MT2 translation models covering 33 language pairs and five Chinese dialect pairs with glossary, context and style-guided workflows.