Gemini 3.x, 2.5, 2.0, and 1.5 series: 2M token context, ultra-fast Flash tiers, context caching discounts up to 90%, and multimodal native reasoning.
16 models tracked · cheapest input: Text Embedding 004 at $0.02/M · official pricing page
Gemini 3.x spans 2M-token-context Flash tiers to Pro reasoning models, with automatic prompt caching and aggressive batch pricing. Flash models are among the fastest budget options in the market; Pro competes on deep-context reasoning and multimodal input.
gemini-3.1-pro
Input /M
$2.00
Output /M
$12.00
Context
2M
Google 3.x flagship with 2M token context window, recursive multi-agent planning, and native multimodal reasoning at $2.00/$12.00.
gemini-3.7-flash
Input /M
$1.50
Output /M
$6.00
Context
1M
Google's ultra-fast reasoning Flash model with dynamic thinking budget and sub-second TTFT at $1.50/$6.00.
gemini-3.6-flash
Input /M
$1.50
Output /M
$7.50
Context
1M
Google high-throughput volume workhorse at $1.50/$7.50 with 1M context and ultra-fast TTFT.
gemini-3.5-flash-lite
Input /M
$0.30
Output /M
$2.50
Context
1M
Sub-dollar pricing at $0.30/$2.50 for high-volume enterprise ingestion, classification, and OCR extraction.
gemini-2.5-pro
Input /M
$1.25
Output /M
$10.00
Context
1M
Gemini 2.5 flagship model with 1M context window and enhanced complex reasoning at $1.25/$10.00.
gemini-2.5-flash
Input /M
$0.30
Output /M
$2.50
Context
1M
High-speed 2.5 generation Flash model with built-in thinking trace at $0.30/$2.50 per 1M tokens.
gemini-2.5-flash-lite
Input /M
$0.10
Output /M
$0.40
Context
1M
Ultra-budget 2.5 tier at $0.10/$0.40 per 1M tokens with 1M context window for high-volume tasks.
gemini-2.0-flash
Input /M
$0.10
Output /M
$0.40
Context
1M
High-speed multimodal workhorse: sub-second latency, 1M context, and native audio/video understanding.
gemini-2.0-flash-lite
Input /M
$0.075
Output /M
$0.30
Context
1M
Lowest cost Gemini 2.0 tier at $0.075/$0.30 per 1M tokens with 1M context window for extreme scale.
gemini-1.5-pro
Input /M
$1.25
Output /M
$5.00
Context
2.1M
Pioneered 2 Million token context window: processes entire codebases and hours of video in a single turn.
gemini-1.5-flash
Input /M
$0.075
Output /M
$0.30
Context
1M
Lightweight, cost-efficient legacy model with 1M context window and native multimodal support.
google/gemma-4-31b-it
Input /M
$0.25
Output /M
$0.75
Context
128K
Google's premier open weights model delivering near-Pro reasoning on consumer server hardware.
google/gemma-2-9b-it
Input /M
$0.20
Output /M
$0.20
Context
8.2K
Ultra-compact open model built on Google's modern transformer architecture at $0.20/$0.20 per 1M tokens.
text-embedding-004
Input /M
$0.02
Output /M
$0.00
Context
8.2K
Google's embedding model with 768 dimensions and 8K token context at $0.02/1M tokens.
gemini-3.1-flash-lite
Input /M
$0.25
Output /M
$1.50
Context
1M
Google multimodal Flash Lite route for low-cost, high-throughput workloads.
gemini-3.5-flash
Input /M
$1.50
Output /M
$9.00
Context
1M
Google multimodal Flash route for reasoning, coding, and high-volume agent workloads.
Other providers
All Google prices verified Aug 28, 2026. Prices change frequently — each model page links to the authoritative source.