Head-to-head showdown: Perplexity Sonar Deep Research ($2.00 in / $8.00 out per 1M) vs Gemini 3.5 Flash (OpenRouter) ($1.50 in / $9.00 out per 1M). Perplexity Sonar Deep Research is 1.1× cheaper across standard token mixes, with 128K vs 1M context windows.
sonar-deep-research · perplexity
gemini-3.5-flash · google
Benchmark
| Workload Scenario | Perplexity Sonar Deep Research | Gemini 3.5 Flash (OpenRouter) | Price Delta |
|---|---|---|---|
| 1M input tokens (raw text) | $2.00 | $1.50 | +$0.50 |
| 1M output tokens (generation) | $8.00 | $9.00 | −$1.00 |
| 1M tokens · 70% input / 30% output mix | $3.80 | $3.75 | +$0.05 |
| Standard chat turn (4K in / 800 out, 50% cached) | $0.0144 | $0.0105 | +$0.0039 |
| Monthly scale (10K requests / day) | $4,320.00 | $3,150.00 | +$1,170.00 |
Negative difference = Perplexity Sonar Deep Research is cheaper. Positive = Gemini 3.5 Flash (OpenRouter) is cheaper.
Scaling Curve
Total cost of a token volume at 70/30 input/output split (uncached). Log-log scale.
Verdict
Gemini 3.5 Flash (OpenRouter) has the cheaper input rate, while Perplexity Sonar Deep Research has the cheaper output rate. The crossover happens when output makes up about 33% of your total tokens.
Below that share (retrieval, summarization, extraction — lots of context in, little text out) Gemini 3.5 Flash (OpenRouter) is cheaper. Above it (generation, translation, coding — long completions) Perplexity Sonar Deep Research wins.
FAQ
On input, Gemini 3.5 Flash (OpenRouter) is cheaper ($1.50/M vs $2.00/M). On output, Perplexity Sonar Deep Research is cheaper ($8.00/M vs $9.00/M). For workloads where more than 33% of tokens are output, the output-cheaper model wins overall.
A chat-style request (4,000 input + 800 output tokens, 50% cached) costs $0.0144 on Perplexity Sonar Deep Research and $0.0105 on Gemini 3.5 Flash (OpenRouter) — Gemini 3.5 Flash (OpenRouter) is 1.4× more expensive for that workload.
Perplexity Sonar Deep Research supports 128,000 tokens (16.4K max output); Gemini 3.5 Flash (OpenRouter) supports 1,048,576 (65.5K max output). Gemini 3.5 Flash (OpenRouter) fits 8.2× more context, which matters for long documents and agents.