Process asynchronous LLM workloads at 50% off standard rates. Compare real-time vs. 24-hour batch turnaround costs across flagship and balanced models.
Half price on both input and output tokens across OpenAI, Anthropic, and Google batch queues.
Bypass strict real-time rate limits (TPM/RPM) with dedicated asynchronous execution pools.
Results delivered as completed JSONL files within 24 hours — often completing in under 1 hour.
| Model | Provider | Standard In | Batch In (50% off) | Standard Out | Batch Out (50% off) | Savings Example (10M In / 2M Out) |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | openai | $5.00 | $2.50 | $30.00 | $15.00 | Save $55.00 / batch |
| GPT-5.6 Terra | openai | $2.00 | $1.00 | $12.00 | $6.00 | Save $22.00 / batch |
| GPT-5.6 Luna | openai | $0.20 | $0.10 | $1.20 | $0.60 | Save $2.20 / batch |
| GPT-5.6 Cyber | openai | $12.50 | $6.25 | $75.00 | $37.50 | Save $137.50 / batch |
| GPT-5.5 Standard | openai | $5.00 | $2.50 | $30.00 | $15.00 | Save $55.00 / batch |
| GPT-5.5 Pro | openai | $30.00 | $15.00 | $180.00 | $90.00 | Save $330.00 / batch |
| GPT-5.4 Workhorse | openai | $2.50 | $1.25 | $15.00 | $7.50 | Save $27.50 / batch |
| GPT-5.4 mini | openai | $0.75 | $0.375 | $4.50 | $2.25 | Save $8.25 / batch |
| GPT-5.4 nano | openai | $0.20 | $0.10 | $1.25 | $0.625 | Save $2.25 / batch |
| GPT-5.3 Codex | openai | $1.75 | $0.875 | $14.00 | $7.00 | Save $22.75 / batch |
| o3-pro (Frontier Reasoning) | openai | $20.00 | $10.00 | $80.00 | $40.00 | Save $180.00 / batch |
| o3 (Reasoning Frontier) | openai | $10.00 | $5.00 | $40.00 | $20.00 | Save $90.00 / batch |
| o3-mini | openai | $1.10 | $0.55 | $4.40 | $2.20 | Save $9.90 / batch |
| o4-mini | openai | $1.10 | $0.55 | $4.40 | $2.20 | Save $9.90 / batch |
A Batch API allows developers to submit non-urgent requests asynchronously (typically as JSONL files) and receive completions within 24 hours at a 50% flat discount compared to standard real-time endpoints.
OpenAI, Anthropic, and Google Vertex AI all provide first-party batch endpoints offering ~50% savings on input, output, and cached-input token rates.
Use Batch APIs for offline pipelines: dataset enrichment, backfill summaries, synthetic data generation, nightly document indexing, and bulk translation. Use Real-time APIs for synchronous user-facing chat, interactive coding, and live webhooks.
Yes. Batch queues typically have separate, significantly higher throughput and token-per-minute (TPM) capacity limits than real-time interactive tiers, making them ideal for high-volume jobs without rate limit errors.