Methodology
Transparent assumptions are more useful than false precision. This page explains what the calculators include and what they do not.
Catalog prices are represented in USD per one million tokens. A base estimate multiplies input tokens by the input rate and output tokens by the output rate, then adds the two amounts. Cached input is separated when a record includes a published cached-input rate.
Monthly projections multiply a representative request by requests per user, users and the selected number of days. These are workload scenarios, not invoices.
Token counts depend on the provider tokenizer, content, language and API serialization. The calculators provide estimates and display tokenizer information where available. Exact billing can differ because providers may count system messages, tool calls, images, cached content or reasoning tokens differently.
A workload split such as 70% input and 30% output is an illustrative scenario. Replace it with measurements from your own application when planning spend.
Some providers apply different rates for long contexts, batch or Flex processing, off-peak windows, cache writes, cache reads, reasoning tokens, regions or service tiers. The standard calculator uses listed base rates unless a pricing rule is represented explicitly. Read each model's note and confirm the linked provider documentation.
Each record in the public dataset includes a last-checked date and source URL. The catalog is manually maintained, not a real-time provider feed. Model records can contain conservative family defaults when a provider does not publish a separate specification; those caveats are shown on model pages.
The catalog currently contains 167 model records across 26 providers. Report corrections to contact@aitokenusagecalculator.com with the model, provider source and the relevant change.
Benchmark pages show third-party composite indices (intelligence, coding, agentic, 0-100) published by Artificial Analysis and exposed through OpenRouter's live model catalog. Each entry carries a source URL and verification date, matching the pricing catalog's freshness discipline. Only models with a published composite are scored; missing values are shown as "no verified score" rather than estimated. Vendor-published task scores (e.g., SWE-bench, GPQA) are displayed separately and labeled "vendor-reported" — they are never merged into the composite axes.
Fit scores weight the three composites per use case (for example, coding 0.6 / intelligence 0.3 / agentic 0.1 for code assistants; agentic 0.6 / coding 0.2 / intelligence 0.2 for agentic workflows). Value divides the fit score by the cost of 1,000 typical requests for that use case (its token mix and cacheable share). Fit weights are editorial and published here so every verdict is reproducible.
Refresh benchmark data with node scripts/fetch-openrouter-benchmarks.mjs; invariants (slug exists, scores 0-100, ISO verification dates) are enforced by node scripts/validate-data.mjs.
Use the model catalog and source links to compare options, then validate projected spend with provider documentation and a measured workload before committing to a production budget.