Every developer has encountered the mismatch between estimated token counts and real API invoices. Here is the technical reality behind Byte-Pair Encoding, tokenizer vocabularies, and payload overheads.
OpenAI's newer o200k_base has a 200,000-token vocabulary, compressing non-English text and code 15–30% more densely than older 100K tokenizers like cl100k_base. The same text consumes fewer tokens on newer models.
When you send a structured messages array (`role: 'system'`, `content: '...'`), providers add 3–7 framing tokens per message for role tags and turn delimiters. A 10-turn conversation incurs ~50 hidden framing tokens.
English words are usually 1 token. But Chinese, Japanese, Korean, Arabic, and Hindi characters often split into 2–4 tokens per character on older tokenizers, causing unexpected 3× price inflation for internationalized features.
Special characters, quotation marks, camelCase variable names, and whitespace indents split into individual sub-word tokens. Minifying JSON payloads before passing them to the prompt can save 15–25% in tokens.
The 1 word = 1.33 tokens heuristic only holds for plain English prose. For source code (indentation, braces, variable camelCase), JSON structures, numbers, and non-Latin alphabets (CJK, Arabic, Cyrillic), 1 word can easily expand into 3 to 10 tokens.
No. OpenAI uses `o200k_base` for GPT-5/o-series and `cl100k_base` for GPT-4. Anthropic uses a proprietary Claude tokenizer with different vocabulary splits. Google uses SentencePiece vocabularies. The exact same paragraph can produce 10–25% different token counts across providers.
In Byte-Pair Encoding (BPE), whitespace is merged with adjacent characters into specific tokens. Multi-space indentation in programming languages (e.g. 4 spaces vs 1 tab) can consume separate tokens if not aligned with common vocabulary merges.
Use the provider's official tokenizer library (e.g. `tiktoken` for OpenAI models, `@anthropic-ai/tokenizer` for Claude) locally in your Node.js or Python backend before making requests.