Due to Byte-Pair Encoding (BPE) character splits, Portuguese text generates 1.16× more tokens than equivalent English text. 1,000 Portuguese words consume ~1,546 tokens in GPT-4o mini.
Quick answer
For 1,000 Portuguese words, GPT-4o mini is modeled at 1,546 tokens—1.16× the English baseline. That is approximately $0.000232 as input, before any cache discount.
Method & trust
The language page applies the published language multiplier to a 1,333-token English baseline, then applies the model's input/output rates. Tokenizers differ, so benchmark representative text before production budgeting.
Relative to English baseline (1.0×).
Price to send 1,000 words of Portuguese text into context.
Discounted price for repeated Portuguese system context.
| Word Count Scale | Portuguese Tokens | Portuguese Input Cost | English Equivalent Cost | Tokenization Penalty |
|---|---|---|---|---|
| 1,000 words (Short Article) | 1,546 | $0.000232 | $0.0002 | +$0.000032 |
| 10,000 words (Whitepaper / Report) | 15,463 | $0.002319 | $0.002 | +$0.00032 |
| 50,000 words (Book / Corpus) | 77,314 | $0.0116 | $0.009998 | +$0.0016 |
| 100,000 words (Enterprise Repository) | 154,628 | $0.0232 | $0.02 | +$0.003199 |
Marginally higher token count than English due to morphological endings.
Standard pricing models apply with near-English token parity.
Most LLM tokenizers are primarily trained on English-heavy web datasets. Non-Latin characters in Portuguese (Latin) split across multiple Byte-Pair Encoding (BPE) sub-word or multi-byte UTF-8 tokens, requiring approximately 1.16× more tokens to encode the exact same semantic meaning as English.
Standard pricing models apply with near-English token parity. In addition, enabling prompt caching on static Portuguese instructions or documentation saves 75–90% on input token rates.
Yes. GPT-4o mini has strong multilingual comprehension and generation capabilities in Portuguese (Português). The difference is purely computational and financial due to sub-word token splits.