Why Hindi costs three times more tokens than English
The same sentence can cost far more tokens in Hindi or Tamil than in English, because most tokenizers were built by counting mostly-English text.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The same sentence can cost two to four times more tokens in Hindi than in English.
Think about paying for a taxi by the number of turns it makes, not by distance. Two routes cover the same ground. One is a straight road. The other zigzags through narrow lanes. The zigzag route costs more, for the identical trip.
That is what happens to non-English languages inside many AI systems. Same meaning, same length of idea, a higher bill. The reason is how the text gets chopped into pieces before the model ever sees it.
Why it exists
As tokenization already covered, a model never reads raw letters. It reads tokens instead — chunks of text pulled from a fixed vocabulary. Those chunks were learned by counting which pieces appear together most often in the training text.
Most widely-used tokenizers were built by counting mostly-English web text. Common English words earned their own single token. Devanagari, Tamil and other non-Latin scripts saw far less training data in that count. They earned far fewer whole-word tokens. Words in those scripts break into more, smaller pieces instead.
More pieces means more tokens for the identical sentence. APIs charge per token. Output speed is roughly per-token. Context windows are measured in tokens too. A less efficient tokenizer quietly costs more, runs slower, and fits less, for the same meaning.
How it works
"The weather is nice today." --> 8 tokens (English-trained tokenizer)
same sentence, Hindi --> 29 tokens (same tokenizer, unfair split)
same sentence, Tamil --> 75 tokens (same tokenizer, worse still)A tokenizer trained with real, substantial multilingual data closes much of this gap. One trained mostly on English does not, no matter how good the model built on top of it is.
Where you have already seen it
- A Hindi or Kannada prompt hitting a length limit sooner than an equivalent English prompt would.
- Slower replies for non-English conversations, since output speed tracks token count.
- Higher API bills for the same conversation, conducted in a different language.
Remember this
- Token count for the same meaning varies sharply by language, purely from how the vocabulary was trained.
- This is not a law of nature — it reflects what text the tokenizer happened to be trained on.
- A multilingual-trained tokenizer narrows this gap; an English-trained one used for everything does not.
What to learn next
- Tokenization — the mechanism producing this gap in the first place.
- XLM-RoBERTa and multilingual encoders — a tokenizer trained across 100 languages from the start.
- Context windows — where a higher token count becomes a hard limit, not only a cost.
Developer — Code and libraries.
This measures the gap directly, comparing an English-trained tokenizer against a multilingual one, on the exact same three sentences.
Setup
pip install tokenizersMeasuring the real gap between two real tokenizers
from tokenizers import Tokenizer
sentences = {
"English": "The weather is nice today.",
"Hindi": "आज मौसम अच्छा है।",
"Tamil": "இன்று வானிலை நன்றாக உள்ளது.",
}
for name in ["roberta-base", "xlm-roberta-base"]:
tok = Tokenizer.from_pretrained(name)
print(f"== {name} ==")
for lang, text in sentences.items():
n = len(tok.encode(text).ids)
print(f" {lang:8} {n:3d} tokens ({len(text)} characters)")== roberta-base == English 8 tokens (26 characters) Hindi 29 tokens (17 characters) Tamil 75 tokens (27 characters) == xlm-roberta-base == English 8 tokens (26 characters) Hindi 7 tokens (17 characters) Tamil 9 tokens (27 characters)
Both downloads together total under 3 MB — these are tokenizer files only, no model weights.
Line by line
roberta-base is an English-trained, byte-level BPE tokenizer — the same family GPT-style models use. On English it costs 8 tokens. On Hindi, a 17-character sentence, it costs 29 tokens — more tokens than characters, because unfamiliar Devanagari bytes get split into small, individually meaningless fragments.
xlm-roberta-base was trained on 100 languages, Hindi included in real volume, and the gap almost disappears: 7 tokens for that same Hindi sentence, close to the English count.
Tamil tells a subtler story. Even the multilingual tokenizer needs 9 tokens against English's 8 — a much smaller gap than roberta-base's 75, but not zero. Tamil had real but comparatively thinner representation than Hindi inside XLM-RoBERTa's training mix, matching what XLM-RoBERTa already showed about uneven coverage even inside one "multilingual" vocabulary.
Common mistakes
Assuming "multilingual tokenizer" means "equally efficient for every language." As Tamil shows here, even a good multilingual tokenizer can leave real, measurable gaps between languages — smaller than an English-only tokenizer's gap, but not eliminated.
Estimating cost with len(text) / 4, a rule of thumb that only holds for English. Applied to Hindi or Tamil against an English-trained tokenizer, this badly underestimates the real token count, sometimes by several times over.
Comparing token counts across two different tokenizers as if they were the same currency. A model's context-window limit, and its billing, are both defined in that specific model's own tokens. Always measure with the actual tokenizer the target model uses.
Try it yourself
Add a fourth sentence in a language you can read, translated to carry the same meaning as the other three. Run it through both tokenizers and compute each one's ratio against the English token count. A ratio far above 1.0 on roberta-base and close to 1.0 on xlm-roberta-base is the pattern this lesson describes, measured on a language of your choice.
What to learn next
- XLM-RoBERTa and multilingual encoders — the tokenizer scoring far better on Hindi here, and why.
- Tokenization — the byte-pair-encoding mechanism producing every number in this lesson.
- Context windows — turning a higher token count into a real, hard limit on how much fits.
Researcher — Mathematics and papers.
Quantifying the gap: fertility
Fertility is the standard metric here, already introduced in tokenization: the mean number of tokens per word (or per whitespace-separated unit) for a given language under a fixed tokenizer. A fertility of 1.0 means, on average, one token per word — the efficient case. Higher fertility means more fragmentation, more tokens for the same content.
fertility(L) = ( total tokens produced for corpus in language L ) / ( total words in that corpus )Petrov, La Malfa, Torr & Bibi (2023) measured this systematically across parallel corpora — identical content, translated into many languages — and found the ratio between the most- and least-efficiently-tokenized language for a given tokenizer exceeded 15 for some tokenizer and language combinations. The developer demo above reproduces the same phenomenon directly, at smaller scale: roberta-base's Hindi-to-English token ratio is roughly 3.6, Tamil's roughly 9.4, on the three sentences measured.
Why fertility varies by script, not only by "resource level"
Two separate mechanisms compound here. First, training-data volume: a tokenizer's vocabulary is learned by frequency counting over its training corpus, so scripts underrepresented in that corpus earn fewer dedicated vocabulary entries, exactly as shown for roberta-base above.
Second, byte-level encoding overhead, independent of training data entirely. Byte-level BPE (Radford et al., 2019, GPT-2) operates over UTF-8 bytes, not Unicode characters. A single Devanagari or Tamil character occupies 3 bytes in UTF-8; a single ASCII character occupies 1. Before any merges are learned at all, non-Latin scripts start at up to 3 times the raw sequence length of equivalent ASCII text. Merges recover much of this only if training data density for that script was high enough to learn efficient multi-byte merges — which is exactly why xlm-roberta-base, using SentencePiece over normalized text rather than raw UTF-8 bytes, and trained with real Hindi volume, avoids most of this overhead for Hindi specifically.
Downstream consequences, concretely
| Effect | Mechanism |
|---|---|
| Higher price | API billing is per token |
| Higher latency | Autoregressive decoding time is roughly linear in output token count |
| Smaller usable context | The context window is measured in tokens, not characters or words |
| Worse quality | More tokens per idea means less semantic content per attention operation, straining the model's effective context budget |
Key references
- Petrov, A., La Malfa, E., Torr, P. & Bibi, A. (2023). Language Model Tokenizers Introduce Unfairness Between Languages. arXiv:2305.15425
- Radford, A. et al. (2019). Language Models are Unsupervised Multitask Learners. — GPT-2, byte-level BPE.
- Conneau, A. et al. (2020). Unsupervised Cross-lingual Representation Learning at Scale. arXiv:1911.02116 — XLM-RoBERTa's SentencePiece vocabulary training.
- Ahia, O. et al. (2023). Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models. arXiv:2305.13707 — quantifies the direct dollar-cost impact of this gap for real API pricing.
Current state and open problems
Newer frontier model tokenizers have widened their vocabularies substantially, in part specifically to improve multilingual fertility — a wider vocabulary can afford more non-English whole-word entries without sacrificing English efficiency. This measurably helps, without closing the gap for lower-resource scripts entirely, since vocabulary allocation still tracks training-data volume even inside a larger budget.
Byte-level and tokenizer-free architectures (ByT5, MEGABYTE, and the byte-latent line of research already discussed in tokenization) remove the fertility gap by construction, since every language pays the same per-byte cost. None has yet displaced subword tokenization at frontier model scale, so the fairness gap measured in this lesson remains a live, current cost for anyone building or using non-English AI products today.
What to learn next
- Tokenization — the full mechanism, including byte-pair encoding and the tokenizer-free research direction mentioned above.
- XLM-RoBERTa and multilingual encoders — a concrete example of training-data volume shaping fertility, measured earlier in this section.
- Context windows — the hard limit this fertility gap runs directly into.