Multilingual and Indic NLP

Why Hindi costs three times more tokens than English

The same sentence can cost far more tokens in Hindi or Tamil than in English, because most tokenizers were built by counting mostly-English text.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The same sentence can cost two to four times more tokens in Hindi than in English.

Think about paying for a taxi by the number of turns it makes, not by distance. Two routes cover the same ground. One is a straight road. The other zigzags through narrow lanes. The zigzag route costs more, for the identical trip.

That is what happens to non-English languages inside many AI systems. Same meaning, same length of idea, a higher bill. The reason is how the text gets chopped into pieces before the model ever sees it.

Why it exists

As tokenization already covered, a model never reads raw letters. It reads tokens instead — chunks of text pulled from a fixed vocabulary. Those chunks were learned by counting which pieces appear together most often in the training text.

Most widely-used tokenizers were built by counting mostly-English web text. Common English words earned their own single token. Devanagari, Tamil and other non-Latin scripts saw far less training data in that count. They earned far fewer whole-word tokens. Words in those scripts break into more, smaller pieces instead.

More pieces means more tokens for the identical sentence. APIs charge per token. Output speed is roughly per-token. Context windows are measured in tokens too. A less efficient tokenizer quietly costs more, runs slower, and fits less, for the same meaning.

How it works

  "The weather is nice today."   -->  8 tokens   (English-trained tokenizer)
  same sentence, Hindi           -->  29 tokens  (same tokenizer, unfair split)
  same sentence, Tamil           -->  75 tokens  (same tokenizer, worse still)

A tokenizer trained with real, substantial multilingual data closes much of this gap. One trained mostly on English does not, no matter how good the model built on top of it is.

Where you have already seen it

  • A Hindi or Kannada prompt hitting a length limit sooner than an equivalent English prompt would.
  • Slower replies for non-English conversations, since output speed tracks token count.
  • Higher API bills for the same conversation, conducted in a different language.

Remember this

  • Token count for the same meaning varies sharply by language, purely from how the vocabulary was trained.
  • This is not a law of nature — it reflects what text the tokenizer happened to be trained on.
  • A multilingual-trained tokenizer narrows this gap; an English-trained one used for everything does not.

What to learn next

Developer — Code and libraries.

This measures the gap directly, comparing an English-trained tokenizer against a multilingual one, on the exact same three sentences.

Setup

bash
pip install tokenizers

Measuring the real gap between two real tokenizers

token_cost_compare.py
from tokenizers import Tokenizer

sentences = {
    "English": "The weather is nice today.",
    "Hindi":   "आज मौसम अच्छा है।",
    "Tamil":   "இன்று வானிலை நன்றாக உள்ளது.",
}

for name in ["roberta-base", "xlm-roberta-base"]:
    tok = Tokenizer.from_pretrained(name)
    print(f"== {name} ==")
    for lang, text in sentences.items():
        n = len(tok.encode(text).ids)
        print(f"  {lang:8} {n:3d} tokens  ({len(text)} characters)")
Output
== roberta-base ==
  English    8 tokens  (26 characters)
  Hindi     29 tokens  (17 characters)
  Tamil     75 tokens  (27 characters)
== xlm-roberta-base ==
  English    8 tokens  (26 characters)
  Hindi      7 tokens  (17 characters)
  Tamil      9 tokens  (27 characters)

Both downloads together total under 3 MB — these are tokenizer files only, no model weights.

Line by line

roberta-base is an English-trained, byte-level BPE tokenizer — the same family GPT-style models use. On English it costs 8 tokens. On Hindi, a 17-character sentence, it costs 29 tokens — more tokens than characters, because unfamiliar Devanagari bytes get split into small, individually meaningless fragments.

xlm-roberta-base was trained on 100 languages, Hindi included in real volume, and the gap almost disappears: 7 tokens for that same Hindi sentence, close to the English count.

Tamil tells a subtler story. Even the multilingual tokenizer needs 9 tokens against English's 8 — a much smaller gap than roberta-base's 75, but not zero. Tamil had real but comparatively thinner representation than Hindi inside XLM-RoBERTa's training mix, matching what XLM-RoBERTa already showed about uneven coverage even inside one "multilingual" vocabulary.

Common mistakes

Assuming "multilingual tokenizer" means "equally efficient for every language." As Tamil shows here, even a good multilingual tokenizer can leave real, measurable gaps between languages — smaller than an English-only tokenizer's gap, but not eliminated.

Estimating cost with len(text) / 4, a rule of thumb that only holds for English. Applied to Hindi or Tamil against an English-trained tokenizer, this badly underestimates the real token count, sometimes by several times over.

Comparing token counts across two different tokenizers as if they were the same currency. A model's context-window limit, and its billing, are both defined in that specific model's own tokens. Always measure with the actual tokenizer the target model uses.

Try it yourself

Add a fourth sentence in a language you can read, translated to carry the same meaning as the other three. Run it through both tokenizers and compute each one's ratio against the English token count. A ratio far above 1.0 on roberta-base and close to 1.0 on xlm-roberta-base is the pattern this lesson describes, measured on a language of your choice.

What to learn next

Researcher — Mathematics and papers.

Quantifying the gap: fertility

Fertility is the standard metric here, already introduced in tokenization: the mean number of tokens per word (or per whitespace-separated unit) for a given language under a fixed tokenizer. A fertility of 1.0 means, on average, one token per word — the efficient case. Higher fertility means more fragmentation, more tokens for the same content.

text
fertility(L) = ( total tokens produced for corpus in language L ) / ( total words in that corpus )

Petrov, La Malfa, Torr & Bibi (2023) measured this systematically across parallel corpora — identical content, translated into many languages — and found the ratio between the most- and least-efficiently-tokenized language for a given tokenizer exceeded 15 for some tokenizer and language combinations. The developer demo above reproduces the same phenomenon directly, at smaller scale: roberta-base's Hindi-to-English token ratio is roughly 3.6, Tamil's roughly 9.4, on the three sentences measured.

Why fertility varies by script, not only by "resource level"

Two separate mechanisms compound here. First, training-data volume: a tokenizer's vocabulary is learned by frequency counting over its training corpus, so scripts underrepresented in that corpus earn fewer dedicated vocabulary entries, exactly as shown for roberta-base above.

Second, byte-level encoding overhead, independent of training data entirely. Byte-level BPE (Radford et al., 2019, GPT-2) operates over UTF-8 bytes, not Unicode characters. A single Devanagari or Tamil character occupies 3 bytes in UTF-8; a single ASCII character occupies 1. Before any merges are learned at all, non-Latin scripts start at up to 3 times the raw sequence length of equivalent ASCII text. Merges recover much of this only if training data density for that script was high enough to learn efficient multi-byte merges — which is exactly why xlm-roberta-base, using SentencePiece over normalized text rather than raw UTF-8 bytes, and trained with real Hindi volume, avoids most of this overhead for Hindi specifically.

Downstream consequences, concretely

EffectMechanism
Higher priceAPI billing is per token
Higher latencyAutoregressive decoding time is roughly linear in output token count
Smaller usable contextThe context window is measured in tokens, not characters or words
Worse qualityMore tokens per idea means less semantic content per attention operation, straining the model's effective context budget

Key references

  • Petrov, A., La Malfa, E., Torr, P. & Bibi, A. (2023). Language Model Tokenizers Introduce Unfairness Between Languages. arXiv:2305.15425
  • Radford, A. et al. (2019). Language Models are Unsupervised Multitask Learners. — GPT-2, byte-level BPE.
  • Conneau, A. et al. (2020). Unsupervised Cross-lingual Representation Learning at Scale. arXiv:1911.02116 — XLM-RoBERTa's SentencePiece vocabulary training.
  • Ahia, O. et al. (2023). Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models. arXiv:2305.13707 — quantifies the direct dollar-cost impact of this gap for real API pricing.

Current state and open problems

Newer frontier model tokenizers have widened their vocabularies substantially, in part specifically to improve multilingual fertility — a wider vocabulary can afford more non-English whole-word entries without sacrificing English efficiency. This measurably helps, without closing the gap for lower-resource scripts entirely, since vocabulary allocation still tracks training-data volume even inside a larger budget.

Byte-level and tokenizer-free architectures (ByT5, MEGABYTE, and the byte-latent line of research already discussed in tokenization) remove the fertility gap by construction, since every language pays the same per-byte cost. None has yet displaced subword tokenization at frontier model scale, so the fairness gap measured in this lesson remains a live, current cost for anyone building or using non-English AI products today.

What to learn next

  • Tokenization — the full mechanism, including byte-pair encoding and the tokenizer-free research direction mentioned above.
  • XLM-RoBERTa and multilingual encoders — a concrete example of training-data volume shaping fertility, measured earlier in this section.
  • Context windows — the hard limit this fertility gap runs directly into.