XLM-RoBERTa and multilingual encoders
XLM-RoBERTa is one BERT-style model trained on 100 languages together, so Hindi, Tamil and English share the same understanding engine.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
XLM-RoBERTa is one shared model trained on 100 languages at once. It reads Hindi, Tamil and English with the same understanding engine.
Picture a classroom in Bengaluru. Students speak Hindi at home, Tamil with grandparents, English at school. One teacher explains photosynthesis. Every student understands, regardless of which language carried the explanation.
XLM-RoBERTa is that teacher. Except it has 100 languages, and the lesson is meaning itself.
Why it exists
Before this, teams trained one BERT-style model per language. That worked well for English, where the internet has endless text to learn from.
It worked badly for Odia, Assamese or Sindhi. There is not enough text in any one of those languages alone to train a strong model from scratch.
XLM-RoBERTa's answer: stop training one model per language. Train one model on all of them together, using web text (Common Crawl) in 100 languages at once. A rare language borrows patterns learned mostly from richer ones — grammar, word order, the general shape of sentences.
This does not make every language equally strong. It makes weak languages less weak than training them alone would.
How it works
Hindi sentence -----\
Tamil sentence ------> ONE shared encoder ----> same 768-number space
English sentence -----/ (shared vocabulary) for every languageOne vocabulary, built by counting subword pieces across all 100 languages together (the tokenizer from tokenization, trained multilingually). One transformer encoder, the same architecture as BERT. One training goal: guess the word hidden behind a mask, in whichever language the sentence happens to be in.
Nothing in the architecture is told "this sentence is Hindi." The model works it out from the subword pieces it sees, the same way it works out everything else.
Where you have already seen it
- Search engines that return a relevant Hindi page for an English query, and vice versa.
- Content moderation that catches harmful text in many languages using one flagging model, not a hundred separate ones.
- IndicBERT and similar Indian-language models, which build on the same "train together" idea, tuned specifically for Indian scripts.
Remember this
- XLM-RoBERTa is one model, one shared vocabulary, trained on 100 languages together — not 100 separate models.
- Low-resource languages benefit the most, since they borrow structure learned from data-rich languages.
- It never sees an explicit "language" label. It infers language from the text itself.
What to learn next
- Cross-lingual transfer — training on English labels, running the model on Hindi text.
- BERT — the single-language ancestor this architecture is built from.
- Tokenization — how the shared vocabulary gets built in the first place.
Developer — Code and libraries.
You do not need the full model to see the interesting part. The tokenizer alone shows how one vocabulary copes with three unrelated scripts.
Setup
pip install transformersTokenizing three languages with one shared vocabulary
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("xlm-roberta-base")
sentences = {
"English": "The weather is nice today.",
"Hindi": "आज मौसम अच्छा है।",
"Tamil": "இன்று வானிலை நன்றாக உள்ளது.",
}
for lang, text in sentences.items():
ids = tok(text)["input_ids"]
pieces = tok.convert_ids_to_tokens(ids)
print(f"{lang:8} {len(ids)} tokens {pieces}")English 8 tokens ['<s>', '▁The', '▁weather', '▁is', '▁nice', '▁today', '.', '</s>'] Hindi 7 tokens ['<s>', '▁आज', '▁मौसम', '▁अच्छा', '▁है', '।', '</s>'] Tamil 9 tokens ['<s>', '▁இன்று', '▁வா', 'னி', 'லை', '▁நன்றாக', '▁உள்ளது', '.', '</s>']
This command downloads only the tokenizer files — about 14 MB on disk, not the 1.1 GB full model. Loading the model itself for real predictions is a separate, much larger download.
Line by line
▁ is not an underscore. It is SentencePiece's marker for "a space came before this piece." ▁today means the word "today" started with a space in the original text.
Hindi came out efficient — 7 tokens, close to English's 8. Hindi has serious representation in the training data, so common words like "मौसम" (weather) survive as single pieces.
Tamil broke apart — "வானிலை" (weather) split into three pieces: ▁வா, னி, லை. Tamil script sees less training data than Hindi inside this shared vocabulary, so fewer whole Tamil words earned their own slot.
That gap is the honest state of things. A shared multilingual vocabulary does not treat every one of its 100 languages equally, even though it was trained on all of them.
Common mistakes
Assuming "multilingual" means "equally good at every language." It means "trained on many languages together," not "uniformly skilled." Coverage tracks how much text of that language existed in the training crawl.
Expecting multilingual-base to beat a good monolingual model in a well-resourced language. For Hindi or English specifically, a dedicated monolingual model trained with the same compute usually wins. XLM-RoBERTa's advantage shows up hardest for languages with little text of their own.
Forgetting the vocabulary is huge. 250,002 subword entries, because 100 languages need far more pieces than one language does. That inflates the embedding layer's parameter count before a single transformer block runs.
Try it yourself
Add a fourth sentence in a language with less web presence — Nepali or Sindhi are good tests. Compare its token count per character against Hindi's. A higher ratio means the shared vocabulary gave that script fewer whole-word pieces to work with.
What to learn next
- Multilingual sentence embeddings — turning this shared space into comparable sentence vectors.
- Why Hindi costs three times more tokens than English — the same coverage gap, measured in cost.
- Attention — the mechanism running inside every layer of this encoder.
Researcher — Mathematics and papers.
Architecture and objective
XLM-RoBERTa (Conneau et al., 2020) follows the RoBERTa recipe (Liu et al., 2019): a BERT-style bidirectional transformer encoder, trained with masked language modelling only — no next-sentence-prediction objective, dynamic masking each epoch.
loss = -sum over masked positions i of log P(x_i | x_{context})x_iis the true token at a masked position.x_{context}is every unmasked token in the sequence, attended to bidirectionally.Pis produced by a softmax over the shared 250,002-entry SentencePiece vocabulary (Kudo & Richardson, 2018), trained jointly across all languages in the corpus.
Two released sizes: base (12 layers, 768 hidden, ~270M parameters) and large (24 layers, 1024 hidden, ~550M parameters). Both figures include the oversized embedding table the shared vocabulary requires.
Training data and language balancing
Pretraining used CC-100 (Conneau et al., 2020), text filtered from Common Crawl into 100 languages by a language-identification classifier, totalling roughly 2.5 TB.
Raw language frequencies in web text are wildly unequal. English dominates by orders of magnitude over most Indian languages. Sampling proportional to raw frequency would starve every low-resource language of gradient signal.
The fix is temperature sampling. A language with natural probability p_l is instead sampled with probability proportional to p_l^alpha, where alpha = 0.3. This exponent flattens the distribution, upweighting rare languages and downweighting English relative to its raw share, without fully equalizing them.
The curse of multilinguality
The paper's central empirical finding: for a fixed model capacity, adding more languages to pretraining improves low-resource language performance up to a point, then degrades performance across the board, including on high-resource languages.
This is a capacity allocation problem. A fixed parameter budget must represent every language's grammar and vocabulary. Beyond some number of languages, per-language capacity falls faster than cross-lingual transfer compensates for it. Scaling model size upward (base to large) delays, but does not remove, this ceiling.
Key references
- Conneau, A. et al. (2020). Unsupervised Cross-lingual Representation Learning at Scale. arXiv:1911.02116 — XLM-RoBERTa.
- Liu, Y. et al. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692
- Devlin, J. et al. (2019). BERT. arXiv:1810.04805
- Conneau, A. et al. (2018). XNLI: Evaluating Cross-lingual Sentence Representations. arXiv:1809.05053 — the standard cross-lingual evaluation benchmark used to measure this family of models.
Current state and open problems
Later work attacks the curse of multilinguality directly rather than accepting it. mDeBERTa-v3 (He et al., 2021) improves the pretraining objective itself (replaced-token detection instead of masking) and gets more out of a comparable parameter budget.
A separate line drops the "one model, every language" premise for Indian languages specifically. IndicBERT and the AI4Bharat family (see the AI4Bharat and IndicNLP stack) restrict the vocabulary to Indic scripts plus English, trading global coverage for deeper per-language capacity where it is needed most.
Whether to pool languages or specialise remains an open, workload-dependent decision. Pooling wins when a language has near-zero labelled data of its own. Specialising wins once a language, or a language family, has enough text and compute to support a dedicated model.
What to learn next
- Cross-lingual transfer — the practical payoff of the shared representation space.
- The AI4Bharat and IndicNLP stack — the specialised alternative to a 100-language model.
- Attention — the mechanism computed inside every one of these 12 or 24 layers.