Train in English, run in Hindi
Cross-lingual transfer trains a classifier on English text and gets it working on Hindi text it never saw, because both languages share one embedding space.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Cross-lingual transfer trains a model on labelled English text. It then runs on Hindi text it has never seen labelled — and it still works.
Think about learning to recognise a raga by ear. Once you know the feeling of Yaman, you recognise it on a sitar, a violin, or a flute. The instrument changed. The pattern underneath did not.
Cross-lingual transfer counts on the same idea. A model trained to spot "this is a complaint" in English learns a pattern of meaning. It does not memorise English words.
Why it exists
Labelled data is expensive. Someone has to read thousands of examples and mark each one correctly.
Companies usually have plenty of labelled English data and almost none in Hindi, Tamil, Marathi or Bengali. Building a fresh labelled dataset for every language does not scale.
XLM-RoBERTa and similar models solve this by placing every language's sentences into the same numerical space during pretraining. A sentence in Hindi and its English translation land close together, without either language being explicitly aligned during training.
That shared space is what makes transfer possible. Teach the model a task using English coordinates. Hindi sentences already sit near the right coordinates, because pretraining put them there.
How it works
Training: English sentence --> shared space --> "complaint" or "praise"
^
| (same space, no extra training)
|
At use time: Hindi sentence --> shared space --> model still decides correctlyThe classifier sitting on top never sees Hindi during training. It only ever learns "which region of this shared space means complaint." Hindi sentences about complaints happen to land in that same region.
Where you have already seen it
- Content moderation trained mostly on English data, catching abuse in other languages too.
- Spam filters built on English complaints, still flagging Hindi and Hinglish spam.
- Sentiment tools launched for English first, working passably on other languages from day one, before any local labelling happens.
Remember this
- Cross-lingual transfer trains on one language and evaluates on another, using a shared multilingual embedding space.
- It works because the pretrained encoder already placed matching-meaning sentences near each other, across languages.
- It is never as strong as training directly on labelled target-language data. It is what you use before that data exists.
What to learn next
- Multilingual sentence embeddings — a closer look at the shared space itself.
- XLM-RoBERTa and multilingual encoders — the model that builds this shared space.
- Text classification — the single-language version of the task being transferred here.
Developer — Code and libraries.
This demo trains a classifier on eight English sentences and tests it on four Hindi sentences it has never seen. No Hindi ever touches training.
Setup
pip install sentence-transformers scikit-learnThe first run downloads sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2, about 460 MB, cached after that.
Train on English, test on Hindi
from sentence_transformers import SentenceTransformer
from sklearn.linear_model import LogisticRegression
model = SentenceTransformer("sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2")
# Training data: English only. Label 1 = question, 0 = statement.
train_texts = [
"What time does the store close?",
"How much does this cost?",
"Where is the nearest station?",
"Can you help me with this form?",
"The store closes at nine tonight.",
"This costs five hundred rupees.",
"The station is two blocks away.",
"She helped me fill the form.",
]
train_labels = [1, 1, 1, 1, 0, 0, 0, 0]
# Test data: Hindi only. Never seen during training.
test_texts_hi = [
"दुकान कितने बजे बंद होती है?", # What time does the shop close?
"यह कितने का है?", # How much is this?
"दुकान रात नौ बजे बंद होती है।", # The shop closes at nine at night.
"यह पाँच सौ रुपये का है।", # This costs five hundred rupees.
]
test_labels_hi = [1, 1, 0, 0]
X_train = model.encode(train_texts, normalize_embeddings=True)
X_test = model.encode(test_texts_hi, normalize_embeddings=True)
clf = LogisticRegression()
clf.fit(X_train, train_labels)
preds = clf.predict(X_test)
correct = 0
for text, true, pred in zip(test_texts_hi, test_labels_hi, preds):
ok = "correct" if pred == true else "WRONG"
correct += pred == true
print(f"{pred} (true={true}, {ok}) {text}")
print(f"\n{correct}/{len(test_labels_hi)} correct, trained on English only")1 (true=1, correct) दुकान कितने बजे बंद होती है? 1 (true=1, correct) यह कितने का है? 0 (true=0, correct) दुकान रात नौ बजे बंद होती है। 0 (true=0, correct) यह पाँच सौ रुपये का है। 4/4 correct, trained on English only
Line by line
The classifier is ordinary logistic regression. It has no idea what language it is looking at. It only sees 384-number vectors and learns a boundary between them. All the multilingual understanding happened earlier, inside model.encode.
normalize_embeddings=True scales every vector to length 1. This makes cosine similarity and a linear classifier boundary behave consistently across sentences of different lengths.
Four Hindi sentences is a toy demo, not a benchmark. This shows the mechanism working, not a measured accuracy number you can quote. Real evaluation needs hundreds of examples per language and a held-out test set.
Common mistakes
Assuming transfer accuracy matches the source language's accuracy. It never does. Expect a real drop when moving from English to a language the encoder saw less of during pretraining.
Skipping evaluation in the target language entirely. "It should work, the model is multilingual" is a guess, not a measurement. Always score on real target-language examples before shipping.
Forgetting that transfer quality depends on the encoder, not the classifier. A better multilingual encoder improves transfer more than a fancier classifier on top of a weak one does.
Try it yourself
Add a fifth pair of labels — say, "polite request" versus "direct command" — with English training examples and a couple of Hindi test sentences. Rerun and see whether the same shared space carries this new distinction across languages too.
What to learn next
- Multilingual sentence embeddings — what
model.encodeis actually producing here. - Question in Hindi, documents in English — the same shared space used for search instead of classification.
- Text classification — the general version of the classifier trained above.
Researcher — Mathematics and papers.
The transfer setup
Given a labelled source-language dataset D_s = {(x_i, y_i)} and an unlabelled or sparsely-labelled target language, cross-lingual transfer trains a task head f on phi(x_i) for x_i drawn from D_s, where phi is a frozen or fine-tuned multilingual encoder.
f* = argmin over f of sum over (x_i, y_i) in D_s of loss( f(phi(x_i)), y_i )At inference, f* is applied to phi(x_t) for target-language input x_t, with no target-language examples in the training sum. Transfer succeeds to the extent that phi maps semantically equivalent sentences in source and target languages to nearby points, so the same decision boundary in phi's output space is meaningful for both.
Zero-shot versus few-shot transfer
Zero-shot: f never sees target-language examples, as in the demo above. This is the cheapest setting and the weakest ceiling.
Few-shot: a small number of labelled target examples are added to fine-tuning, either mixed into D_s or used for a lightweight adapter update. Even 50 to 100 target examples typically close a substantial part of the zero-shot gap, since they correct systematic encoder biases the linear head cannot see from source-language data alone.
Why the gap exists
XNLI (Conneau et al., 2018) is the standard benchmark for measuring this gap directly: the same natural language inference task, translated into 15 languages, evaluated after training only on English.
The consistent finding across XLM-R and its predecessors: zero-shot accuracy in the target language trails supervised English accuracy by a margin that grows as the target language's pretraining data shrinks and as its script or grammar diverges further from the source language. High-resource, typologically close pairs (English to German) transfer far better than low-resource, distant pairs (English to Swahili or Urdu).
Key references
- Conneau, A. et al. (2018). XNLI: Evaluating Cross-lingual Sentence Representations. arXiv:1809.05053
- Conneau, A. et al. (2020). Unsupervised Cross-lingual Representation Learning at Scale. arXiv:1911.02116
- Hu, J. et al. (2020). XTREME: A Massively Multilingual Multi-task Benchmark. arXiv:2003.11080 — broadens evaluation beyond XNLI to nine task families across 40 languages.
- Lauscher, A. et al. (2020). From Zero to Hero: On the Limitations of Zero-Shot Cross-Lingual Transfer with Transformers. arXiv:2005.04903 — quantifies how few target examples are needed to close most of the gap.
Current state and open problems
Transfer quality correlates with pretraining data volume for the target language, but not perfectly. Lauscher et al. (2020) show typological distance from the source language, script differences, and morphological complexity all contribute independently.
Adapter-based methods (MAD-X, Pfeiffer et al., 2020) now let a single base encoder host lightweight, swappable per-language modules, narrowing the gap for genuinely low-resource languages without full model fine-tuning per language.
The open question is not whether transfer works. It reliably does, in the sense demonstrated above. The open question is how far it can be pushed before target-language data becomes unavoidable — and that answer differs by task, language pair, and how much the task depends on fine-grained syntax versus coarse topical meaning.
What to learn next
- XLM-RoBERTa and multilingual encoders — the encoder whose shared space this section formalises.
- Multilingual sentence embeddings — the sentence-level version of the same representation.
- Attention — the mechanism computing the representations transfer depends on.