Multilingual sentence embeddings
Multilingual sentence embeddings turn a whole sentence, in any of many languages, into one list of numbers that sits near the same-meaning sentence in every other language.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A multilingual sentence embedding turns a whole sentence into a list of numbers. The same idea lands in nearly the same spot, no matter which language wrote it.
Picture a large signboard at a railway station, written in Hindi, English and Tamil. All three lines point to the same platform. The words differ completely. The direction they give you does not.
Multilingual embeddings work the same way. "How do I reset my password?" in English lands almost on top of its Hindi translation. Not one word matches, yet the location does.
Why it exists
A single-language embedding, like the ones in Embeddings, only ever compares sentences written in the same language. Compare an English sentence with a Hindi one and the numbers mean nothing next to each other.
Products need to search, match and cluster text across languages. A support search bar should work whether the user types in English or Hindi. A duplicate-question detector needs to work the same way, across a multilingual forum. That requires one shared number-space, not one space per language.
Multilingual sentence embedding models are trained specifically for this. Translated sentence pairs land close together. Unrelated sentences land far apart, regardless of which language either one is written in.
How it works
"How do I reset my password?" (English) --\
>-- same neighbourhood
"मैं अपना पासवर्ड कैसे रीसेट करूं?" (Hindi) --/ in number-space
"The weather is nice today." (English) ----> far away, different neighbourhoodDistance in this space is measured with cosine similarity, a number from -1 to 1. Close to 1 means "point in the same direction." Close to 0 means "unrelated." Negative means "opposite."
Where you have already seen it
- "Related questions" sections on multilingual help sites, surfacing a match written in a different language than you typed.
- Duplicate detection on multilingual community forums.
- Search bars that return the right result whether you type in English or your own language.
Remember this
- A multilingual embedding model turns a sentence into numbers. Translations of the same sentence land near each other.
- Cosine similarity, not exact word matching, is how "near" gets measured.
- The model was trained specifically to align translations. A generic multilingual model, without that training, will not do this reliably.
What to learn next
- Question in Hindi, documents in English — using this exact idea to build search.
- Embeddings — the single-language foundation this extends.
- Train in English, run in Hindi — the same shared space, used for classification instead of comparison.
Developer — Code and libraries.
Setup
pip install sentence-transformersThis downloads sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2, about 460 MB, cached after the first run. It supports over 50 languages.
Comparing a sentence across languages
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2")
sentences = [
"How do I reset my password?", # English
"मैं अपना पासवर्ड कैसे रीसेट करूं?", # Hindi, same meaning
"The weather is nice today.", # English, unrelated meaning
]
labels = ["en_password", "hi_password", "en_weather"]
emb = model.encode(sentences, normalize_embeddings=True)
sim = emb @ emb.T
print("shape of one embedding:", emb.shape)
print()
header = " " + " ".join(f"{l:>12}" for l in labels)
print(header)
for label, row in zip(labels, sim):
print(f"{label:14}" + " ".join(f"{v:12.3f}" for v in row))shape of one embedding: (3, 384)
en_password hi_password en_weather
en_password 1.000 0.870 -0.070
hi_password 0.870 1.000 -0.024
en_weather -0.070 -0.024 1.000Line by line
Every sentence becomes a 384-number vector, regardless of length or language. emb.shape confirms that: 3 sentences, 384 numbers each.
emb @ emb.T computes every pairwise cosine similarity in one matrix multiply, because the vectors were already normalised to length 1. Dot product of two unit vectors equals cosine similarity directly.
The English and Hindi password questions score 0.870. Same meaning, different script, different word order, different vocabulary — the model still places them next to each other. The unrelated weather sentence scores near zero against both, exactly as it should.
Common mistakes
Comparing raw embeddings without normalising them. Cosine similarity via dot product only works correctly on unit-length vectors. Skip normalize_embeddings=True and your "similarity" numbers stop meaning what you think they mean.
Assuming any multilingual model does this. XLM-RoBERTa on its own, without extra training for sentence similarity, produces embeddings that cluster poorly by meaning. This model was specifically fine-tuned on translation pairs to align languages — that fine-tuning is the entire point.
Comparing sentence embeddings across model versions or checkpoints. Two different model versions place sentences in unrelated coordinate systems. Never mix embeddings from different models in one comparison.
Try it yourself
Add a third language you can read — French, Marathi, Spanish — with a translation of the password question. Check whether its similarity score to the English original stays near 0.85, or drops. A lower score is a real signal about how well this particular language is covered.
What to learn next
- Question in Hindi, documents in English — ranking multiple documents by this same similarity score.
- Hugging Face — the ecosystem these models are usually downloaded and run through.
- XLM-RoBERTa and multilingual encoders — the base encoder this model builds on.
Researcher — Mathematics and papers.
Training objective
paraphrase-multilingual-MiniLM-L12-v2 is trained with multilingual knowledge distillation (Reimers & Gurevych, 2020): a strong monolingual teacher model, already good at English sentence similarity, guides a multilingual student.
loss = || teacher(x_en) - student(x_en) ||^2 + || teacher(x_en) - student(x_L) ||^2x_enis an English sentence,x_Lits translation into target languageL.teacheris a fixed, pretrained English sentence-embedding model.studentis the multilingual model being trained, initialised from a multilingual encoder such as XLM-RoBERTa or multilingual MiniLM.
The second term is the important one. It forces the student to place x_L's embedding wherever the teacher already places x_en's embedding, for every language L in the parallel training data. This is what direct masked-language-model pretraining alone does not guarantee — MLM pretraining aligns languages loosely, at the level of shared syntax and word co-occurrence, not at the level of matching whole-sentence meaning precisely enough for retrieval.
Why this differs from plain XLM-RoBERTa embeddings
Pooling the raw hidden states of a masked-language-model encoder, with no further training, produces embeddings dominated by frequency and superficial lexical overlap rather than semantic similarity — a well-documented weakness of BERT-family embeddings (Reimers & Gurevych, 2019, on the monolingual case).
Distillation against a similarity-trained teacher corrects this by training directly on the metric the model will be judged on: cosine similarity between semantically equivalent pairs, across language boundaries.
Complexity
Encoding n sentences costs one forward pass per sentence through a 12-layer transformer, O(n * L * d^2) for sequence length L and hidden width d, the same cost profile as any transformer encoder pass — see attention for the per-layer breakdown. Comparing n sentences pairwise afterwards costs O(n^2 * d), negligible next to encoding for typical n.
Key references
- Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv:1908.10084
- Reimers, N. & Gurevych, I. (2020). Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation. arXiv:2004.09813
- Feng, F. et al. (2022). Language-agnostic BERT Sentence Embedding (LaBSE). arXiv:2007.01852 — an alternative trained with a translation-ranking objective across 109 languages.
- Wang, L. et al. (2024). Multilingual E5 Text Embeddings. arXiv:2402.05672 — a current widely-used retrieval-oriented multilingual embedding family.
Current state and open problems
Coverage across the model's supported languages is uneven, tracking the volume of parallel (translated) training data available per language pair — plentiful for Hindi, thin for many other Indian languages such as Santali or Bodo.
Retrieval-specific multilingual embeddings (E5, BGE-M3) now dominate production use over pure similarity models like this one, trained with contrastive objectives directly on query-document pairs rather than paraphrase pairs. The distinction matters: a model tuned for "these two sentences mean the same thing" is not automatically well tuned for "this short query should retrieve this long document," which is the actual shape of the task in cross-lingual retrieval.
What to learn next
- Question in Hindi, documents in English — the retrieval-shaped version of this problem.
- XLM-RoBERTa and multilingual encoders — the base encoder family these models are built from.
- Attention — the mechanism computing the representations this model pools into a single vector.