The Embedding Family

Embedding size and Matryoshka truncation

A bigger embedding vector usually captures more, but Matryoshka training builds vectors that can be safely cut shorter on demand, trading some accuracy for real savings in storage and speed.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A bigger embedding vector usually captures more detail. Matryoshka training makes a vector you can safely cut shorter, whenever you need to save space.

Think about Russian nesting dolls. A big doll opens to reveal a smaller, complete doll inside. That one opens to reveal an even smaller complete doll inside it. Each smaller doll is not a broken fragment of the big one. It is a smaller, whole doll in its own right.

A Matryoshka embedding is built the same way. A 768-number vector contains, within its first 128 numbers, a smaller vector that is still useful on its own. Not a random, broken slice — a genuinely smaller, complete version of the same idea.

Why it exists

Bigger embedding vectors generally capture finer distinctions. But they cost more: more storage per vector, more memory, slower comparison at search time. Multiply that cost by millions of stored documents, and 768 numbers versus 128 becomes a real, expensive decision.

The obvious fix — chop a big vector down to its first 128 numbers — usually does not work well. Most embedding models were never trained with that shortening in mind. The numbers scattered across all 768 positions share the workload fairly evenly. Cut off most of them, and you throw away information spread across the whole vector. None of it was packed neatly at one end.

Matryoshka Representation Learning, introduced in 2022, trains a model differently. It deliberately teaches the early numbers to carry the most important, most general information. That way, cutting the vector short becomes a graceful, intentional trade-off, rather than an accident.

How it works

   A Matryoshka-trained 384-number vector:

     first  32 numbers  -> already a usable, coarse summary on its own
     first  64 numbers  -> noticeably better than 32
     first 128 numbers  -> noticeably better than 64
     all   384 numbers  -> the full, most detailed version

   Each prefix is a smaller, still-meaningful version of the whole —
   not a random, unusable fragment.

During training, the model is scored — and its numbers adjusted — not only on the full-length vector. Several shortened prefixes get scored too, at the same time. That repeated pressure is what makes early truncation safe later, at search time.

Where you have already seen it

  • Modern AI search products, offering a "fast" and "accurate" mode built on the same underlying vectors, at different lengths.
  • Newer embedding APIs, letting you request a shorter vector directly, trading some accuracy for real storage savings.
  • Mobile and on-device AI search, where a shorter vector is often the only workable option at all.

Remember this

  • A bigger embedding vector usually captures more detail, but costs more to store and compare.
  • Naively chopping a normal embedding vector short usually loses information unevenly and unpredictably.
  • Matryoshka training deliberately front-loads the important information, so a shortened prefix stays genuinely useful.

What to learn next

  • Embedding quantization — a different, complementary way of shrinking the same vectors.
  • Sentence-transformers — the kind of model these truncated vectors usually come from.
  • PCA — a much older, unrelated technique for reducing dimensions, worth contrasting against this one.

Developer — Code and libraries.

Testing naive truncation — cutting a normal, non-Matryoshka-trained model's vectors short after the fact — to see honestly how far it can go before it breaks.

Setup

bash
pip install torch transformers

all-MiniLM-L6-v2 was not trained with Matryoshka Representation Learning. This experiment tests naive truncation on purpose, to show what you get without it — see the researcher block for how genuine Matryoshka-trained models differ.

Truncating embeddings and watching search quality change

truncation_test.py
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel

tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model.eval()

def encode(sentences):
    enc = tok(list(sentences), padding=True, return_tensors="pt")
    with torch.no_grad():
        out = model(**enc).last_hidden_state
    mask = enc["attention_mask"].unsqueeze(-1).float()
    pooled = (out * mask).sum(1) / mask.sum(1).clamp(min=1e-9)
    return F.normalize(pooled, dim=1)

corpus = [
    "the chai stall serves hot tea every morning",
    "fresh coffee beans delivered to your door",
    "the stock market fell sharply this afternoon",
    "quarterly earnings beat analyst expectations",
    "the football match ended in a draw",
    "the cricket team won the tournament final",
    "heavy rain flooded the streets overnight",
    "the monsoon arrived a week early this year",
]
query = "will it rain today"

full = encode(corpus)               # shape [8, 384]
query_vec = encode([query])         # shape [1, 384]

def top1(doc_vecs, q_vec):
    sims = (doc_vecs @ q_vec.T).squeeze(1)
    return int(torch.argmax(sims))

full_top1 = top1(full, query_vec)
print(f"query: {query!r}")
print(f"dim 384 (full) top-1: {corpus[full_top1]!r}\n")

for dim in [64, 32, 16, 8, 4, 2]:
    truncated = F.normalize(full[:, :dim], dim=1)          # slice, then re-normalize
    q_truncated = F.normalize(query_vec[:, :dim], dim=1)
    result = top1(truncated, q_truncated)
    correct = "correct" if result == full_top1 else "WRONG"
    print(f"dim {dim:3} top-1: {corpus[result]!r:55} [{correct}]")
Output
query: 'will it rain today'
dim 384 (full) top-1: 'heavy rain flooded the streets overnight'

dim  64 top-1: 'heavy rain flooded the streets overnight'              [correct]
dim  32 top-1: 'heavy rain flooded the streets overnight'              [correct]
dim  16 top-1: 'heavy rain flooded the streets overnight'              [correct]
dim   8 top-1: 'the monsoon arrived a week early this year'            [WRONG]
dim   4 top-1: 'the monsoon arrived a week early this year'            [WRONG]
dim   2 top-1: 'fresh coffee beans delivered to your door'             [WRONG]

Line by line

Truncation holds up surprisingly well down to 16 dimensions — a 24x reduction from 384 — for this small, easy example. That is not a guarantee. It reflects that this particular model happened to front-load a fair amount of general signal into its earlier dimensions, as an accident of how it was trained, not a promise.

At 8 dimensions, the result quietly degrades rather than collapsing entirely. "The monsoon arrived a week early" is not the correct top match, but it is at least topically related to rain — a soft failure, not a random one, because some general signal about the topic survived even that much cutting.

At 2 dimensions, the result becomes genuinely unrelated — "fresh coffee beans delivered to your door" has nothing to do with rain. This is the honest floor: naive truncation, on a model never trained for it, eventually produces nonsense, and there was no guarantee about exactly where that would happen going in.

Common mistakes

Assuming any embedding model can be safely truncated. As shown above, it can work reasonably well by chance on an easy example, and it can also fail. Without a model specifically trained for Matryoshka-style truncation, there is no reliable way to know in advance how much accuracy you are giving up at a given cut length.

Truncating without re-normalizing. F.normalize(full[:, :dim], dim=1) is not optional — slicing a normalized vector down to fewer dimensions leaves it no longer unit length, which breaks the assumption that a plain dot product equals cosine similarity.

Choosing a truncated length based on one easy test example. The 8-dimension result above happened to still be topically close. A harder query, needing a finer distinction, could fail far earlier. Always test truncation against a representative, reasonably difficult evaluation set, not one convenient example.

Try it yourself

Change the query to something requiring a finer distinction — for instance "cricket score update" instead of the broader "will it rain today" — and re-run. Expect the truncated ranking to break down at a longer length than 8 dimensions, since distinguishing "cricket team won" from other, more loosely related sentences needs more of the vector's detail preserved.

What to learn next

  • Embedding quantization — shrinking vectors a different way, and often combined with truncation.
  • PCA — an older dimensionality-reduction technique, contrasted with Matryoshka training in the researcher block.
  • Sentence-transformers — where these embeddings come from before any truncation happens.

Researcher — Mathematics and papers.

Matryoshka Representation Learning

Kusupati, Bhatt, Baik et al. (2022), Matryoshka Representation Learning, train a single embedding to be simultaneously useful at multiple nested dimensionalities by modifying the training loss to be a weighted sum of losses computed at several prefix lengths:

text
L_Matryoshka = sum over m in M of  c_m * L(f(x)[0:m], y)
  • M is a chosen set of nested dimensions, e.g. {8, 16, 32, 64, 128, ..., d}.
  • f(x)[0:m] is the first m coordinates of the full embedding f(x).
  • L is the task loss (classification cross-entropy in the original paper; a contrastive loss for sentence embedding applications).
  • c_m are per-dimension weights, often uniform, sometimes tuned to prioritize particular target lengths.

Because every prefix length is included in every training step's gradient, information genuinely useful at the shortest included length is pushed toward the earliest coordinates — not as an emergent accident, the way it might loosely happen in an ordinarily trained model, but as a direct consequence of the training objective itself.

Contrast with the developer block's naive truncation

The developer block's all-MiniLM-L6-v2 was trained with a single loss on the full 384-dimensional output only. Its apparent robustness to truncation down to 16 dimensions on the tested example is not a property the training process targeted — it likely reflects that mean-pooled transformer embeddings tend to carry a moderate amount of generically useful signal spread broadly across dimensions, similar to how the leading principal components of many trained representations capture disproportionate variance even without explicit training pressure to do so. This is a weaker, statistical tendency, not a guarantee, which is precisely why the failure at 8 dimensions in the developer block is expected rather than surprising.

Relationship to PCA and nested dropout

Matryoshka training is conceptually related to two earlier ideas. PCA (see PCA) also produces components ordered by importance, but as an unsupervised post-hoc linear transformation of a fixed representation, not something the base representation itself was trained to support — PCA can be applied to any embedding after the fact, at the cost of an extra transform and without task-specific optimization of what "importance" means. Nested dropout (Rippel, Gelbart & Adams, 2014) is Matryoshka's more direct conceptual ancestor: it randomly truncates a representation's dimensions during training with probabilities that grow with depth, encouraging exactly the same ordered-importance property, applied originally to autoencoder representations rather than contrastive sentence embeddings.

Production adoption

Several widely used embedding APIs now ship Matryoshka-trained models explicitly, allowing a caller to request a shorter output vector directly, with published accuracy-versus-length trade-off curves: OpenAI's text-embedding-3 family (2024), Nomic's nomic-embed-text-v1.5, and Mixedbread AI's mxbai-embed-large-v1, among others. Kusupati et al.'s original paper reported that an ImageNet classifier's 8-dimensional Matryoshka prefix could match the accuracy of a similarly-sized independently trained low-dimensional model, while retaining the option to use the full representation when higher accuracy is worth the cost — the same representation genuinely serves both regimes, rather than requiring two separately trained and maintained models.

Complexity and the deployment trade-off

Truncation itself is O(1) — a slice operation, no recomputation. The value of Matryoshka training is entirely in avoiding the need to train, store, and maintain multiple separately-sized models for different deployment tiers (a fast, coarse tier and a slow, precise tier). A single Matryoshka-trained model, encoded once at full length and stored once, can serve both a cheap approximate-search first pass (using a short truncated prefix) and an accurate re-ranking pass (using the full vector) without any additional encoding cost — exactly the pattern used in the two-stage retrieval systems increasingly common in production search.

Key references

  • Rippel, O., Gelbart, M. A. & Adams, R. P. (2014). Learning Ordered Representations with Nested Dropout. ICML. The conceptual predecessor.
  • Kusupati, A. et al. (2022). Matryoshka Representation Learning. NeurIPS. arXiv:2205.13147
  • OpenAI (2024). New Embedding Models and API Updates. Announcement of the text-embedding-3 family's native dimension-reduction support.

Current state and open problems

Matryoshka training is now common practice for newly released embedding models, but it is not universal, and the developer block's result is a useful reminder that a model's published architecture or paper does not automatically tell you whether truncation is safe — it must either be explicitly documented as Matryoshka-trained, or empirically verified per model. An open practical question is how to choose the right cut length for a given application without exhaustive per-length evaluation, since the accuracy-versus-length curve is model- and domain-dependent, and the failure point demonstrated in the developer block (a soft degradation, then a sharper one) is a typical shape but not a universal, precisely predictable one.

What to learn next

  • Embedding quantization — a complementary compression technique, often stacked with Matryoshka truncation.
  • PCA — the older, unsupervised dimensionality-reduction technique this lesson's approach is often compared against.
  • Sentence-transformers — the base training setup Matryoshka's multi-length loss is typically added onto.