The Embedding Family

Sentence-transformers

Sentence-transformers turns a whole sentence into one vector built specifically so that similar meaning lands close together, which is what makes fast semantic search possible.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Sentence-transformers turns a whole sentence into one vector, built specifically so that sentences with similar meaning end up close together.

Think about a librarian who has read every book in a huge library. Without opening a single one, they can tell you which two books cover the most similar topics. That instant sense of "these two belong near each other" is what a sentence-transformer gives a computer. For text, instead of books.

BERT is powerful at reading a sentence. But it was not originally built to produce one clean vector that plays well with simple distance comparisons. Sentence-transformers is a small but important fix to that gap. It is the engine behind most modern search-by-meaning tools.

Why it exists

Say you want to search ten thousand support tickets for the ones closest in meaning to a brand new ticket. With plain BERT, the honest way to compare two texts is to feed both into the model together. Let it judge their relationship directly. Accurate, but far too slow to repeat ten thousand times.

Sentence-transformers, introduced in 2019 as Sentence-BERT, solves this differently. It trains a model to produce one good vector per sentence. Simple distance between two such vectors then reliably reflects how similar their meanings are. Each comparison becomes a fast calculation, not a full model run.

How it works

Feed a sentence in, and get one fixed-length list of numbers out. Usually a few hundred numbers, whether the sentence was five words or fifty.

   "cheap phone under 15000"          -> [ 0.02, -0.11, 0.44, ... ]  (384 numbers)
   "best budget smartphones this year" -> [ 0.03, -0.09, 0.41, ... ]  (384 numbers)
   "how to boil rice perfectly"        -> [-0.22,  0.31, 0.02, ... ]  (384 numbers)

   Compare the lists with cosine similarity:

     phone sentence  <-> smartphone sentence  -> HIGH similarity (similar meaning)
     phone sentence  <-> rice sentence         -> LOW similarity (unrelated)

Two sentences that mean similar things land close together, even sharing no exact word. That is the whole payoff — meaning-based comparison, at a speed plain BERT was never built to deliver.

Where you have already seen it

  • "Search that understands you." A phone-shop query returns results with none of your exact words, matched purely by meaning.
  • "Find similar questions" on forums and support sites, matching a new question to previously answered ones.
  • Chatbots that read your documents, matching your question to the closest paragraph. The retrieval half of a RAG system.

Remember this

  • Sentence-transformers builds one vector per whole sentence, specifically tuned so that similar meaning lands close together.
  • It is dramatically faster than comparing sentences directly with plain BERT, because comparison becomes simple vector maths.
  • It is the standard engine behind semantic search, question matching, and the retrieval step in most RAG systems.

What to learn next

  • Pooling strategies — the exact technique used to squeeze many word vectors into one sentence vector.
  • Vector databases — where millions of these sentence vectors get stored and searched at speed.
  • What is RAG? — the system that puts sentence-transformer search in front of an LLM.

Developer — Code and libraries.

The sentence-transformers library wraps a model and a pooling step behind one call: SentenceTransformer(model_name).encode(sentences). Under the hood, for a model like all-MiniLM-L6-v2, that call runs the model and then averages its token vectors together — mean pooling, covered in full in pooling strategies.

The code below reproduces exactly that, one layer down, using the underlying transformers library directly. It is worth seeing once at this level, because it shows precisely what "encode a sentence" actually computes.

Setup

bash
pip install torch transformers

all-MiniLM-L6-v2 downloads once, about 90MB, and is small enough to run comfortably on CPU.

Building sentence vectors and searching by meaning

semantic_search.py
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel

MODEL_NAME = "sentence-transformers/all-MiniLM-L6-v2"
tok = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModel.from_pretrained(MODEL_NAME)
model.eval()

def embed(sentences):
    enc = tok(list(sentences), padding=True, truncation=True, return_tensors="pt")
    with torch.no_grad():
        hidden = model(**enc).last_hidden_state              # [batch, tokens, 384]
    mask = enc["attention_mask"].unsqueeze(-1).float()
    pooled = (hidden * mask).sum(1) / mask.sum(1).clamp(min=1e-9)  # mean of real tokens
    return F.normalize(pooled, dim=1)                         # unit length, so dot = cosine

corpus = [
    "cheap phone under 15000",
    "best budget smartphones this year",
    "how to boil rice perfectly",
    "easy rice recipes for beginners",
    "flight delays at the airport today",
]
query = "affordable mobile phones"

corpus_vecs = embed(corpus)
query_vec = embed([query])

scores = (corpus_vecs @ query_vec.T).squeeze(1)
ranking = torch.argsort(scores, descending=True)

print(f"query: {query!r}\n")
for i in ranking:
    print(f"  {scores[i].item():.3f}  {corpus[i]}")
Output
query: 'affordable mobile phones'

  0.745  best budget smartphones this year
  0.693  cheap phone under 15000
  0.115  easy rice recipes for beginners
  0.034  how to boil rice perfectly
  0.028  flight delays at the airport today

Neither phone sentence shares a single word with the query "affordable mobile phones" beyond none at all — "affordable" versus "cheap"/"budget", "mobile phones" versus "smartphones"/"phone". Yet both rank far above the unrelated sentences. This is meaning-based matching, not word matching — exactly what BM25 cannot do on its own.

Line by line

padding=True pads every sentence in the batch to the same token length, since sentences differ in length but tensors need a fixed shape. attention_mask records which tokens are real and which are padding, so padding never leaks into the averaged vector — see the masking step next.

(hidden * mask).sum(1) / mask.sum(1) is mean pooling: sum the real (non-padding) token vectors, divide by how many real tokens there were. This single line is what the sentence-transformers library's .encode() call does internally for this model.

F.normalize(pooled, dim=1) scales every vector to length 1. Once vectors are unit length, a plain dot product equals cosine similarity — corpus_vecs @ query_vec.T above is doing full cosine comparison in one matrix multiply, no separate normalization step needed at search time.

Common mistakes

Comparing un-normalized vectors with dot product. Without the F.normalize step, longer or larger-magnitude vectors would score higher purely from their size, regardless of actual similarity. Always normalize before comparing with a plain dot product, or use cosine similarity directly.

Forgetting model.eval(). Some models behave differently in training mode versus evaluation mode (dropout layers, in particular). Skipping model.eval() before inference can introduce small, unwanted randomness into otherwise deterministic embeddings.

Re-encoding your entire corpus on every single query. In the example above, corpus_vecs should be computed once and reused. Real systems encode the corpus once, store the vectors in a vector database, and only encode the query fresh, every time someone searches.

Try it yourself

Add "where to fix a mobile phone screen" to corpus and re-run. Check whether it ranks near the other phone sentences — this tests whether the model's sense of "similar" survives a fairly different intent (repair, not purchase) sharing the same topic.

What to learn next

Researcher — Mathematics and papers.

The problem SBERT was built to solve

Before Sentence-BERT, comparing two sentences with BERT meant feeding both together as a single input pair and reading a relatedness score from the model — a cross-encoder. This gives strong accuracy, because the model can attend across both sentences jointly, but its cost is quadratic in the size of the collection: finding the most similar pair among n sentences requires O(n^2) full forward passes. Reimers & Gurevych (2019) quote a concrete figure: finding the most similar pair among 10,000 sentences with a BERT cross-encoder would take roughly 65 hours on a modern GPU at the time of writing.

The Sentence-BERT architecture

SBERT instead uses a siamese (or triplet) network: the same BERT encoder is applied independently to each sentence, producing two separate token-level outputs, which are each pooled into a single vector (mean pooling was found to perform best in the paper's ablations — see pooling strategies). Training objectives compare these two independent embeddings, forcing the encoder to place similar sentences close together in the pooled vector space specifically — a property plain, off-the-shelf BERT embeddings do not reliably have, as shown empirically in the same paper: raw BERT [CLS] and average-pooled embeddings performed worse than simple GloVe averaging on semantic textual similarity benchmarks, prior to this fine-tuning.

With embeddings precomputed once per sentence, finding the most similar pair among n sentences drops to O(n) encoding time plus an O(n^2) or, with approximate nearest-neighbour indexing, sub-quadratic comparison over already-computed vectors — the same 10,000-sentence task the paper estimated at 65 hours for a cross-encoder was reported at around 5 seconds with SBERT.

Training objectives

The original paper used three:

  • Classification objective, for labelled pairs (e.g. natural language inference): concatenate (u, v, |u - v|) and pass through a softmax classifier, u and v being the two sentence embeddings.
  • Regression objective, for continuous similarity labels: minimize mean-squared error between cos(u, v) and the gold similarity score.
  • Triplet objective, for (anchor, positive, negative) triples: push the anchor closer to the positive than to the negative by at least a margin.

Later, contrastive in-batch negative objectives (see fine-tuning an embedding model and hard negative mining) became the dominant training approach for large-scale sentence embedding models, largely superseding the original three.

Complexity

Encoding n sentences: O(n * L^2 * d) for typical sequence length L and model width d, dominated by the encoder's per-sentence attention cost, run independently and in parallel across sentences — the crucial difference from a cross-encoder, whose cost scales with pairs, not individual sentences.

Key references

  • Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. EMNLP-IJCNLP. arXiv:1908.10084
  • Devlin, J. et al. (2018). BERT. arXiv:1810.04805. The base encoder architecture SBERT fine-tunes.
  • Wang, W. et al. (2020). MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. arXiv:2002.10957. The distillation technique behind the small, fast all-MiniLM-L6-v2 model used throughout this section.

Current state and open problems

Sentence-transformer-style models are now trained on much larger contrastive datasets than the original SBERT paper used — often hundreds of millions of mined pairs — and the field's active research has shifted toward hard negative mining quality, multilingual coverage, and embedding dimension efficiency rather than the base architecture, which has remained stable since 2019. The MTEB benchmark (Muennighoff et al., 2023) is now the standard way models in this family are compared, covering retrieval, clustering, classification and similarity tasks across dozens of datasets, precisely because single-task benchmarks like the original semantic textual similarity sets proved too narrow to reflect real downstream use.

What to learn next