The Embedding Family

CLS, mean and max pooling

Pooling squeezes a sentence's many word vectors down into one vector, and the method chosen for that squeeze changes how well the result actually captures meaning.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Pooling squeezes many word vectors from one sentence into a single vector. How you squeeze changes how good the result is.

Think about summarising a cricket team's performance with one number. You could take the score of whoever batted first. You could average every batter's score. You could take the single highest score of the day. Each choice tells a different story about the same match. None of them is automatically the "right" one.

A BERT-style model outputs one vector per word, not one for the whole sentence. Pooling combines those word vectors into a single sentence vector. Exactly like the cricket example, different pooling methods tell noticeably different stories.

Why it exists

Once a transformer model reads a sentence, it hands back one vector per token. That is dozens of them, for anything longer than a few words. Sentence-transformers needs exactly one vector per sentence, to keep comparison fast and simple.

Getting from "many vectors" to "one vector" is not automatic. The method chosen matters more than it sounds like it should. Pick badly, and even a strong underlying model can produce disappointing sentence comparisons.

How it works

Three common choices, each using the same underlying word vectors differently.

   Sentence:  "the chai was hot"
   Word vectors:  [the]  [chai]  [was]  [hot]

   CLS pooling:   use only the model's special starting position's vector
                  -> ignores "chai", "was", "hot" entirely

   Mean pooling:  average all four word vectors together
                  -> every word contributes equally

   Max pooling:   take the single strongest value, position by position,
                  across all four word vectors
                  -> the loudest signal wins, wherever it came from

CLS pooling relies on one special marker token, placed at the start of every BERT-family input. It was designed for classification, not for summarising meaning on its own.

Mean pooling averages every real word's vector together, giving each word an equal vote.

Max pooling takes, for each position in the vector, whichever word contributed the largest value. It is a "loudest word wins" summary.

Where you have already seen it

  • Every semantic search tool, silently, since some pooling method sits behind every sentence vector it produces.
  • Recommendation systems that compare product descriptions or article summaries as single vectors.
  • Any tool built on sentence-transformers, where the library picks a pooling method for you, based on the model.

Remember this

  • Pooling turns many word vectors into one sentence vector, and the method used genuinely changes the result's quality.
  • CLS uses one special position's vector. Mean averages every word. Max keeps only the strongest signal per position.
  • The best choice depends on how the specific model was trained — there is no universally correct pooling method.

What to learn next

Developer — Code and libraries.

All three pooling methods, computed from the same underlying model output, compared on whether they can tell a similar sentence pair from an unrelated one.

Setup

bash
pip install torch transformers

CLS vs mean vs max, measured directly

pooling.py
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel

tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model.eval()

sentences = [
    "the chai was hot and sweet",
    "the tea was warm and sugary",
    "the stock market fell sharply today",
]

enc = tok(sentences, padding=True, return_tensors="pt")
with torch.no_grad():
    out = model(**enc).last_hidden_state             # [batch, tokens, 384]
mask = enc["attention_mask"].unsqueeze(-1).float()    # 1 for real tokens, 0 for padding

def cls_pool(hidden):
    return hidden[:, 0]                               # the [CLS] token's own vector

def mean_pool(hidden, mask):
    summed = (hidden * mask).sum(1)
    counts = mask.sum(1).clamp(min=1e-9)
    return summed / counts                            # average of the real tokens only

def max_pool(hidden, mask):
    hidden = hidden.masked_fill(mask == 0, -1e9)       # padding must never win the max
    return hidden.max(1).values

pooled = {
    "cls":  F.normalize(cls_pool(out), dim=1),
    "mean": F.normalize(mean_pool(out, mask), dim=1),
    "max":  F.normalize(max_pool(out, mask), dim=1),
}

for name, vecs in pooled.items():
    sim_similar = F.cosine_similarity(vecs[0:1], vecs[1:2]).item()    # chai vs tea sentence
    sim_different = F.cosine_similarity(vecs[0:1], vecs[2:3]).item()  # chai vs stock market
    gap = sim_similar - sim_different
    print(f"{name:5}  similar-pair={sim_similar:.3f}   unrelated-pair={sim_different:.3f}   gap={gap:.3f}")
Output
cls    similar-pair=0.673   unrelated-pair=0.595   gap=0.078
mean   similar-pair=0.439   unrelated-pair=-0.024   gap=0.463
max    similar-pair=0.810   unrelated-pair=0.718   gap=0.092

Line by line

"gap" measures the useful signal: how much higher the similar pair scores than the unrelated pair. A larger gap means the pooling method is doing a better job of separating "these mean similar things" from "these do not". Mean pooling's gap of 0.463 dwarfs CLS pooling's 0.078 here.

This specific result is not a general law about CLS being weak — it is specific to this model, and that matters. all-MiniLM-L6-v2 was trained with mean pooling as its actual training objective. Its CLS token was never optimised to summarise a sentence on its own, so it underperforms here for a specific, understandable reason. A model trained specifically to use its CLS token well — some cross-encoder and classification-tuned models are — would show the opposite pattern.

Max pooling scores everything higher in absolute terms, but its gap is not much better than CLS's. Taking the single strongest value per dimension tends to inflate similarity scores generally, including for unrelated sentences, which is why its "unrelated-pair" score (0.718) is high too — not because it found real similarity, but because "loudest signal per position" is a noisier summary than an average.

Common mistakes

Assuming mean pooling is always the right choice. It performs best here because of how this specific model was trained. Using mean pooling on a model trained with CLS-based objectives can give worse results than using CLS pooling on that same model — always check what the model card recommends.

Mixing pooling methods between training and inference. If a model was fine-tuned using mean pooling, encoding new sentences with CLS pooling at inference time silently produces a different, uncalibrated vector space. This is a quiet, hard-to-debug source of poor search results in production.

Forgetting to mask padding before max pooling. Without masked_fill(mask == 0, -1e9), padding positions — which are not real words — can win the max in short sentences padded to match a longer one in the same batch, corrupting the result in a way that depends on batch composition, not only sentence content.

Try it yourself

Change the sentences to three that are all closely related in topic but phrased very differently — for instance three different ways of saying a flight was delayed — and compare the three gap scores again. Expect mean pooling's advantage to persist, since it is a property of how this particular model was trained, not of any one sentence pair.

What to learn next

Researcher — Mathematics and papers.

The three methods, formally

Given per-token hidden states H in R^{n x d} for a sequence of n tokens and an attention mask m in {0,1}^n:

text
CLS pooling:   v = H_0
Mean pooling:  v = (sum over i of m_i * H_i) / (sum over i of m_i)
Max pooling:   v = elementwise-max over i (where m_i = 1) of H_i
  • H_0 is the hidden state at the special classification-token position, present in BERT-family models by construction.
  • Mean pooling is a masked average over the sequence dimension; max pooling is a masked elementwise maximum.

Why the choice affects quality, structurally

Reimers & Gurevych (2019) ran a controlled ablation across all three pooling strategies on their SBERT training setup, and found mean pooling outperformed both CLS and max pooling on semantic textual similarity benchmarks by a clear margin, motivating its use as the SBERT default. Their finding, replicated by the developer block's measurement above, is best understood through what each token's representation is actually optimized to carry:

[CLS] is trained, in vanilla BERT pretraining, on the next sentence prediction auxiliary objective — a binary classification task, not a similarity-preserving embedding objective. Nothing in standard pretraining pushes the CLS vector's cosine geometry to reflect sentence meaning; it only needs to be linearly separable by a classification head, a considerably weaker requirement than "closer in cosine distance means more similar in meaning."

Mean pooling, by contrast, aggregates signal from every token equally, which empirically correlates better with meaning-level similarity even without specific pretraining pressure toward that outcome — plausibly because averaging is a low-variance estimator of "the sentence's overall content," while any single position's vector, CLS included, is a high-variance summary dependent on architecture-specific quirks of that position.

Anisotropy and pooling interact

Following on from static vs contextual embeddings's discussion of Ethayarajh (2019)'s anisotropy findings: since raw transformer hidden states occupy a narrow cone in vector space, pooling methods that are more sensitive to a small number of dominant, high-magnitude dimensions (CLS, and to a lesser extent max pooling) inherit that distortion more directly than mean pooling, which averages the distortion out across many positions. Timkey & van Schijndel (2021)'s rogue-dimension standardization, discussed in the same lesson, was shown to improve CLS-pooled similarity quality substantially — suggesting the CLS-versus-mean gap is partly an artifact of a small number of problematic dimensions, not an intrinsic property of the CLS token itself.

What changes after fine-tuning

Once a model is fine-tuned with a contrastive or triplet objective directly on the pooled output (see fine-tuning an embedding model), the gap between pooling methods narrows considerably, because the fine-tuning gradient now flows directly into whichever pooling output was chosen, shaping that specific representation toward the similarity objective. This is why some fine-tuned models — cross-encoders trained explicitly to use CLS for a classification-style relevance score — perform perfectly well with CLS pooling, while general-purpose base checkpoints, evaluated zero-shot, reliably favour mean pooling in the published literature.

Complexity

All three pooling methods are O(n * d), a single pass over the sequence — negligible relative to the O(n^2 * d) encoder forward pass that produces H in the first place. Pooling choice affects representation quality, never runtime cost.

Key references

  • Devlin, J. et al. (2018). BERT. arXiv:1810.04805. Defines the CLS token and next-sentence-prediction pretraining.
  • Reimers, N. & Gurevych, I. (2019). Sentence-BERT. EMNLP-IJCNLP. arXiv:1908.10084. The controlled pooling ablation.
  • Ethayarajh, K. (2019). How Contextual are Contextualized Word Representations? EMNLP-IJCNLP. arXiv:1909.00512
  • Timkey, W. & van Schijndel, M. (2021). All Bark and No Bite. EMNLP. arXiv:2109.04404

Current state and open problems

Mean pooling remains the practical default across most current open sentence embedding models, but it is not universal — some newer architectures use a learned attention-weighted pooling (a small trainable layer that decides per-token weights, rather than a fixed uniform average), and some retrieval-specific models append a dedicated pooling token trained end-to-end for exactly this purpose, sidestepping the CLS-versus-mean question entirely by not reusing either mechanism. The open question is less about which fixed pooling formula wins in general, and more about whether pooling should be a fixed post-hoc operation at all, versus something learned jointly with the rest of the model for a specific downstream objective.

What to learn next