Semantic Search and Reranking

Learned sparse retrieval with SPLADE

SPLADE builds a sparse vector where every dimension is an actual vocabulary word, letting a masked language model quietly expand a short text with related words it never wrote at all.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

SPLADE builds a word-based search vector. It also fills in related words you never wrote, using a language model's own vocabulary knowledge.

Write a shopping list — "milk, bread" — and hand it to someone who knows your habits well. A thoughtful helper might pencil in "curd, paneer" nearby. Those items tend to go with milk in your usual shopping, even though you never wrote them down.

SPLADE builds exactly this kind of expanded, penciled-in list for a piece of text, automatically. Every entry on the list is a real, readable word from the vocabulary. A dense embedding, by contrast, is only numbers that mean nothing on their own. But the list includes more than the words originally written. It expands using everything a language model has learned about which words keep company with which.

Why it exists

Keyword search matches only literal words. Miss the exact wording, and you miss the match. Dense semantic search fixes that, but its vectors are opaque numbers. You cannot look at one and say "this is about waterproofing," the way you can with a word list.

SPLADE, introduced by Formal, Piwowarski & Clinchant in 2021, sits between the two. It builds a sparse vector, like keyword search, where almost every dimension is zero. Every non-zero dimension corresponds to one real vocabulary word. But the values come from a language model. It can activate "sandals" for a document that only ever said "shoes." It learned those words tend to appear in related contexts.

How it works

text: "affordable running shoes with good cushioning"
        |
        v
run it through a masked-language-model head
(the same head trained to predict a hidden word from context)
        |
        v
this gives every vocabulary word a score for this text --
most words get pushed to zero, but a handful of RELATED
words, some never written in the original text, activate too

result:  shoes (2.96)  running (2.90)  cushion (2.88)
         walking (2.46)   <- never appeared in the text!
         sneakers, sandals, boots  <- also never written, also activated

The list stays human-readable. Every activated dimension is a real word. It quietly covers vocabulary the original text never used, closing exactly the gap keyword search struggles with.

A real example you have seen

Search "sneakers" on a shopping site, and it also surfaces listings that only ever say "running shoes" or "trainers." No exact word needs to overlap at all. A learned sparse model like SPLADE is one common way that expansion happens, sitting right inside the search pipeline. It stays fast and interpretable in a way a fully dense vector search is not.

Remember this

  • SPLADE builds a sparse vector where every dimension is a real, readable vocabulary word, unlike a dense embedding.
  • The model activates related words the text never literally contained, using what it learned during language-model pretraining.
  • It keeps the interpretability and infrastructure-friendliness of keyword search, with some of semantic search's ability to bridge different wordings.

What to learn next

Developer — Code and libraries.

This uses distilbert-base-uncased's existing masked-language-model head to demonstrate the SPLADE mechanism directly — this is not a real, specially-trained SPLADE checkpoint. A production SPLADE model is fine-tuned with a contrastive ranking loss plus explicit sparsity regularisation; an off-the-shelf, un-fine-tuned MLM head produces the same shape of output without the same retrieval-quality guarantees. Treat the term weights below as illustrative, not as calibrated relevance scores.

Setup

bash
pip install transformers torch

distilbert-base-uncased downloads once, roughly 260 MB. Outputs verified with transformers 5.6.2 and torch 2.5.1 on CPU.

Building a sparse vector, and finding query-document overlap

splade_demo.py
import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM

tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased")
model = AutoModelForMaskedLM.from_pretrained("distilbert-base-uncased")
model.eval()

SKIP = set(tokenizer.all_special_tokens) | {".", ",", ";", "?", "!", "with", "and", "the", "a", "of", "for"}


def splade_vector(text):
    inputs = tokenizer(text, return_tensors="pt")
    with torch.no_grad():
        logits = model(**inputs).logits[0]              # (seq_len, vocab_size)
    weights = torch.log1p(torch.relu(logits))            # log(1 + relu(x)) per token per vocab entry
    sparse, _ = weights.max(dim=0)                        # max-pool over sequence positions
    return sparse


def top_terms(vec, k=8):
    values, idx = vec.topk(k + len(SKIP))
    out = []
    for v, i in zip(values, idx):
        tok = tokenizer.convert_ids_to_tokens([i.item()])[0]
        if tok in SKIP:
            continue
        out.append((tok, round(v.item(), 2)))
        if len(out) == k:
            break
    return out


doc = "affordable running shoes with good cushioning"
vec = splade_vector(doc)
print("input text:", repr(doc))
print("non-zero dimensions:", int((vec > 0).sum()), "out of", vec.numel(), "vocabulary entries")
print("top activated terms (includes words not in the original text):")
for term, weight in top_terms(vec):
    print(f"  {term:<12} {weight}")

query = "cheap jogging sneakers"
q_vec = splade_vector(query)
overlap = torch.minimum(vec, q_vec)
nz = overlap.nonzero().squeeze(-1)
contribs = [(tokenizer.convert_ids_to_tokens([i.item()])[0], round(overlap[i].item(), 2)) for i in nz]
contribs = [(t, c) for t, c in contribs if t not in SKIP]

print(f"\nquery: {query!r}")
print("shared activated terms and their contribution to the match score:")
for term, c in sorted(contribs, key=lambda x: -x[1])[:8]:
    print(f"  {term:<12} {c}")
print(f"sparse dot-product score: {overlap.sum().item():.2f}")
Output
input text: 'affordable running shoes with good cushioning'
non-zero dimensions: 6470 out of 30522 vocabulary entries
top activated terms (includes words not in the original text):
  shoes        2.96
  running      2.9
  good         2.88
  cushion      2.88
  ##ing        2.75
  feet         2.48
  walking      2.46
  affordable   2.45

query: 'cheap jogging sneakers'
shared activated terms and their contribution to the match score:
  shoes        2.75
  sandals      2.43
  walking      2.42
  boots        2.4
  sneakers     2.33
  running      2.31
  shoe         2.31
  barefoot     2.17
sparse dot-product score: 2162.54

Neither "sandals," "boots" nor "barefoot" appear anywhere in the original document text, yet all three activated strongly enough to contribute to the match score against a query about "sneakers" — the model expanded the document's vocabulary using its own learned sense of which words belong to the same world as "shoes" and "running."

The walkthrough

torch.log1p(torch.relu(logits)) is the whole SPLADE transform. relu zeroes out every negative logit (words the model actively thinks are unlikely given this context), and log1p compresses the remaining positive values so a handful of very confident words cannot completely dominate the vector.

weights.max(dim=0) pools across token positions. A word can be strongly implied by any position in the sequence — max-pooling takes each vocabulary word's single strongest activation across the whole input, rather than averaging it down across positions where it was less relevant.

torch.minimum(vec, q_vec) computes term-by-term overlap. For every vocabulary word, this takes whichever of the query's or document's activation is smaller — a word only contributes to the match score in proportion to how strongly both sides activated it, which is the sparse-vector equivalent of a dot product when most dimensions are zero on at least one side.

SKIP filters punctuation and a handful of stop words for a readable demo. A real, trained SPLADE model suppresses these itself, as part of its training objective — shown here as a manual filter purely to keep the printed output legible.

Common mistakes

Treating this untrained-for-the-task MLM head as retrieval-quality output. As stated up front, this demonstrates the computation, not calibrated relevance. A real SPLADE checkpoint is trained with a ranking loss over labelled query-document pairs specifically so its activations reflect genuine relevance, not only generic language-model plausibility.

Ignoring how many dimensions activate. 6,470 out of 30,522 vocabulary entries activated for one six-word sentence here — far denser than a typical trained SPLADE vector, which is explicitly regularised toward much greater sparsity during training. This matters directly for storage and search cost at scale.

Comparing the sparse dot-product score across different document lengths without normalisation. Longer documents naturally activate more terms and accumulate a higher raw score, independent of relevance — production sparse-retrieval systems normalise for this.

Try it yourself

Change the document to a completely different sentence, such as "a laptop with a fast processor and long battery life", and re-run the query "cheap jogging sneakers" against it. Confirm the shared-terms overlap collapses toward nothing — a useful sanity check that the expansion mechanism responds to genuine topical relatedness, not to any input at all.

What to learn next

Researcher — Mathematics and papers.

The SPLADE representation, formally

For a document (or query) with token logits l_{i,j} — the masked-language-model logit for vocabulary entry j at sequence position i — SPLADE computes a single sparse vector w (all entries >= 0) over the vocabulary V:

text
w_j = max over i=1..N of  log(1 + ReLU(l_{i,j}))

Where N is sequence length. Relevance between query q and document d is then the sparse dot product w_q . w_d, computed exactly as with BM25 or any classical inverted-index score — which is precisely why SPLADE integrates directly with existing inverted-index search infrastructure rather than requiring a dense vector index at all.

Training: ranking loss plus FLOPS regularisation

An untrained MLM head, as used in the developer block, produces a plausible-shaped vector with no relevance calibration. A real SPLADE model (Formal, Piwowarski & Clinchant, 2021; SPLADE v2, Formal, Lassance, Piwowarski & Clinchant, 2021) is trained with two objectives combined:

  1. A contrastive ranking loss over (query, positive document, negative document) triples, pushing w_q . w_d+ above w_q . w_d- — the same broad family of objective used to train dense bi-encoders, applied here to sparse vectors instead.
  2. FLOPS regularisation (Paria et al., 2020), an L1-style penalty on the average activation per vocabulary dimension across a batch, explicitly pushing the model toward far greater sparsity than an unregularised MLM head naturally produces — directly addressing the "too many activated dimensions" issue observed in the developer block's untrained demonstration.

Later versions (SPLADE-doc, SPLADE++) apply asymmetric strategies — expanding documents more aggressively than queries, or vice versa — since document-side expansion can be precomputed offline while query-side expansion adds to online latency.

Why sparse-and-learned, rather than dense, in some deployments

Because every dimension of a SPLADE vector corresponds to an actual vocabulary term, the resulting representation is directly compatible with an inverted index — the same data structure BM25 relies on — rather than requiring a dedicated ANN vector index such as FAISS. This gives SPLADE a genuine infrastructure advantage in organisations with mature inverted-index search systems already in production, alongside a degree of interpretability dense embeddings do not offer: you can inspect exactly which terms drove a given match, as demonstrated directly in the developer block.

Key references

  • Formal, T., Piwowarski, B. & Clinchant, S. (2021). SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. arXiv:2107.05720
  • Formal, T., Lassance, C., Piwowarski, B. & Clinchant, S. (2021). SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval. arXiv:2109.10086
  • Paria, B. et al. (2020). Minimizing FLOPs to Learn Efficient Sparse Representations. ICLR.

Current state

SPLADE and related learned sparse models are used in production both as standalone first-stage retrievers over inverted indexes and as one signal fused with dense retrieval in hybrid systems, following the same reciprocal rank fusion pattern covered earlier in this section. Their appeal remains the combination that neither pure BM25 nor pure dense retrieval offers alone: learned, context-aware term weighting and expansion, on infrastructure that still looks, to the rest of the system, like ordinary keyword search.

What to learn next