The Embedding Family

Hard negative mining

Hard negative mining picks training examples that look deceptively similar but are actually wrong, because those near-misses teach an embedding model far more than plainly unrelated examples do.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Hard negative mining picks "wrong answer" examples that look deceptively similar to the right one. Those near-misses teach a model far more than obvious mismatches do.

Think about studying for an exam using flashcards. A card asking "capital of France, or a banana?" teaches you nothing. No one would ever confuse those. A card asking "capital of France, or Lyon?" forces you to actually know. Both are real French cities, and only one is right.

Training an embedding model works the same way. Showing it a plainly wrong example teaches it almost nothing. Showing it a deceptively close wrong example — a hard negative — is where the real learning happens.

Why it exists

Fine-tuning an embedding model needs both correct pairs and incorrect pairs to learn from. The easiest way to get incorrect pairs is to grab random unrelated text and call it a negative example.

The problem: a strong model, even before fine-tuning, already scores random unrelated text as strongly dissimilar. Training on random negatives past a certain point stops teaching it anything new. It is being tested on questions it can already answer.

Hard negatives fix this. They are wrong answers deliberately chosen to be topically close, or to share vocabulary with the right answer. These are the examples that push a model to sharpen its sense of meaning. They do more than confirm what it already knows.

How it works

   Anchor:  "cannot login to my account"

   Random negative:  "the chai stall opens at seven in the morning"
                      -> plainly unrelated, model already knows this
                      -> teaches almost nothing

   Hard negative:     "my account was created successfully and is now active"
                       -> shares the word "account", same general topic
                       -> but means something close to the OPPOSITE
                       -> forces the model to learn a finer distinction

A hard negative is not a trick question with no right answer. It is a genuinely wrong answer, sharing enough surface similarity with the right one — vocabulary, topic, structure. A model must truly understand meaning to reject it.

Where you have already seen it

  • Search engines that avoid "almost right" results. Rankings are trained partly on hard negatives, so near-miss pages stop outranking the correct one.
  • Product search on e-commerce sites, trained to tell apart very similar-looking listings.
  • Any retrieval system behind a chatbot, where a wrong-but-similar paragraph is far more dangerous than a plainly unrelated one.

Remember this

  • A hard negative looks deceptively similar to the right answer. It is not a plainly unrelated example.
  • Random negatives stop teaching a model much once it has learned the easy, obvious distinctions.
  • Hard negatives force a model to learn the finer distinctions that actually matter in practice.

What to learn next

Developer — Code and libraries.

Measuring, directly, why a hard negative provides more training signal than a random one — using a real embedding model's own similarity scores.

Setup

bash
pip install torch transformers

Comparing a random negative against a hard negative

hard_negatives.py
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel

tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model.eval()

def encode(sentences):
    enc = tok(list(sentences), padding=True, return_tensors="pt")
    with torch.no_grad():
        out = model(**enc).last_hidden_state
    mask = enc["attention_mask"].unsqueeze(-1).float()
    pooled = (out * mask).sum(1) / mask.sum(1).clamp(min=1e-9)
    return F.normalize(pooled, dim=1)

anchor = "cannot login to my account"
positive = "unable to sign in to my account"
random_negative = "the chai stall opens at seven in the morning"
hard_negative = "my account was created successfully and is now active"

vecs = encode([anchor, positive, random_negative, hard_negative])
a, p, rn, hn = vecs[0], vecs[1], vecs[2], vecs[3]

sim_pos = F.cosine_similarity(a, p, dim=0).item()
sim_rand = F.cosine_similarity(a, rn, dim=0).item()
sim_hard = F.cosine_similarity(a, hn, dim=0).item()

print(f"anchor:            {anchor!r}")
print(f"positive:          {positive!r}")
print(f"random negative:   {random_negative!r}")
print(f"hard negative:     {hard_negative!r}")
print()
print(f"cosine(anchor, positive)        = {sim_pos:.3f}")
print(f"cosine(anchor, random negative) = {sim_rand:.3f}   margin to positive: {sim_pos - sim_rand:.3f}")
print(f"cosine(anchor, hard negative)   = {sim_hard:.3f}   margin to positive: {sim_pos - sim_hard:.3f}")
Output
anchor:            'cannot login to my account'
positive:          'unable to sign in to my account'
random negative:   'the chai stall opens at seven in the morning'
hard negative:     'my account was created successfully and is now active'

cosine(anchor, positive)        = 0.901
cosine(anchor, random negative) = 0.050   margin to positive: 0.851
cosine(anchor, hard negative)   = 0.524   margin to positive: 0.377

Line by line

The random negative scores 0.050 — already about as far from the positive's 0.901 as it could realistically get, before any training on this pair happens. A contrastive loss computed on this pair has almost nothing left to push against; the model already separates them cleanly.

The hard negative scores 0.524 — a genuinely uncomfortable middle ground. It is closer to the anchor than a random sentence has any right to be, purely from sharing vocabulary and topic, despite describing close to the opposite situation. This is exactly the kind of example a contrastive loss can still learn from — there is real distance left to push.

Margin to the positive tells the training story directly. The random negative's margin (0.901 − 0.050 = 0.851) is already comfortably large. The hard negative's margin (0.901 − 0.524 = 0.377) is much smaller — and a smaller, harder-won margin is where a gradient-based training step has genuine work left to do.

Common mistakes

Mining hard negatives so aggressively that some are actually mislabelled positives. If a "hard negative" is close enough to the anchor that a human reviewer would call it a correct match, training on it as a negative actively teaches the model something false. Manual spot-checking of mined hard negatives is standard practice for exactly this reason.

Using only hard negatives and dropping random ones entirely. A training batch made entirely of hard negatives can overcorrect, making the model overly cautious about superficially similar text in general. Most production training pipelines mix a majority of in-batch random negatives with a smaller number of deliberately mined hard ones.

Mining hard negatives using the same model you are about to fine-tune. This can create a feedback loop where the model's own current blind spots go unchallenged, since it will not surface a hard negative pattern it does not yet recognise as similar. A common fix is mining hard negatives with a different, independent method — commonly BM25, or a separately trained retriever — rather than the model being improved.

Try it yourself

Write your own anchor sentence and two candidate negatives — one you expect to be "random" and one you expect to be "hard" — and run them through encode. Check whether the hard negative's cosine similarity to the anchor genuinely lands in the uncomfortable middle range demonstrated above, roughly 0.4 to 0.6, rather than at either extreme.

What to learn next

Researcher — Mathematics and papers.

Why hard negatives matter, in loss-function terms

For the multiple negatives ranking loss / InfoNCE objective described in fine-tuning an embedding model:

text
L = -log( exp(sim(a, p) / tau) / sum over j of exp(sim(a, n_j) / tau) )
  • The gradient magnitude with respect to a given negative n_j scales with exp(sim(a, n_j) / tau) relative to the sum — a negative that already scores very low similarity contributes an exponentially small term, and correspondingly a near-zero gradient. A negative sitting close to the decision boundary (moderate similarity) contributes a much larger gradient term.

This is a direct, formal restatement of the developer block's empirical observation: the random negative (sim = 0.050) sits deep in the region where exp(sim/tau) is negligible, contributing little to the loss or its gradient, while the hard negative (sim = 0.524) sits in the region where the gradient is largest. Training predominantly on random negatives, especially as a model improves, drives average negative similarity down, which — by this exact mechanism — drives the average per-example gradient magnitude down too, a form of vanishing training signal specific to contrastive objectives once the easy cases are solved.

Mining strategies used in practice

BM25-mined hard negatives (used in DPR, Karpukhin et al., 2020): retrieve top-k lexically similar but non-relevant passages using BM25 as a cheap, independent similarity signal, decoupled from the embedding model being trained.

ANCE — Approximate Nearest Neighbour Negative Contrastive Learning (Xiong et al., 2020): periodically re-index the training corpus using the current embedding model's own vectors during training, and mine hard negatives from the model's own current nearest-neighbour mistakes. This directly targets the model's present blind spots, at the cost of requiring a re-indexing step interleaved with training — more expensive than static BM25 mining, but shown to produce measurably stronger retrievers in the original paper.

Cross-encoder-mined negatives: use a slower, more accurate cross-encoder model (see sentence-transformers for the cross-encoder-versus-bi-encoder distinction) to score and select hard negatives for training a faster bi-encoder — a form of knowledge distillation applied specifically to negative selection, common in recent retrieval training pipelines (e.g. the E5 and BGE model families' published training recipes).

The false-negative risk, formalized

Mining aggressively for hard negatives increases the risk of false negatives — candidates labelled negative that are, in fact, valid matches the mining process failed to recognise as such. Qu et al. (2021), RocketQA, identified this specifically as a limiting factor in dense passage retrieval training, and proposed denoising the negative pool using a cross-encoder to filter out likely false negatives before training, reporting measurable retrieval quality gains from this filtering step alone, holding the rest of the pipeline fixed.

Complexity of mining itself

Static BM25 mining is a one-time O(corpus size) indexing cost, then O(log(corpus size))-ish retrieval per query using an inverted index. ANCE-style dynamic mining requires periodic re-embedding and re-indexing of the full corpus during training — O(corpus size) encoder forward passes, repeated on a schedule (e.g. every few thousand training steps) — a substantial additional training cost relative to static mining, justified when the quality gain from fresher hard negatives outweighs it.

Key references

  • Karpukhin, V. et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. arXiv:2004.04906
  • Xiong, L. et al. (2020). Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. arXiv:2007.00808
  • Qu, Y. et al. (2021). RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering. arXiv:2010.08191
  • Oord, A. van den, Li, Y. & Vinyals, O. (2018). Representation Learning with Contrastive Predictive Coding. arXiv:1807.03748. The underlying InfoNCE gradient behaviour discussed above.

Current state and open problems

Hard negative mining quality is currently one of the largest sources of measurable difference between competing open-source embedding models on the MTEB leaderboard, frequently more influential than base architecture choice. The unresolved tension is between mining harder (better training signal, per the gradient argument above) and mining safer (fewer false negatives, per RocketQA's finding) — these pull in opposite directions, and the optimal balance point appears to depend on corpus characteristics in ways not yet reducible to a simple, universal rule. Curriculum-style approaches — starting training with easier negatives and progressively introducing harder ones as the model improves — are an active area of practical experimentation, echoing curriculum learning ideas from the broader machine learning literature.

What to learn next