Fine-tuning an embedding model on your own data
Fine-tuning teaches a general embedding model your specific vocabulary, by showing it pairs of text that should be considered similar until it starts agreeing with you.
- 12 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Fine-tuning teaches a general embedding model your own vocabulary. It works by showing examples of text that should count as similar, until the model agrees.
Think about a new employee, fluent in English but new to your company. They understand normal conversation perfectly. But they do not yet know that "the blue form" and "the intake request" mean the same document. Give them a week around your team's own language, and they pick it up.
A pretrained sentence-transformer is that fluent new employee. Fine-tuning is the week of exposure. It shows the model pairs of text from your domain that should count as similar. Over time, its vector space starts to reflect your vocabulary, not only general English.
Why it exists
A general-purpose embedding model trains on broad text — Wikipedia, books, forums, news. It has never seen your support ticket abbreviations, or your product's internal nicknames. It does not know how your users actually phrase a problem.
Ask it whether "payment failed" and "transaction declined" mean the same thing. A general model usually does well — these are common phrases. Ask it about your own jargon, and it can miss connections a teammate would catch instantly.
Fine-tuning closes that gap. Training an embedding model from nothing needs enormous amounts of data. Instead, you nudge an already-capable model toward your domain, using far less data and far less compute.
How it works
Show the model pairs of text that should be considered similar, over and over. Adjust its numbers slightly each time, so similar pairs end up closer together.
Training pair: "payment failed" <-> "transaction declined"
Before fine-tuning: vectors sit a moderate distance apart
After many examples: the model nudges them closer together
Repeat across many domain-specific pairs, and the whole
vector space reshapes slightly around your vocabulary.Crucially, the model is not only told "these two are close" in isolation. It is also shown other, unrelated text in the same batch. That way it learns to push similar pairs together and keep everything else apart. This two-sided pressure is what reshapes the space usefully, instead of collapsing everything together.
Where you have already seen it
- Company-specific chatbots and internal search, fine-tuned on that company's own documents and support history.
- Legal and medical AI tools, fine-tuned on domain-specific terminology general models handle poorly out of the box.
- E-commerce search, fine-tuned on a specific catalogue's product names and abbreviations.
Remember this
- Fine-tuning nudges a general embedding model toward your specific domain, using far less data than training from nothing.
- It works by showing the model pairs of text that should be similar, alongside other text that should not.
- The result is a vector space still built on general language understanding, but reshaped around your vocabulary.
What to learn next
- Hard negative mining — choosing better "unrelated" examples to make this training far more effective.
- Sentence-transformers — the base model architecture usually being fine-tuned.
- Fine-tuning — the broader picture of adapting a pretrained model to new data.
Developer — Code and libraries.
A real, small contrastive fine-tuning loop — training an embedding model on a handful of domain pairs, and measuring whether it actually moved a held-out pair the model never saw during training.
Setup
pip install torch transformersFine-tuning on a tiny domain-specific dataset
import time
import torch
import torch.nn.functional as F
from torch.optim import AdamW
from transformers import AutoTokenizer, AutoModel
torch.manual_seed(0)
tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
# domain jargon pairs: (support ticket phrase, its plain-English meaning)
# a general-purpose model has never specifically been trained on this pairing
pairs = [
("payment failed", "transaction declined"),
("cannot login", "unable to sign in"),
("app crashes on startup", "app closes immediately after opening"),
("refund not received", "money not credited back"),
("otp not received", "verification code not delivered"),
("order stuck in processing", "order status not updating"),
]
# held out — never shown to the model during fine-tuning
test_pair = ("payment declined at checkout", "card was rejected while paying")
def encode(sentences):
enc = tok(list(sentences), padding=True, return_tensors="pt")
out = model(**enc).last_hidden_state
mask = enc["attention_mask"].unsqueeze(-1).float()
pooled = (out * mask).sum(1) / mask.sum(1).clamp(min=1e-9)
return F.normalize(pooled, dim=1)
with torch.no_grad():
before = F.cosine_similarity(*encode(test_pair), dim=0).item()
optimizer = AdamW(model.parameters(), lr=2e-5)
anchors = [a for a, _ in pairs]
positives = [p for _, p in pairs]
t0 = time.time()
for epoch in range(8):
optimizer.zero_grad()
a_vecs = encode(anchors)
p_vecs = encode(positives)
# in-batch negatives: every OTHER positive in the batch acts as a negative
logits = a_vecs @ p_vecs.T * 20.0 # temperature-scaled cosine similarities
targets = torch.arange(len(pairs)) # the correct positive is on the diagonal
loss = F.cross_entropy(logits, targets)
loss.backward()
optimizer.step()
if epoch % 2 == 0:
print(f"epoch {epoch} loss={loss.item():.4f}")
print(f"training time: {time.time() - t0:.1f}s")
with torch.no_grad():
after = F.cosine_similarity(*encode(test_pair), dim=0).item()
print()
print(f"held-out pair: {test_pair}")
print(f"cosine similarity before fine-tuning: {before:.3f}")
print(f"cosine similarity after fine-tuning: {after:.3f}")epoch 0 loss=0.4323
epoch 2 loss=0.0445
epoch 4 loss=0.0149
epoch 6 loss=0.0046
training time: 1.1s
held-out pair: ('payment declined at checkout', 'card was rejected while paying')
cosine similarity before fine-tuning: 0.611
cosine similarity after fine-tuning: 0.631Training time depends on your CPU; expect roughly the same order of magnitude, a few seconds at most, since this is only 6 pairs and 8 short training steps.
Line by line
a_vecs @ p_vecs.T * 20.0 builds a 6-by-6 grid of similarity scores between every anchor and every positive in the batch. targets = torch.arange(len(pairs)) says the correct match for anchor i is positive i — the diagonal of that grid. F.cross_entropy then pushes each anchor's own positive higher than every other positive in the batch, which is exactly what makes the other 5 positives act as free negative examples, with no separate negative dataset needed. This exact technique is called in-batch negatives, or a multiple negatives ranking loss.
The training loss collapses fast — 0.43 down to 0.005 in 8 steps — because 6 pairs is a tiny dataset the model can essentially memorise. This is the honest, important detail: near-zero training loss does not mean the model generalised well. It means it fit these 6 examples almost exactly.
The held-out pair only moved from 0.611 to 0.631. That is a real, measurable improvement — but a modest one, especially set against how completely the training loss collapsed. This gap between "training loss looks perfect" and "held-out improvement is modest" is the whole point of testing on a pair the model never saw.
Common mistakes
Judging fine-tuning success from training loss alone. As shown directly above, a training loss near zero on 6 pairs says almost nothing about how the model will behave on the thousandth new example it meets in production. Always hold out real examples and measure them separately.
Fine-tuning on far too little data and expecting production-grade results. Six pairs is enough to demonstrate the mechanism correctly, which is this lesson's goal. Real fine-tuning runs for a production system typically use many hundreds to many thousands of pairs — public benchmark training sets like MS MARCO use hundreds of thousands.
Forgetting to freeze anything, and overfitting the whole network on tiny data. The script above fine-tunes every parameter in the model. On a genuinely small dataset, this risks damaging the model's general-purpose ability to handle text unrelated to your domain — a known trade-off, usually managed with a lower learning rate, fewer epochs, or freezing early layers.
Try it yourself
Add three or four more domain-specific pairs of your own choosing to pairs, covering a topic unrelated to the payment/login examples already there — this diversifies what the in-batch negatives look like. Re-run and compare the held-out pair's improvement to the result above.
What to learn next
- Hard negative mining — replacing these random in-batch negatives with far more useful, deliberately chosen ones.
- Sentence-transformers — the base architecture this fine-tuning loop is adapting.
- LoRA — a lighter-weight way to fine-tune large models without updating every parameter.
Researcher — Mathematics and papers.
The objective: multiple negatives ranking loss
The training loop above implements Multiple Negatives Ranking Loss (Henderson et al., 2017), also known as in-batch negative sampling or the batch-softmax form of contrastive learning:
L = -(1/N) * sum over i=1..N of log( exp(sim(a_i, p_i) / tau) / sum over j=1..N of exp(sim(a_i, p_j) / tau) )a_i,p_iare thei-th anchor and its true positive;Nis the batch size.sim(x, y)is cosine similarity between the two pooled, normalized embeddings.tauis a temperature parameter (1/20 = 0.05in the developer block, via the* 20.0scaling) controlling how sharply the softmax distinguishes close scores — a lower temperature makes the model penalise near-misses more aggressively.
This is a direct application of the InfoNCE objective (Oord, Li & Vinyals, 2018), originally developed for self-supervised representation learning in the contrastive predictive coding framework, and now the standard training objective across essentially all modern dense retrieval and sentence embedding models.
Why in-batch negatives are efficient
For a batch of N pairs, in-batch negative sampling produces N * (N-1) negative comparisons using only N forward passes — every other example's positive doubles as a negative for free, at zero extra encoding cost. This is why batch size matters directly for training quality in this setup: a larger batch supplies more negatives per gradient step, at the cost of more memory. Large-scale retrieval model training (e.g. Karpukhin et al.'s DPR, Wang et al.'s E5 family) explicitly reports batch size as one of the most influential hyperparameters for exactly this reason.
Full fine-tuning versus parameter-efficient adaptation
The script above updates every parameter in the encoder — full fine-tuning. At larger scale, this is often replaced with parameter-efficient methods such as LoRA (Hu et al., 2021), which inserts small trainable low-rank matrices into the frozen base model, updating a small fraction of total parameters. For embedding model adaptation specifically, full fine-tuning remains common at the scale of the base models typically used (under ~500M parameters), since the cost difference is smaller than for large generative language models, where LoRA's savings are more decisive.
Catastrophic forgetting
Fine-tuning every parameter on a narrow, small dataset risks catastrophic forgetting — degrading the model's general-domain performance while improving its narrow-domain performance, since gradient updates driven entirely by domain-specific pairs can overwrite representations useful for text the fine-tuning set never covered. Standard mitigations: a low learning rate (2e-5 in the developer block is already conservative by design), few epochs, mixing in some general-domain pairs alongside domain-specific ones, or freezing early encoder layers and fine-tuning only the later layers plus pooling.
Data scale in practice
Public sentence embedding training sets used for major open models are dramatically larger than the toy example above: the all-MiniLM-L6-v2 model itself was trained on over a billion sentence pairs, drawn from more than 30 different data sources (Reimers et al., sentence-transformers training data documentation). Domain-specific fine-tuning on top of such a base model typically needs far less — often thousands to low tens of thousands of pairs are enough to shift a strong base model meaningfully toward a new domain, precisely because the base model's general language understanding does not need to be relearned, only adjusted.
Key references
- Henderson, M. et al. (2017). Efficient Natural Language Response Suggestion for Smart Reply. arXiv:1705.00652. Introduces the in-batch negative / multiple negatives ranking loss used in the developer block.
- Oord, A. van den, Li, Y. & Vinyals, O. (2018). Representation Learning with Contrastive Predictive Coding. arXiv:1807.03748. The InfoNCE objective this loss is a special case of.
- Reimers, N. & Gurevych, I. (2019). Sentence-BERT. arXiv:1908.10084
- Hu, E. et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685
- Karpukhin, V. et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. arXiv:2004.04906. Large-scale application of this exact training pattern to retrieval.
Current state and open problems
The remaining hard problem in embedding fine-tuning is not the loss function — in-batch InfoNCE-style contrastive training is now standard and well understood — but data quality: how positive pairs are mined or labelled, and crucially, how negative examples are chosen. Random in-batch negatives, as used in the developer block, are frequently too easy once a model has learned the coarse structure of a domain, providing diminishing training signal as the model improves. This is precisely the gap hard negative mining addresses, and it is currently where most of the measurable quality difference between competing production embedding models actually comes from, more so than architecture differences.
What to learn next
- Hard negative mining — fixing the weakest part of the training loop shown here.
- LoRA — parameter-efficient fine-tuning, for adapting larger models at lower cost.
- Pooling strategies — the pooling step this whole training loop optimizes through.