Sentiment, Opinion and Text Mining

Few-shot text classification with SetFit

Few-shot classification trains a working classifier from a tiny number of labelled examples per class, by building on sentence embeddings instead of starting from scratch.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Few-shot classification trains a working classifier from a tiny number of labelled examples. Sometimes as few as four or eight per category.

Picture teaching someone to sort mail into "personal" and "business" piles. You do not need to show them a thousand letters first. Show them four or five clear examples of each, and most people pick up the pattern immediately.

Few-shot text classification aims for that same efficiency. Traditional classifiers often need hundreds or thousands of labelled examples. Few-shot methods aim to work with a mere handful, by starting from a model that already understands language well.

Why it exists

Zero-shot classification needs no labelled examples at all, but its accuracy has a ceiling. Full fine-tuning gets the best accuracy, but needs a large labelled dataset most teams do not have on day one.

Few-shot classification sits deliberately in between. A handful of labelled examples, often gathered in an afternoon, can noticeably beat zero-shot accuracy. It skips the weeks of labelling a full fine-tune would need.

This matters most in exactly the situation many real projects start in. A new classification task, a new set of categories, and only a small starting batch of labelled examples.

How it works

4 labelled examples per class (8 total)
              |
              v
Turn each example into a sentence embedding (a list of numbers capturing meaning)
              |
              v
Train a small, simple classifier on top of those 8 embeddings
              |
              v
Classifier ready — apply it to new, unseen text

The heavy lifting was already done by whatever model produces the embeddings. That model trained long before, on a huge amount of general text. The few-shot step only has to learn a lightweight boundary on top of embeddings that already capture meaning well.

Where you have already seen it

  • A support team starting a new ticket category with only a dozen labelled examples. Working accuracy from day one, refined as more labels arrive.
  • Content moderation tools bootstrapped quickly for a new, specific policy. No time to wait for a large labelled dataset before launch.
  • Research tools classifying a narrow, specialised topic with few available examples. Sometimes only a handful of labelled cases exist at all.
  • Startups validating a classification idea before committing to full-scale data labelling. Few-shot results decide whether the bigger investment is worth making.

Remember this

  • Few-shot classification needs only a handful of labelled examples per category, not hundreds.
  • It works by reusing a pretrained embedding model, then training a small classifier on top.
  • It sits between zero-shot, which needs no examples, and full fine-tuning, which needs many.

What to learn next

Developer — Code and libraries.

The library most associated with this technique is called SetFit. At the time of writing, it pins an older version of transformers and fails to import against a current install. Below is what SetFit does internally, written directly: embed each example, then train a simple classifier on top. This runs on any current setup, and doubles as a clear look under the hood.

Setup

bash
pip install transformers torch scikit-learn

Training a classifier on eight examples

fewshot.py
import torch
from transformers import AutoTokenizer, AutoModel
from sklearn.linear_model import LogisticRegression

tokenizer = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")

def embed(texts):
    tokens = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
    with torch.no_grad():
        output = model(**tokens)
    mask = tokens["attention_mask"].unsqueeze(-1)
    summed = (output.last_hidden_state * mask).sum(1)
    counts = mask.sum(1).clamp(min=1e-9)
    return (summed / counts).numpy()

# 4 labelled examples per class. This is the entire training set.
train_texts = [
    "My card was charged twice for the same order.",
    "I want a refund, the invoice amount looks wrong.",
    "The subscription fee increased without any warning.",
    "Please cancel my payment plan and refund this month.",
    "The app crashes every time I open the settings page.",
    "I can't log in, it keeps saying invalid password.",
    "Video calls freeze after about two minutes for me.",
    "The export button does nothing when I click it.",
]
train_labels = ["billing", "billing", "billing", "billing",
                "technical", "technical", "technical", "technical"]

clf = LogisticRegression(max_iter=1000)
clf.fit(embed(train_texts), train_labels)

test_texts = [
    "You billed me for a plan I already cancelled last week.",
    "The app keeps freezing whenever I try to upload a photo.",
]
predictions = clf.predict(embed(test_texts))

for text, label in zip(test_texts, predictions):
    print(f"{label:10s} <- {text}")
Output
billing    <- You billed me for a plan I already cancelled last week.
technical  <- The app keeps freezing whenever I try to upload a photo.

Line by line

embed() is the same mean-pooling function used in earlier lessons. It turns each sentence into a fixed-size vector. This is the "understanding language" part, and none of it was trained here.

LogisticRegression().fit(...) is the entire "few-shot" training step. Eight embeddings in, eight labels in, and a working linear boundary comes out. This is genuinely fast: milliseconds, not hours.

Both test predictions were correct, and neither test sentence closely resembles its matching training examples word-for-word. "You billed me for a plan I already cancelled" shares almost no exact wording with any billing example. Yet the embedding space placed it close enough to the billing cluster to classify correctly.

This is what makes few-shot work at all. The embeddings already encode meaning, not only surface wording, from their original large-scale training. Eight examples only have to teach the boundary between two already-well-separated regions of that space.

Common mistakes

Trying pip install setfit and hitting an import error. At the time of writing, the setfit package pins an older transformers version. It fails with an ImportError on a current install. Either pin a compatible transformers version specifically for that environment. Or use the direct embed-plus-classifier approach shown above, which needs no extra package at all.

Using too few examples per class, or examples that are too similar to each other. Four near-identical billing examples teach the model less than four genuinely varied ones. Diversity in the small training set matters more than raw count.

Expecting few-shot accuracy to match a fully fine-tuned model. It will not, reliably. Few-shot is a strong starting point, not a finish line. Collect more labelled data as it becomes available, and retrain.

Forgetting to check class balance. Four examples of one class and forty of another will bias LogisticRegression toward the larger class. Keep the few-shot set balanced across categories, even when balance is easy to lose track of at such small scale.

Try it yourself

Add a third category, "account access", with four new examples about locked accounts or forgotten passwords, and retrain. Test whether the small classifier separates all three categories cleanly, or whether "technical" and "account access" start to blur together.

What to learn next

Researcher — Mathematics and papers.

The SetFit two-stage training procedure

SetFit (Tunstall et al., 2022) formalises the pattern shown in the developer block into two explicit stages. Stage one fine-tunes a sentence-transformer using contrastive learning. Pairs of examples from the same class get pulled together in embedding space. Pairs from different classes get pushed apart, generated from the small labelled set by sampling positive and negative pairs. Stage two trains a lightweight classification head, typically logistic regression, on the resulting fine-tuned embeddings.

The developer block's code implements only an approximation of stage two, using off-the-shelf embeddings without stage one's contrastive fine-tuning step. Full SetFit's stage one measurably improves few-shot accuracy over this approximation. It adapts the embedding space itself to the specific classes at hand. It does not only adjust the boundary drawn on top.

Why contrastive fine-tuning helps with so little data

Standard classifier fine-tuning needs one gradient update per labelled example. Contrastive fine-tuning instead generates a combinatorial number of training pairs from the same small labelled set. With n examples, there are O(n^2) possible pairs, each usable as a contrastive training signal. This is the core trick behind SetFit's data efficiency. It multiplies a small labelled set's effective training signal, with no more raw labels needed.

Comparison to prompt-based few-shot learning

A different family of methods works with no training step at all. In-context few-shot learning (Brown et al., 2020) places labelled examples directly in an LLM's prompt at inference time. It re-pays the cost of processing those examples on every single call, though. SetFit instead pays a one-time training cost, then runs cheaply per prediction afterward.

Sample efficiency, measured

Tunstall et al. (2022) report SetFit reaching accuracy competitive with full fine-tuning, on several benchmarks, using as few as 8 examples per class. It substantially outperforms standard fine-tuning of a comparably sized model under the same tiny data budget. The gap narrows as available labelled data grows, and standard fine-tuning eventually overtakes few-shot methods once enough labelled data exists.

Where few-shot methods break down

Performance degrades on tasks requiring fine distinctions the base embedding model was never trained to separate well. Subtle sentiment gradations, or technical jargon absent from its pretraining data, are typical cases. Domain-adaptive pretraining, covered in continued pretraining on your own domain, addresses this directly. But it reintroduces exactly the data and compute cost few-shot methods exist to avoid.

Key references

  • Tunstall, L. et al. (2022). Efficient Few-Shot Learning Without Prompts (SetFit). arXiv:2209.11055
  • Reimers, N. & Gurevych, I. (2019). Sentence-BERT. arXiv:1908.10084
  • Brown, T. et al. (2020). Language Models are Few-Shot Learners. arXiv:2005.14165

Current state and open problems

Both SetFit-style embedding-based few-shot learning and LLM in-context few-shot learning remain in active, simultaneous use. The choice comes down to latency and cost, not one method outright winning. SetFit-style methods win where per-prediction cost and latency matter, once trained. In-context methods win instead where the label set changes constantly and retraining overhead is unacceptable.

The open problem is exactly what the developer block surfaced as a real, current issue. The machine learning ecosystem moves fast. Today's convenient library can quietly break on tomorrow's dependency versions, especially for smaller, less actively maintained packages. Understanding what a library like SetFit does internally is what lets a practitioner recover when it breaks. Knowing only how to call it is not enough.

What to learn next