Sentiment, Opinion and Text Mining
Zero-shot text classification
Zero-shot classification sorts text into categories it was never specifically trained on, working from the category names alone.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Zero-shot classification sorts text into categories it was never specifically trained on, using only the category names.
Picture handing someone a fruit they have never seen before. Ask them to guess: is it more like a citrus fruit, or a berry? They were never taught this exact fruit. They still reason it out, from what citrus and berries generally look and taste like.
Zero-shot text classification works the same way. Give it labels it has never specifically trained on, like "refund request" or "compliment." It reasons out the best fit anyway.
Why it exists
Normal text classification needs labelled examples for every category, sometimes hundreds of them, before it can sort anything. Gathering that data takes real time and money.
Sometimes you do not have the luxury of waiting. A new product launches, and support tickets need sorting into categories that did not exist yesterday. Zero-shot classification handles this immediately, with no training step at all.
It trades some accuracy for enormous flexibility. Change the category list, and it works instantly on the new list, no retraining required.
How it works
Text: "The delivery was four days late and the box arrived crushed."
Labels: ["billing problem", "delivery problem", "product quality", "praise"]
|
v
Score how well each label fits, without ever training on this exact task
|
v
delivery problem: 0.91 product quality: 0.06 billing problem: 0.02 praise: 0.01Under the hood, the model is not really "classifying" in the traditional sense. It checks whether the text logically supports a sentence built from each label. It picks whichever one it supports best.
Where you have already seen it
- New support categories that work the moment they are added, with no retraining delay. A direct sign of zero-shot classification running underneath.
- Content tagging tools that let you type in your own custom tags. The system was never trained on your specific tag names in advance.
- Rapid prototyping tools that classify text by any labels you type in. No dataset needed to try an idea out immediately.
- Research tools sorting documents by topics defined after the tool was built. The categories came later; the model did not need retraining to handle them.
Remember this
- Zero-shot classification needs only the category names, no labelled training examples.
- It trades some accuracy for the ability to work on brand-new categories instantly.
- It works by checking how well the text logically supports each label, not by memorised examples.
What to learn next
- Natural language inference — the exact mechanism this technique is built on.
- Few-shot text classification with SetFit — the next step up, once a handful of labelled examples exist.
- Text classification — the general task this is one flexible way to approach.
Developer — Code and libraries.
Below, the same small NLI-based model from earlier lessons scores a support ticket. It checks against four labels it was never specifically trained to recognise.
Setup
pip install transformers torchThe first run downloads MoritzLaurer/deberta-v3-xsmall-zeroshot-v1.1-all-33, about 150 MB, a compact model purpose-built for zero-shot classification tasks like this one.
Classifying with custom labels
from transformers import pipeline
classifier = pipeline(
"zero-shot-classification",
model="MoritzLaurer/deberta-v3-xsmall-zeroshot-v1.1-all-33",
)
# No training data at all -- only the raw text and the label names we invented.
ticket = "My washing machine stopped spinning halfway through the cycle."
labels = ["refund request", "appliance repair", "delivery complaint", "compliment"]
result = classifier(ticket, candidate_labels=labels)
for label, score in zip(result["labels"], result["scores"]):
print(f"{score:.3f} {label}")0.963 appliance repair 0.030 delivery complaint 0.004 compliment 0.003 refund request
Line by line
candidate_labels is a plain Python list, typed directly into the code. No dataset, no training loop, no fine-tuning step happened before this call.
Behind the scenes, each label becomes a sentence. The model checks whether the ticket text logically supports a sentence like "this example is about appliance repair." It does this for every label, then ranks them by how strongly each is supported.
appliance repair won decisively, at 0.963. The ticket describes a broken washing machine, and no other label fits nearly as well. This kind of clear-cut case is exactly where zero-shot classification shines.
The scores sum to 1.0. By default, this pipeline treats the labels as mutually exclusive alternatives. They compete for one shared probability budget, not independent yes/no questions.
Common mistakes
Writing vague or overlapping label names. "issue" and "problem" as two separate labels will confuse the model, since they mean nearly the same thing. Clear, distinct labels give sharper, more useful results.
Treating zero-shot accuracy as equal to a properly trained classifier's. Zero-shot is remarkably capable for its cost. Still, a model fine-tuned on your own labelled data will generally outperform it, given enough examples to train on.
Forgetting multi_label=True when labels are not mutually exclusive. A ticket can reasonably be both "urgent" and "billing related" at once. Without multi_label=True, the pipeline forces the scores to compete against each other and sum to 1.0.
Using single, ambiguous words as labels instead of fuller phrases. "bank" alone is ambiguous, a financial institution or a riverbank. A fuller label like "a complaint about a bank account" gives the model far more to reason with.
Try it yourself
Add multi_label=True to the classifier(...) call and rerun. Compare the new scores to the original. Notice they no longer need to sum to 1.0, since each label is now judged independently.
What to learn next
- Natural language inference — the full underlying mechanics of this technique.
- Few-shot text classification with SetFit — what to reach for once a handful of labelled examples exist.
- Reranking the top 50 — a related technique, scoring relevance rather than category fit.
Researcher — Mathematics and papers.
Zero-shot classification as entailment
The standard implementation, following Yin, Hay & Roth (2019), reframes classification as textual entailment. For text x and candidate label l, construct a hypothesis h_l from a template, such as "This example is {l}." Then score:
score(l) = P(entailment | premise = x, hypothesis = h_l)Labels are ranked by this entailment probability. With multi_label=False, the default, scores across labels are renormalised to sum to 1. This frames the task as single-label selection among the candidates. With multi_label=True, each label's entailment score is evaluated independently against a fixed "not entailment" alternative, correctly supporting non-exclusive labels.
This means zero-shot classification requires no classification head specific to the target task at all. It reuses an NLI model's existing entailment head, trained once on general-purpose NLI data. That head then handles an unbounded set of label sets it never saw during training.
Why this works: NLI as a general reasoning primitive
NLI training data, such as MNLI (Williams, Nangia & Bowman, 2018), covers a broad range of relationships between sentence pairs. It spans many topics and domains. A model trained on this general task learns something closer to general semantic reasoning, not one narrow classification task. That is what lets it transfer to arbitrary new label sets at inference time, with no additional training.
The hypothesis template matters more than it looks
The exact wording of the template measurably affects accuracy. "This example is {}." is a common default template. Domain-specific templates, such as "This customer support ticket is about {}.", often improve accuracy by matching the target text's style. This is prompt engineering, applied to classification rather than generation.
Calibration limits
Zero-shot classification scores are frequently miscalibrated. A 0.96 score is not necessarily a reliable 96% chance of correctness. A properly calibrated, task-specific classifier's score would be. Confidence tends to concentrate near label names lexically or semantically close to the input text, regardless of true task-relevant fit. Check this bias directly against labelled validation examples, before trusting scores at face value in production.
Zero-shot versus few-shot versus fine-tuned: the accuracy-effort curve
Zero-shot needs no labelled data and works instantly, at the lowest accuracy ceiling of the three. Few-shot classification with SetFit needs a handful of examples per class and meaningfully improves on zero-shot. Full fine-tuning, covered generally in fine-tuning, needs the most labelled data and delivers the highest ceiling. Choosing among the three depends on two things. How much labelled data is actually available, and how much accuracy the application genuinely needs.
Key references
- Yin, W., Hay, J. & Roth, D. (2019). Benchmarking Zero-shot Text Classification. arXiv:1909.00161
- Williams, A., Nangia, N. & Bowman, S. (2018). A Broad-Coverage Challenge Corpus for Sentence Understanding (MNLI). arXiv:1704.05426
- Laurer, M. et al. (2023). Less Annotating, More Classifying. Political Analysis — on NLI-based zero-shot classification for social science applications.
Current state and open problems
LLMs can be prompted directly with a task description and a label list. This offers a further zero-shot option beyond the NLI-entailment approach. It often reasons better on ambiguous cases, at a higher per-item cost than a small dedicated NLI model.
The open problem is calibration. No widely adopted method reliably converts a zero-shot classifier's raw score into a trustworthy probability, across arbitrary, unforeseen label sets. Every deployment still needs its own validation against real labelled examples. Only then can its scores be trusted for automated decisions.
What to learn next
- Natural language inference — the full formal treatment of the mechanism this entire lesson depends on.
- Few-shot text classification with SetFit — the next rung up the accuracy-effort ladder.
- Reliability diagrams and calibration error — measuring the calibration gap named directly above.