Natural language inference
Natural language inference checks whether one sentence entails, contradicts, or says nothing about another, and models trained for it become evaluators for other tasks.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Natural language inference checks if one sentence must be true, must be false, or is unrelated, given another sentence.
Someone tells you "it is raining outside". You instantly know "the ground is wet" is probably true. You know "the sun is shining brightly" is probably false. You cannot tell if "traffic will be heavy" is true or false at all.
Those three reactions have names: entailment (must be true), contradiction (must be false) and neutral (cannot tell). That three-way judgement is natural language inference, often shortened to NLI.
Why it exists
Early NLP models were good at spotting matching words. They were bad at reasoning about what a sentence actually implies. "A dog is sleeping" and "no dog is sleeping" share almost every word, yet mean opposite things.
NLI was built as a stress test for exactly that gap. A model cannot pass an NLI test by matching words. It has to work out the logical relationship between two sentences.
That test turned out to be useful for more than testing. Models trained on NLI became reusable tools for checking whether one piece of text is supported by another. That skill is now central to catching AI-generated misinformation.
How it works
Premise: "A man is playing guitar on stage."
Hypothesis: "A person is performing music."
-> ENTAILMENT (this must be true, given the premise)
Hypothesis: "The stage is empty."
-> CONTRADICTION (this cannot be true, given the premise)
Hypothesis: "The man is wearing a red hat."
-> NEUTRAL (the premise never says either way)A model trained for NLI reads both sentences together and picks one of the three labels. It is not checking for shared words. It is checking whether the second sentence logically follows from the first.
Where you have already seen it
- Fact-checking tools. Checking if a claim is supported, contradicted, or unaddressed by a trusted source document.
- Chatbot guardrails. Verifying that a chatbot's answer is actually backed by the document it was supposed to read, not invented.
- Summarisation quality checks. Testing if every sentence in a summary is entailed by the original article, catching invented details.
Remember this
- NLI has three outcomes: entailment (must be true), contradiction (must be false), neutral (cannot tell).
- It tests logical relationship, not word overlap, which word-counting metrics cannot do.
- Models trained on NLI get reused as general-purpose "does this text support that claim?" checkers.
What to learn next
- Text classification — the general technique NLI models are a specific case of.
- BERT — the model family most NLI classifiers are built on.
- Hallucination — the problem NLI-based fact-checking tools are built to catch.
Developer — Code and libraries.
This uses a small distilbert model already fine-tuned for NLI, so no training is needed. It downloads once, then runs on CPU.
Setup
pip install transformers torchTesting all three outcomes
from transformers import pipeline
classifier = pipeline("text-classification", model="typeform/distilbert-base-uncased-mnli")
pairs = [
("A man is playing guitar on stage.", "A person is performing music."),
("A man is playing guitar on stage.", "The stage is empty."),
("A man is playing guitar on stage.", "The man is wearing a red hat."),
]
for premise, hypothesis in pairs:
result = classifier(f"{premise} [SEP] {hypothesis}")[0]
print(f"label={result['label']:14} score={result['score']:.3f} hypothesis={hypothesis!r}")The first run downloads the model, about 270MB. It prints a short loading report; only the three result lines below matter.
label=ENTAILMENT score=0.999 hypothesis='A person is performing music.' label=CONTRADICTION score=1.000 hypothesis='The stage is empty.' label=CONTRADICTION score=0.997 hypothesis='The man is wearing a red hat.'
The first two match what you would expect. The third does not. "The man is wearing a red hat" was never confirmed or denied by the premise. The correct label is neutral, but the model called it a confident contradiction.
Line by line
f"{premise} [SEP] {hypothesis}" joins the two sentences with the model's separator token. NLI models take a pair of sentences, not one, and this is how that pair gets packed into a single input string for this particular checkpoint.
pipeline("text-classification", ...) loads both the tokenizer and the trained classification head in one call. The model outputs one of three labels, plus a confidence score for that label.
Common mistakes
Trusting a high confidence score without checking the label makes sense. As the third example shows, the model was 99.7% confident, and wrong. A high score means the model is sure, never that the model is right.
Confusing neutral with contradiction. They are different failure modes for a fact-checking system. Neutral means "not addressed", contradiction means "actively wrong". Conflating them produces misleading fact-checks.
Using an English-only NLI model on other languages. Most small NLI checkpoints, including this one, were fine-tuned only on English data. Running them on Hindi or Tamil text gives unreliable, often silently wrong, labels.
Try it yourself
Rewrite the third hypothesis to something the premise genuinely rules out, for example "No one is on the stage.". Rerun the script.
That should score confidently as contradiction, correctly this time. Compare it against the red-hat case above, and notice how similar the model's confidence looks for a genuine contradiction and a mistaken one. That gap is exactly why the honesty rule matters here: confidence alone cannot tell you which is which.
What to learn next
- BERT — the encoder architecture behind most NLI classifiers.
- Using an LLM to grade text — a heavier alternative when three fixed labels are not enough.
- What is RAG — a system design that leans on NLI-style checking to catch ungrounded answers.
Researcher — Mathematics and papers.
Task formulation
NLI is a three-way classification problem over sentence pairs:
f(premise, hypothesis) -> {entailment, contradiction, neutral}Framed this way, NLI reduces to standard supervised classification: encode the pair, feed the representation to a softmax head over three classes, train with cross-entropy against human-annotated labels.
The interesting part is the label definition itself. Entailment is defined as: a human reading the premise would judge the hypothesis as most likely true. This is deliberately a human plausibility judgement, not strict formal logical entailment, which makes the datasets tractable to annotate at scale but also introduces the annotation artefacts discussed below.
Datasets
SNLI (Bowman et al., 2015) used image captions as premises, with crowdworkers writing entailed, contradicted and neutral hypotheses for each. 570k pairs, a single narrow domain (photo captions).
MultiNLI (Williams, Nangia & Bowman, 2018) extended this across ten genres of written and spoken English, testing whether models generalise beyond photo captions specifically. It is the standard NLI training set today, and the one behind the checkpoint used in the developer block.
ANLI (Nie et al., 2020) took a different approach: humans adversarially wrote examples specifically designed to fool a current best model, iterated over several rounds. This produces a harder, less exploitable benchmark than static crowdsourced data.
Known annotation artefacts
Gururangan et al. (2018), Annotation Artifacts in Natural Language Inference Data, showed something striking. A classifier reading only the hypothesis, with no access to the premise, could still beat chance on SNLI and MultiNLI. Certain words correlate with certain labels regardless of context: negation in the hypothesis with contradiction, vague generalisations with neutral.
This means NLI benchmark accuracy alone can overstate genuine reasoning ability, since a model can partially solve the task using surface statistics of the hypothesis. ANLI's adversarial collection process was designed specifically to reduce this exploit.
NLI as a general-purpose tool
Because entailment checking is exactly "is claim B supported by text A", trained NLI models get reused directly for:
- Zero-shot classification: reframe an arbitrary label as a hypothesis ("this text is about sports"), and use entailment probability as the classification score, an approach popularised by Yin, Hay & Roth (2019).
- Factual consistency checking in summarisation, treating the source document as premise and each summary sentence as hypothesis.
- Retrieval-augmented generation verification, checking whether a generated answer is entailed by retrieved passages before it is shown to a user.
Key references
- Bowman, S., Angeli, G., Potts, C. & Manning, C. (2015). A Large Annotated Corpus for Learning Natural Language Inference. EMNLP (SNLI).
- Williams, A., Nangia, N. & Bowman, S. (2018). A Broad-Coverage Challenge Corpus for Sentence Understanding. arXiv:1704.05426 (MultiNLI).
- Gururangan, S. et al. (2018). Annotation Artifacts in Natural Language Inference Data. arXiv:1803.02324
- Nie, Y. et al. (2020). Adversarial NLI: A New Benchmark for Natural Language Understanding. arXiv:1910.14599
- Yin, W., Hay, J. & Roth, D. (2019). Benchmarking Zero-shot Text Classification. arXiv:1909.00161
Current state and open problems
Dedicated NLI classifiers remain in use as a lightweight, cheap-to-run building block. This matters most for factual consistency checking, where running a full LLM judge on every sentence would cost far more.
The open problem is the one the developer block's red-hat example exposed directly. NLI models stay overconfident on genuinely neutral cases, defaulting to entailment or contradiction instead of admitting uncertainty. Calibrating these scores, so a 99% confidence is actually right 99% of the time, remains unresolved for anyone using them as an automated filter.
What to learn next
- BERTScore — a different way to compare meaning, without a fixed three-way label.
- Reliability diagrams and calibration error — measuring exactly the overconfidence problem raised above.
- Legal RAG and fabricated citations — a real-world use of entailment-style checking to catch unsupported claims.