Sentiment, Opinion and Text Mining
Stance detection
Stance detection asks whether a piece of text is for, against or neutral on a specific claim, which is a different question from whether the text sounds positive or negative.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Stance detection asks whether a piece of text is for, against or neutral on a specific claim.
Picture a heated dinner table debate about a new policy. Some people argue calmly for it. Some argue angrily against it. Tone tells you how heated they are. Stance tells you which side they are on.
That is the exact gap plain sentiment analysis misses. A calm, politely worded sentence can still be firmly against something. Stance detection is built to catch that, separately from tone.
Why it exists
Sentiment analysis answers "does this sound positive or negative?" Stance detection answers a different question entirely: "is this for or against a specific claim?"
The two do not always line up. "Remote work destroys team culture" is a negative-sounding sentence, and it is against remote work, so here they agree. But consider: "I miss the office, though I'll admit remote work has been great for my focus." It sounds mostly positive, while arguably favouring the office.
Public debate, policy research and social media analysis all need this distinction. Knowing 60% of comments sound negative tells you far less than knowing 60% are against a specific proposal.
How it works
Claim: "Remote work should be the default for office jobs."
"I've been way more productive since we stopped commuting." -> favour
"Nothing beats solving a hard problem in the same room." -> against
"Our team switched to three days in office, two from home." -> neutral (states a fact, no opinion)Stance detection needs both the text and the claim it is being measured against. The same sentence can support one claim and oppose a differently worded one.
Where you have already seen it
- Fact-checking tools that classify a claim as "supports," "refutes" or "unrelated." A close cousin of stance detection, built on the same idea.
- Political and policy sentiment trackers around an election or a bill. Measuring where public comments actually stand, not only their tone.
- Debate and argument-mining tools used in research. Sorting arguments by side before summarising each one.
- Brand monitoring during a controversy. Distinguishing customers who are against a specific decision from ones who are only unhappy in general.
Remember this
- Stance detection measures for-or-against on a specific claim, not overall tone.
- A calm sentence can be firmly against something; an angry one can be for it.
- The claim being tested has to be stated explicitly; stance is always relative to something.
What to learn next
- Natural language inference — the mechanism this lesson's technique is built directly on.
- Zero-shot text classification — the same underlying method, applied more generally.
- Toxicity and abuse detection — a different, narrower kind of text judgement.
Developer — Code and libraries.
Below, the same zero-shot model from earlier lessons scores each message. Does it favour, oppose, or stay neutral on a specific claim?
Setup
pip install transformers torchScoring stance toward one claim
from transformers import pipeline
classifier = pipeline(
"zero-shot-classification",
model="MoritzLaurer/deberta-v3-xsmall-zeroshot-v1.1-all-33",
)
hypothesis_template = "This text is in {} of remote work being the default for office jobs."
labels = ["favor", "against", "neutral"]
posts = [
"I have been way more productive since we stopped commuting five days a week.",
"Nothing beats being in the same room to actually solve a hard problem fast.",
"Our team switched to three days in office and two from home this month.",
]
for post in posts:
result = classifier(post, candidate_labels=labels, hypothesis_template=hypothesis_template)
top_label, top_score = result["labels"][0], result["scores"][0]
print(f"{top_label:8s} ({top_score:.3f}) {post}")against (0.460) I have been way more productive since we stopped commuting five days a week. against (0.680) Nothing beats being in the same room to actually solve a hard problem fast. neutral (0.491) Our team switched to three days in office and two from home this month.
Line by line
hypothesis_template turns each label into a full sentence: "This text is in favor of remote work being the default..." The model checks whether the original post supports that constructed sentence. This is the same entailment mechanism used for zero-shot classification in earlier lessons.
Read the results honestly, including the wrong one. The second and third results look reasonable. The first one does not. A post celebrating not commuting reads, to a human, as plainly favouring remote work. The model called it "against," at a fairly low 0.460 confidence.
This is a genuine model mistake, left in on purpose. The post never uses words like "remote work" or "default" directly. It only implies its position through "stopped commuting" and "more productive," and the model's reasoning did not connect the dots. Low confidence, 0.460, was at least an honest signal that it was not sure.
Common mistakes
Trusting a low-confidence result as if it were a high-confidence one. The lowest score in the batch above, 0.460, was also the wrong answer. Confidence is not a hard guarantee. Still, a low score is a real signal to treat the label with more suspicion.
Writing a vague hypothesis template. A vague template like "This text is about remote work" gives the model far less to work with. A template spelling out a clear position, like the one used above, works better. Specific, unambiguous templates give noticeably better results.
Assuming stance is symmetric across differently worded claims. A post's stance on one claim says nothing about its stance on a different claim. "The default" and "should be banned" are different claims. Re-run detection separately for each one you care about.
Confusing "neutral" with "unrelated." A post that only states a fact, with no clear position, is genuinely neutral. A post about something else entirely is a different case, and some pipelines conflate the two without you noticing.
Try it yourself
Rewrite the first post to state its position explicitly. Try: "Remote work should absolutely be the default, I've never been more productive." Compare the new stance score against the original, more implicit version.
What to learn next
- Natural language inference — the entailment mechanism underneath this lesson's whole approach.
- Zero-shot text classification — the general version of the technique used here.
- Reliability diagrams and calibration error — understanding why 0.460 was a meaningful warning sign, not noise.
Researcher — Mathematics and papers.
Stance detection as a relation between two texts
Formally, stance detection predicts P(label | text, target), where target is an explicit claim or entity. Ordinary sentiment classification predicts only P(label | text). This target-conditioning is the task's defining structural feature. The same text can receive different labels for different target values, as seen directly in the developer block's hypothesis template.
Mohammad et al. (2016) introduce the SemEval-2016 Task 6 stance dataset. It formalises the standard three-way label scheme: favor, against, neutral, sometimes extended with "unrelated." This remains the dominant framing in the field.
Target-specific versus cross-target stance
Most early stance datasets train and test on the same fixed set of targets. Cross-target stance detection is evaluated on targets unseen during training. It is measurably harder, since target-specific vocabulary and framing do not transfer directly. Augenstein et al. (2016) show conditional encoding helps here. Jointly encoding text and target, rather than encoding each separately, meaningfully improves cross-target generalisation.
Why NLI-based zero-shot stance detection works, and where it fails
Reframing stance as textual entailment, as in the developer block, lets a general-purpose NLI model handle arbitrary, unseen targets. No target-specific training data is needed. The hypothesis_template construction converts a label into a natural-language claim. Stance then becomes a simple question: does the post entail, contradict, or stay neutral toward that claim?
This approach's known failure mode is exactly what the developer block surfaced. A post can support a position without discussing it directly, an implicit stance. Reading that requires real-world inference, which general NLI models handle inconsistently. Explicit stance, where a post states its position directly, performs considerably better under this same method.
Sarcasm and stance interact badly
A post can be sarcastically "in favor" of a position it actually opposes. That compounds sarcasm detection's own difficulty on top of stance detection's. Neither task's standard framing accounts for the other directly. Production systems handling both typically run sarcasm detection first, as a filter, before trusting a stance label.
Evaluation
Standard stance detection is scored with macro-averaged F1, across the favor and against classes specifically. Neutral is commonly excluded from the average, following the SemEval-2016 convention. It is both the most common class in most real data, and the least informative one for downstream use.
Key references
- Mohammad, S. et al. (2016). SemEval-2016 Task 6: Detecting Stance in Tweets. SemEval.
- Augenstein, I. et al. (2016). Stance Detection with Bidirectional Conditional Encoding. arXiv:1606.05464
- Küçük, D. & Can, F. (2020). Stance Detection: A Survey. ACM Computing Surveys.
Current state and open problems
LLMs can be prompted directly with a claim and a text. This now handles stance detection competitively, including reasonable cross-target generalisation, with no task-specific training data at all.
The open problem is exactly the failure the developer block demonstrated: implicit stance inference. It requires real-world knowledge about what a stated fact implies about an unstated position. No current benchmark isolates this failure mode cleanly. Most stance datasets are drawn from naturally opinionated text. Stance there is stated more directly than in this lesson's deliberately implicit example.
What to learn next
- Natural language inference — the full formal mechanics behind this lesson's zero-shot approach.
- Sarcasm and irony detection — the compounding failure mode discussed above.
- Zero-shot text classification — the general technique, beyond the stance-specific framing here.