Sentiment, Opinion and Text Mining
Emotion classification
Emotion classification labels text with a specific feeling, like joy, anger or surprise, instead of only positive or negative.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Emotion classification labels text with a specific feeling, like joy, anger or surprise, not only positive or negative.
Picture two messages: "I won the lottery!" and "I finally got the job." Both are strongly positive. They do not feel the same at all. One is shock, the other is pride.
Sentiment analysis would call both "positive" and stop there. Emotion classification goes further. It asks which feeling, out of a set like joy, anger, fear, sadness and surprise.
Why it exists
"Positive" and "negative" are useful, but thin. Anger and sadness are both negative. Yet they call for completely different responses, from a business or from a person reading them.
A support ticket written in anger needs a different reply than one written in sadness or confusion. Sentiment analysis cannot tell these apart. Emotion classification can.
This matters anywhere the specific feeling changes what happens next. Mental health apps, support routing, and social media monitoring during a public event all qualify.
How it works
"I can't believe I got the internship!" -> excitement, surprise
"The delivery was left outside again." -> neutral (mild annoyance, unstated)An emotion model is trained on text labelled with specific feelings, not only positive or negative. It learns the words, phrasing and punctuation patterns that go with each one.
Where you have already seen it
- Mental health and journalling apps that track mood over time. Not "good day, bad day," but a real emotional pattern.
- Customer support tools that route angry tickets to a senior agent first. Emotion, not only negative sentiment, decides the routing.
- Social media monitoring during a product launch or a crisis. Distinguishing genuine anger from mild disappointment matters for the response.
- Chatbots that adjust their tone based on how you sound. A frustrated user gets a different kind of reply than a confused one.
Remember this
- Emotion classification predicts a specific feeling, not only positive or negative.
- Two emotions can share the same sentiment polarity and still call for very different responses.
- It works the same way as sentiment analysis, only trained on more specific labels.
What to learn next
- Fine-grained sentiment — the closest relative, using a scale instead of named emotions.
- Sarcasm and irony detection — a case where the stated emotion and the real one do not match.
- Text classification — the general technique this lesson specialises.
Developer — Code and libraries.
Below, a model predicts the top emotions in two short messages. It was trained on GoEmotions, a dataset of real Reddit comments labelled across 28 emotion categories.
Setup
pip install transformers torchThe first run downloads SamLowe/roberta-base-go_emotions, about 500 MB.
Predicting emotions in two messages
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="SamLowe/roberta-base-go_emotions",
top_k=3,
)
texts = [
"I can't believe I finally got the internship, I've been waiting for months!",
"The delivery guy left my package outside in the rain again.",
]
for text in texts:
scores = classifier(text)[0]
print(text)
for s in scores:
print(f" {s['label']:12s} {s['score']:.3f}")
print()I can't believe I finally got the internship, I've been waiting for months! excitement 0.654 surprise 0.397 joy 0.045 The delivery guy left my package outside in the rain again. neutral 0.942 annoyance 0.019 approval 0.010
Line by line
top_k=3 returns the three highest-scoring emotions per message, not only the single best one. Real text often carries more than one emotion at once. Keeping the top few is usually more honest than forcing a single label.
The first message correctly reads as excitement and surprise. Both scored high, which makes sense together: unexpected good news genuinely does trigger both feelings at once.
The second message is worth reading twice. Most people would call it annoyed, even angry. The model scored it 0.942 "neutral," with annoyance barely registering at 0.019. This is a real limitation, not a display error.
Why the model likely missed it. GoEmotions is built from Reddit comments. Everyday mild complaints, stated flatly with no exclamation marks or angry words, are underrepresented in that kind of data. Loudly emotional text is not.
Common mistakes
Trusting "neutral" as a confident, correct label. As shown above, "neutral" can also mean "the model could not detect a real but understated emotion." Spot-check neutral-labelled text by hand, especially for anything downstream that matters.
Treating all 28 GoEmotions labels as equally reliable. Emotions with more training examples, like joy and anger, tend to be recognised more reliably. Rarer ones, like grief or pride, do not. Check per-class performance before trusting a rare label.
Assuming emotion labels transfer across text styles. A model trained on Reddit comments may not behave the same way on formal emails, support tickets, or spoken transcripts. Validate on your actual data before deploying.
Picking only the top-1 label when the top few are close. If excitement scores 0.40 and surprise scores 0.38, calling it "excitement" alone throws away a genuinely close second reading.
Try it yourself
Rewrite the delivery message more explicitly. Try: "I am so annoyed, the delivery guy left my package outside in the rain again." Compare the new scores against the original, flatter version.
What to learn next
- Fine-grained sentiment — a scale-based alternative to naming specific emotions.
- Toxicity and abuse detection — a related, narrower task focused on one specific negative signal.
- Reliability diagrams and calibration error — understanding exactly what a confidence score like 0.942 does and does not promise.
Researcher — Mathematics and papers.
Emotion models: categorical versus dimensional
Two competing psychological frameworks underlie computational emotion classification. The categorical model (Ekman, 1992) proposes a small set of universal, discrete emotions: anger, disgust, fear, joy, sadness, surprise. The dimensional model (Russell, 1980) instead represents any emotional state as a point in a continuous space. That space is typically valence (positive-negative) crossed with arousal (calm-excited).
Most NLP emotion classifiers, including the one in the developer block, use an extended categorical scheme. GoEmotions (Demszky et al., 2020) uses 28 categories, considerably finer-grained than Ekman's original six. This reflects the richer emotional vocabulary present in real online text.
Multi-label, not multi-class
Emotion classification is typically framed as multi-label classification. A single text can carry multiple simultaneous emotions. Standard sentiment analysis, by contrast, is multi-class: exactly one label applies. This is reflected directly in the developer block's output. Excitement and surprise both score meaningfully high for the same message. The training objective, independent per-label binary cross-entropy, explicitly allows that.
Known dataset and modelling biases
Class imbalance. GoEmotions' 28 categories are far from evenly represented. "Neutral" and "admiration" dominate by raw count. Categories like "grief" and "pride" have relatively few training examples. This directly explains uneven per-class reliability. It is a standard, measurable property, not a failure specific to any one architecture.
Platform-specific register. Training data drawn from one platform encodes that platform's conventions for expressing emotion: sarcasm markers, capitalisation patterns, emoji use. Reddit for GoEmotions, Twitter for many others. A model trained on one register underperforms measurably on a different one, such as formal writing or spoken transcripts. That is exactly the failure in the developer block's second example.
Understated versus performed emotion. Emotionally charged text online is frequently performed for an audience, with exclamation marks, capitalisation and explicit emotion words. Real understated emotion, the kind common in workplace or support communication, is comparatively underrepresented in most training sets. This biases models toward needing more explicit textual signal than a human reader does.
Evaluation
Multi-label emotion classification is scored with per-label precision, recall and F1. These are then macro-averaged across all labels, treating rare and common emotions equally. This differs from micro-averaging, which weights by label frequency. Micro-averaging can mask poor performance on rare classes entirely, since they contribute little to the total count.
Key references
- Ekman, P. (1992). An Argument for Basic Emotions. Cognition and Emotion.
- Russell, J. (1980). A Circumplex Model of Affect. Journal of Personality and Social Psychology.
- Demszky, D. et al. (2020). GoEmotions. arXiv:2005.00547
- Mohammad, S. & Turney, P. (2013). NRC Emotion Lexicon. Computational Intelligence — a widely used word-level emotion resource, predating deep learning approaches.
Current state and open problems
LLMs can be prompted to classify emotion directly. This now performs competitively with dedicated fine-tuned classifiers on well-studied benchmarks like GoEmotions, particularly when the emotion is stated outright.
The open problem is exactly what the developer block surfaced. Understated, implicit emotional signal, with no explicit emotion words or punctuation, remains difficult for both fine-tuned classifiers and general-purpose LLMs. No current benchmark adequately measures this specific failure mode, since most emotion datasets are themselves drawn from naturally emotionally-expressive text.
What to learn next
- Toxicity and abuse detection — a closely related task, scoring one specific harmful signal rather than a full emotion palette.
- Multi-label classification — the general machinery behind training on datasets like GoEmotions.
- Sarcasm and irony detection — where the literal words actively mislead a model's emotional read.