Sentiment, Opinion and Text Mining

Emotion classification

Emotion classification labels text with a specific feeling, like joy, anger or surprise, instead of only positive or negative.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Emotion classification labels text with a specific feeling, like joy, anger or surprise, not only positive or negative.

Picture two messages: "I won the lottery!" and "I finally got the job." Both are strongly positive. They do not feel the same at all. One is shock, the other is pride.

Sentiment analysis would call both "positive" and stop there. Emotion classification goes further. It asks which feeling, out of a set like joy, anger, fear, sadness and surprise.

Why it exists

"Positive" and "negative" are useful, but thin. Anger and sadness are both negative. Yet they call for completely different responses, from a business or from a person reading them.

A support ticket written in anger needs a different reply than one written in sadness or confusion. Sentiment analysis cannot tell these apart. Emotion classification can.

This matters anywhere the specific feeling changes what happens next. Mental health apps, support routing, and social media monitoring during a public event all qualify.

How it works

"I can't believe I got the internship!"  ->  excitement, surprise
"The delivery was left outside again."   ->  neutral (mild annoyance, unstated)

An emotion model is trained on text labelled with specific feelings, not only positive or negative. It learns the words, phrasing and punctuation patterns that go with each one.

Where you have already seen it

  • Mental health and journalling apps that track mood over time. Not "good day, bad day," but a real emotional pattern.
  • Customer support tools that route angry tickets to a senior agent first. Emotion, not only negative sentiment, decides the routing.
  • Social media monitoring during a product launch or a crisis. Distinguishing genuine anger from mild disappointment matters for the response.
  • Chatbots that adjust their tone based on how you sound. A frustrated user gets a different kind of reply than a confused one.

Remember this

  • Emotion classification predicts a specific feeling, not only positive or negative.
  • Two emotions can share the same sentiment polarity and still call for very different responses.
  • It works the same way as sentiment analysis, only trained on more specific labels.

What to learn next

Developer — Code and libraries.

Below, a model predicts the top emotions in two short messages. It was trained on GoEmotions, a dataset of real Reddit comments labelled across 28 emotion categories.

Setup

bash
pip install transformers torch

The first run downloads SamLowe/roberta-base-go_emotions, about 500 MB.

Predicting emotions in two messages

emotion.py
from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="SamLowe/roberta-base-go_emotions",
    top_k=3,
)

texts = [
    "I can't believe I finally got the internship, I've been waiting for months!",
    "The delivery guy left my package outside in the rain again.",
]

for text in texts:
    scores = classifier(text)[0]
    print(text)
    for s in scores:
        print(f"  {s['label']:12s} {s['score']:.3f}")
    print()
Output
I can't believe I finally got the internship, I've been waiting for months!
  excitement   0.654
  surprise     0.397
  joy          0.045

The delivery guy left my package outside in the rain again.
  neutral      0.942
  annoyance    0.019
  approval     0.010

Line by line

top_k=3 returns the three highest-scoring emotions per message, not only the single best one. Real text often carries more than one emotion at once. Keeping the top few is usually more honest than forcing a single label.

The first message correctly reads as excitement and surprise. Both scored high, which makes sense together: unexpected good news genuinely does trigger both feelings at once.

The second message is worth reading twice. Most people would call it annoyed, even angry. The model scored it 0.942 "neutral," with annoyance barely registering at 0.019. This is a real limitation, not a display error.

Why the model likely missed it. GoEmotions is built from Reddit comments. Everyday mild complaints, stated flatly with no exclamation marks or angry words, are underrepresented in that kind of data. Loudly emotional text is not.

Common mistakes

Trusting "neutral" as a confident, correct label. As shown above, "neutral" can also mean "the model could not detect a real but understated emotion." Spot-check neutral-labelled text by hand, especially for anything downstream that matters.

Treating all 28 GoEmotions labels as equally reliable. Emotions with more training examples, like joy and anger, tend to be recognised more reliably. Rarer ones, like grief or pride, do not. Check per-class performance before trusting a rare label.

Assuming emotion labels transfer across text styles. A model trained on Reddit comments may not behave the same way on formal emails, support tickets, or spoken transcripts. Validate on your actual data before deploying.

Picking only the top-1 label when the top few are close. If excitement scores 0.40 and surprise scores 0.38, calling it "excitement" alone throws away a genuinely close second reading.

Try it yourself

Rewrite the delivery message more explicitly. Try: "I am so annoyed, the delivery guy left my package outside in the rain again." Compare the new scores against the original, flatter version.

What to learn next

Researcher — Mathematics and papers.

Emotion models: categorical versus dimensional

Two competing psychological frameworks underlie computational emotion classification. The categorical model (Ekman, 1992) proposes a small set of universal, discrete emotions: anger, disgust, fear, joy, sadness, surprise. The dimensional model (Russell, 1980) instead represents any emotional state as a point in a continuous space. That space is typically valence (positive-negative) crossed with arousal (calm-excited).

Most NLP emotion classifiers, including the one in the developer block, use an extended categorical scheme. GoEmotions (Demszky et al., 2020) uses 28 categories, considerably finer-grained than Ekman's original six. This reflects the richer emotional vocabulary present in real online text.

Multi-label, not multi-class

Emotion classification is typically framed as multi-label classification. A single text can carry multiple simultaneous emotions. Standard sentiment analysis, by contrast, is multi-class: exactly one label applies. This is reflected directly in the developer block's output. Excitement and surprise both score meaningfully high for the same message. The training objective, independent per-label binary cross-entropy, explicitly allows that.

Known dataset and modelling biases

Class imbalance. GoEmotions' 28 categories are far from evenly represented. "Neutral" and "admiration" dominate by raw count. Categories like "grief" and "pride" have relatively few training examples. This directly explains uneven per-class reliability. It is a standard, measurable property, not a failure specific to any one architecture.

Platform-specific register. Training data drawn from one platform encodes that platform's conventions for expressing emotion: sarcasm markers, capitalisation patterns, emoji use. Reddit for GoEmotions, Twitter for many others. A model trained on one register underperforms measurably on a different one, such as formal writing or spoken transcripts. That is exactly the failure in the developer block's second example.

Understated versus performed emotion. Emotionally charged text online is frequently performed for an audience, with exclamation marks, capitalisation and explicit emotion words. Real understated emotion, the kind common in workplace or support communication, is comparatively underrepresented in most training sets. This biases models toward needing more explicit textual signal than a human reader does.

Evaluation

Multi-label emotion classification is scored with per-label precision, recall and F1. These are then macro-averaged across all labels, treating rare and common emotions equally. This differs from micro-averaging, which weights by label frequency. Micro-averaging can mask poor performance on rare classes entirely, since they contribute little to the total count.

Key references

  • Ekman, P. (1992). An Argument for Basic Emotions. Cognition and Emotion.
  • Russell, J. (1980). A Circumplex Model of Affect. Journal of Personality and Social Psychology.
  • Demszky, D. et al. (2020). GoEmotions. arXiv:2005.00547
  • Mohammad, S. & Turney, P. (2013). NRC Emotion Lexicon. Computational Intelligence — a widely used word-level emotion resource, predating deep learning approaches.

Current state and open problems

LLMs can be prompted to classify emotion directly. This now performs competitively with dedicated fine-tuned classifiers on well-studied benchmarks like GoEmotions, particularly when the emotion is stated outright.

The open problem is exactly what the developer block surfaced. Understated, implicit emotional signal, with no explicit emotion words or punctuation, remains difficult for both fine-tuned classifiers and general-purpose LLMs. No current benchmark adequately measures this specific failure mode, since most emotion datasets are themselves drawn from naturally emotionally-expressive text.

What to learn next