Sentiment, Opinion and Text Mining
Sentiment beyond positive and negative
Fine-grained sentiment rates text on a scale, like one to five stars, instead of forcing every opinion into only positive or negative.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Fine-grained sentiment rates text on a scale, like one to five stars, instead of only positive or negative.
Picture rating a restaurant. "Good" and "bad" are not enough. You want to say it was decent, not amazing, maybe three stars out of five.
Basic sentiment analysis only offers two buckets: positive or negative. Fine-grained sentiment gives you a scale instead, closer to how people actually rate things in real life.
Why it exists
A plain positive-or-negative model throws away real information. "It was okay" and "best purchase of my life" both get labelled "positive," even though they mean very different things.
Businesses need that missing nuance. A product review site wants to know if a rating is trending from 4.2 to 3.8 stars. Saying "still positive" hides that entirely. A two-bucket system cannot show that kind of movement at all.
Fine-grained sentiment keeps the ranking information a plain positive-negative split destroys.
How it works
"It's fine, does the job." -> 3-4 stars (mild)
"Absolutely loved it, best ever!" -> 5 stars (strongly positive)
"Terrible, broke on day two." -> 1 star (strongly negative)A fine-grained model is trained on text that already carries a star rating, commonly pulled from real review sites. It learns which words and phrases tend to go with which rating.
Where you have already seen it
- Amazon, Flipkart and Google star ratings, estimated from review text. Some platforms auto-suggest a star rating based on what you typed.
- App store review dashboards. Developers track average sentiment score over time, not only a positive-versus-negative count.
- Customer feedback tools that show a 1-10 satisfaction trend line. That line needs a scale, not a binary flag.
- Movie and restaurant aggregator scores. Built from thousands of reviews, each scored on a finer scale than "liked it or not."
Remember this
- Fine-grained sentiment uses a scale, commonly one to five stars, not only positive or negative.
- It preserves nuance a two-bucket system throws away.
- It is trained the same way as basic sentiment analysis, only on data with finer labels.
What to learn next
- Aspect-based sentiment analysis — rating different parts of the same review separately.
- Text classification — the general technique fine-grained sentiment is a specific case of.
- Emotion classification — going further than a positive-to-negative scale entirely.
Developer — Code and libraries.
Below, a model trained on real product reviews across five languages predicts a 1-5 star rating directly from review text.
Setup
pip install transformers torchThe first run downloads nlptown/bert-base-multilingual-uncased-sentiment, about 670 MB. This is a full multilingual BERT model. It is noticeably bigger than the tiny models used elsewhere on this site. That is a real cost worth knowing before you download it.
Rating three reviews
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="nlptown/bert-base-multilingual-uncased-sentiment",
)
reviews = [
"It's fine. Does the job, nothing more, nothing less.",
"Absolutely loved it, best purchase I made all year!",
"Terrible. Broke on day two and support never replied.",
]
for text in reviews:
result = classifier(text)[0]
print(f"{result['label']:10s} ({result['score']:.3f}) {text}")4 stars (0.534) It's fine. Does the job, nothing more, nothing less. 5 stars (0.983) Absolutely loved it, best purchase I made all year! 1 star (0.951) Terrible. Broke on day two and support never replied.
Line by line
The output label is a star count, not "positive" or "negative." That is the entire difference from basic sentiment analysis: five labels instead of two.
Look closely at the confidence scores. The strongly positive and strongly negative reviews scored 0.983 and 0.951, both very confident. The mild, middling review scored only 0.534. Even then it landed on "4 stars," arguably generous for "nothing more, nothing less."
That low-confidence middle case is worth sitting with. Fine-grained models are usually far more confident at the extremes than in the middle of the scale. Genuinely mixed or lukewarm feelings are the hardest cases, for a model and for a person.
Common mistakes
Treating the star output as exact, not probabilistic. A 0.534 confidence "4 stars" is not the same as a 0.983 confidence "5 stars." Always check the score, not only the label.
Assuming five-star granularity is always what you need. Sometimes three buckets, negative, neutral, positive, are more reliable and easier to act on than five. More classes means more room for the model to be uncertain between adjacent ones.
Running this large multilingual model when a smaller one would do. If your text is only English, a smaller English-only sentiment model will usually be faster with similar accuracy. Reach for multilingual models specifically when you need multilingual coverage.
Forgetting that star scales are trained on review-style text. This model learned from product and restaurant reviews. A tweet, a support ticket or a doctor's note does not map onto "stars" the same way. None of those were part of its training data.
Try it yourself
Try "It could have been worse, I guess." An ambiguous sentence like this tests where the model's confidence actually drops.
What to learn next
- Aspect-based sentiment analysis — splitting one review into several separately-rated parts.
- Text classification — the underlying technique, in full depth.
- BERT — the architecture this lesson's model is built on.
Researcher — Mathematics and papers.
Ordinal structure, and why plain classification ignores it
Fine-grained sentiment is naturally an ordinal problem. The five classes have a meaningful order, 1 < 2 < 3 < 4 < 5. That is unlike arbitrary labels such as "cat" or "dog." Standard cross-entropy classification, treating each class as unrelated, ignores this. Predicting "1 star" for a true "5 star" review is penalised the same as predicting "4 stars." One of those errors is far smaller in any real sense.
Ordinal regression approaches address this directly. The cumulative link model of McCullagh (1980) models P(Y <= k) as a function of one latent score. It does not treat each class independently. In practice, most production fine-grained sentiment models still use plain classification. The accuracy loss from ignoring ordinality is often smaller than the engineering cost of a custom ordinal loss.
Training data and its biases
Models like the one above are trained on review platforms. Star ratings and review text are paired directly, at large scale. This introduces a specific bias worth naming: review-writing behaviour is not neutral. Extremely satisfied and extremely dissatisfied customers write reviews far more often than mildly satisfied ones. This is the well-documented J-shaped distribution of online reviews (Hu, Pavlou & Zhang, 2009).
The practical consequence is visible directly in the developer block. Models trained on this data become better calibrated at the extremes than in the middle. The middle is comparatively underrepresented.
SST-5 and the standard fine-grained benchmark
Socher et al. (2013) introduced the Stanford Sentiment Treebank. It is labelled at five levels, from very negative to very positive. Both full sentences and their constituent phrases carry a label. SST-5 remains the standard academic benchmark for fine-grained sentiment. Production models differ: they learn from naturally occurring star labels, not careful phrase-level annotation.
Fine-grained sentiment as a regression problem
An alternative framing treats the star rating as a continuous target. It applies regression loss, mean squared error, instead of classification loss. This naturally respects ordinal structure, with no custom loss function needed. The cost: it loses the clean probability distribution over discrete classes a classification head gives you. Some downstream uses genuinely need that.
Key references
- McCullagh, P. (1980). Regression Models for Ordinal Data. Journal of the Royal Statistical Society.
- Socher, R. et al. (2013). Recursive Deep Models for Semantic Compositionality (SST). EMNLP.
- Hu, N., Pavlou, P. & Zhang, J. (2009). Overcoming the J-shaped Distribution of Product Reviews. Communications of the ACM.
Current state and open problems
LLMs, prompted to output a numeric rating directly, now handle fine-grained sentiment competitively with no task-specific fine-tuning. This holds particularly well for common domains like product reviews.
The open problem is calibration in the middle of the scale, the exact weakness surfaced by the developer block's example. No widely adopted training method closes this gap reliably. It follows directly from how naturally occurring review data is distributed, not from a fixable modelling bug.
What to learn next
- Text classification — the general classification machinery fine-grained sentiment specialises.
- Aspect-based sentiment analysis — adding structure sentiment analysis alone does not capture.
- Reliability diagrams and calibration error — the deeper issue behind this lesson's low-confidence middle case.