Build a sentiment analyser
Train a model to tell a happy review from an angry one, read the exact words it learned, then watch it fail on "not good" and on sarcasm.
- 22 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A sentiment analyser reads a piece of writing and decides whether the writer was happy or unhappy.
The analogy you have already lived
Your friend replies to a long message with one word: "fine." You know instantly that things are not fine.
Nobody taught you that rule. You picked it up from thousands of conversations, and now you read tone without effort.
Or think about shopping online. You scroll through twenty reviews without reading any of them properly. Certain words catch your eye — "waste", "broke", "excellent" — and within seconds you have a feeling about the product.
That skimming is what this project automates.
Why this project exists
A popular product might collect fifty thousand reviews. A film gets a hundred thousand tweets on opening weekend. A bank's app gets comments faster than any team can read.
Nobody can read all of it. But somebody needs to know the answer to one question: are people happy or angry, and about what?
Before this, companies hired people to read samples by hand. That was slow, expensive, and by the time the report arrived the problem had grown for three weeks.
Sentiment analysis gives you the shape of the answer in seconds. You still read the individual complaints — the model tells you which pile to read first.
What you are actually going to build
You will write down sixty short reviews. Thirty are from happy customers, thirty from unhappy ones. Each one carries its true label.
The model learns which words lean happy and which lean angry. Then it reads reviews it has never seen and makes a call, with a confidence attached.
The lovely part is that you can open the model up and read its opinion of every single word. You will see, in plain text, that it learned "excellent" is positive and "terrible" is negative. Nobody told it that.
How it works
"Poor build quality, feels cheap, waste of money"
|
v
[ chop into words ]
|
v
poor -> leans angry cheap -> leans angry
waste -> leans angry money -> neutral
|
v
[ add the leanings up ]
|
v
"negative" (96% sure)Nothing here understands English. The model is adding up word leanings. It works surprisingly well, and it fails in ways that will teach you a lot.
Where you have already seen this
- Amazon and Flipkart showing you what reviewers liked and disliked most.
- Zomato and Swiggy summarising thousands of restaurant reviews into themes.
- Election coverage claiming to measure public mood from social media.
- Support systems that push angry tickets to the front of the queue.
The honest part — and this one is big
You are going to break this model on purpose, and it will break easily.
Type "The food was not good at all." Your model will call it positive. It sees the word "good" and counts it as happy. The word "not" sitting right in front changes nothing. The model has no idea that words affect their neighbours.
Now type "This phone is not bad." Your model calls it negative, for the mirror-image reason.
Then try sarcasm: "Great, another two hour delay. Wonderful service." Your model sees "great" and "wonderful" and confidently reports a delighted customer.
These are not bugs you can patch. They come from the model's basic design, which treats a sentence as a loose bag of words with the order thrown away. Fixing them properly needs a model that reads words in sequence, which is what transformers do.
There is a harder problem underneath. Human beings disagree about sentiment. Show ten people the same review and two or three will label it differently. No model can be more accurate than the humans who made the labels.
Remember this
- Sentiment analysis sorts writing into happy or unhappy, from labelled examples.
- It works by learning which words lean which way, then adding them up.
- It cannot handle "not good" or sarcasm, because it ignores word order entirely.
What to learn next
- Text classification — the general method, beyond sentiment.
- Embeddings — representations that know "reimbursed" and "refund" are related.
- BERT — the sequence model that handles the negation cases this project cannot.
Developer — Code and libraries.
The problem, stated precisely
Given a short review, output the probability that the writer was positive. Binary text classification again, exactly like your spam detector — but a much harder label, because sentiment lives in phrasing rather than in vocabulary alone.
Our plan:
- Sixty hand-written reviews, thirty per class.
- TF-IDF features plus logistic regression — the standard strong baseline for text.
- Read the coefficients to see what the model learned.
- Attack it with negation and sarcasm, and understand precisely why it loses.
Setup
pip install scikit-learn numpyNo downloads, no GPU, about three seconds of runtime.
The dataset
Sixty short reviews, typed into the file. Real review corpora reuse a small set of sentiment words heavily — "great", "poor", "terrible", "excellent" appear again and again — so these reviews do the same. A model that learns from word counts needs that repetition to have anything to generalise from.
Be warned that this makes the dataset easier than reality. You will see the consequence in the very first number.
The full build
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_score, train_test_split
from sklearn.metrics import classification_report
DATA = [
# ---- positive (1) ----
("Great phone, the battery is excellent and it never heats up", 1),
("Excellent food, delicious biryani, will order again", 1),
("Great service and the staff were very helpful", 1),
("Love this laptop, fast, quiet and great value", 1),
("The picture quality is excellent and the sound is great", 1),
("Fast delivery, good packaging, very happy with it", 1),
("Comfortable seats and great legroom, a good journey", 1),
("Brilliant film, great acting, I loved every minute", 1),
("Good build quality, feels premium, worth the money", 1),
("Helpful support team, they fixed my problem quickly", 1),
("Very happy with this purchase, highly recommended", 1),
("The camera is great, sharp photos even at night", 1),
("Clean room, lovely view, the staff were polite", 1),
("Runs smoothly, good performance, stays cool", 1),
("Delicious food, generous portions, good value for money", 1),
("Lovely story and excellent acting, a beautiful film", 1),
("Charges fast and the battery lasts all day, love it", 1),
("The teacher explained everything clearly, a great course", 1),
("Quick refund and the support was polite and helpful", 1),
("Beautiful design, feels great in the hand, good grip", 1),
("Good fabric, comfortable fit, happy with the quality", 1),
("Quiet, fast and reliable, excellent for daily work", 1),
("The driver was polite and drove safely, good trip", 1),
("Clear instructions, easy to assemble, very good product", 1),
("Fresh flowers, delivered on time, my mother loved them", 1),
("Bright screen, fast touch response, a great upgrade", 1),
("Good price for excellent quality, genuinely recommended", 1),
("Lovely cafe, friendly staff and excellent coffee", 1),
("The guide was knowledgeable and very helpful throughout", 1),
("Reliable battery backup, good product, works perfectly", 1),
# ---- negative (0) ----
("Terrible phone, the battery is poor and it heats up badly", 0),
("Terrible food, cold biryani, never ordering again", 0),
("Poor service and the staff were rude and unhelpful", 0),
("Hate this laptop, slow, noisy and poor value", 0),
("The picture quality is poor and the sound is terrible", 0),
("Slow delivery, damaged packaging, very unhappy with it", 0),
("Uncomfortable seats and no legroom, a terrible journey", 0),
("Boring film, poor acting, I hated every minute", 0),
("Poor build quality, feels cheap, waste of money", 0),
("Unhelpful support team, they ignored my problem completely", 0),
("Very unhappy with this purchase, not recommended", 0),
("The camera is poor, blurry photos even in daylight", 0),
("Dirty room, terrible view, the staff were rude", 0),
("Lags badly, poor performance, gets very hot", 0),
("Stale food, tiny portions, poor value for money", 0),
("Boring story and terrible acting, an awful film", 0),
("Charges slowly and the battery dies by noon, hate it", 0),
("The teacher rushed everything, a terrible course", 0),
("Slow refund and the support was rude and unhelpful", 0),
("Ugly design, feels cheap in the hand, poor grip", 0),
("Poor fabric, uncomfortable fit, unhappy with the quality", 0),
("Noisy, slow and unreliable, terrible for daily work", 0),
("The driver was rude and drove badly, awful trip", 0),
("Confusing instructions, hard to assemble, very poor product", 0),
("Wilted flowers, delivered late, my mother hated them", 0),
("Dull screen, slow touch response, a terrible upgrade", 0),
("High price for poor quality, not recommended at all", 0),
("Dirty cafe, rude staff and terrible coffee", 0),
("The guide was careless and completely unhelpful throughout", 0),
("Unreliable battery backup, poor product, stopped working", 0),
]
texts = [t for t, _ in DATA]
labels = [y for _, y in DATA]
print("reviews:", len(texts), " positive:", sum(labels))
# TF-IDF down-weights words that appear in nearly every review, such as "the".
# C=10 loosens the default regularisation, which is too strong for 60 rows.
pipe = make_pipeline(TfidfVectorizer(), LogisticRegression(C=10, max_iter=2000))
scores = cross_val_score(pipe, texts, labels, cv=5)
print("5-fold accuracy:", np.round(scores, 3), " mean:", round(float(scores.mean()), 3))
X_tr, X_te, y_tr, y_te = train_test_split(texts, labels, test_size=0.25,
random_state=1, stratify=labels)
pipe.fit(X_tr, y_tr)
print()
print(classification_report(y_te, pipe.predict(X_te),
target_names=["negative", "positive"], digits=3))reviews: 60 positive: 30
5-fold accuracy: [1. 1. 1. 1. 1.] mean: 1.0
precision recall f1-score support
negative 0.875 1.000 0.933 7
positive 1.000 0.875 0.933 8
accuracy 0.933 15
macro avg 0.938 0.938 0.933 15
weighted avg 0.942 0.933 0.933 15A perfect cross-validation score is bad news
[1. 1. 1. 1. 1.] is not a triumph. It is a warning that the task, as I built it, is too easy.
I wrote these sixty reviews in one sitting. Positives lean on "great", "excellent", "good" and "love". Negatives lean on "poor", "terrible", "rude" and "slow". Real customers are far messier: they write mixed reviews, they use words neither side of my vocabulary contains, and they type "gr8".
Whenever a text model scores a perfect 1.0, look for the shortcut before you celebrate. Usually the dataset is templated, or the same content appears in train and test, or a label leaked into a feature. On real review data, a good TF-IDF baseline lands somewhere in the high eighties.
Now read what the model learned
pipe.fit(texts, labels) # refit on everything to inspect the weights
vec = pipe.named_steps["tfidfvectorizer"]
clf = pipe.named_steps["logisticregression"]
names = np.array(vec.get_feature_names_out())
order = np.argsort(clf.coef_[0])
print("vocabulary size:", len(names))
print("most negative words:", list(names[order[:8]]))
print("most positive words:", list(names[order[-8:]][::-1]))
print()
probe = [
"The camera is great and the battery is excellent",
"Poor quality and terrible service",
"The food was not good at all",
"This phone is not bad",
"Great, another two hour delay. Wonderful service.",
]
for t, p in zip(probe, pipe.predict_proba(probe)[:, 1]):
print(f"{'POSITIVE' if p > 0.5 else 'NEGATIVE':9} p={p:.3f} {t}")vocabulary size: 184 most negative words: ['poor', 'terrible', 'rude', 'unhelpful', 'slow', 'unhappy', 'not', 'completely'] most positive words: ['good', 'great', 'excellent', 'helpful', 'fast', 'happy', 'polite', 'lovely'] POSITIVE p=0.877 The camera is great and the battery is excellent NEGATIVE p=0.044 Poor quality and terrible service POSITIVE p=0.636 The food was not good at all NEGATIVE p=0.333 This phone is not bad POSITIVE p=0.878 Great, another two hour delay. Wonderful service.
The word list is the whole point
The model was never handed a dictionary of positive and negative words. It worked out from sixty examples that "terrible" and "rude" mean trouble, and "excellent" and "lovely" mean a happy customer.
Being able to read a model's reasoning is worth a great deal. When a business asks "why did you flag this review?", you can point at four words. Very few modern models allow that.
Notice not sitting in the negative list. It got there from "not recommended" appearing twice in the negative half. The model has learned that this particular string tends to show up in complaints. It has not learned what negation means, and the next three lines prove it.
Three failures, and exactly why each one happens
"The food was not good at all" → POSITIVE at 0.636.
The model sees not (negative weight) and good (strong positive weight) and adds them. Positive wins.
It has no concept of one word reversing the meaning of another. A sentence, to this model, is an unordered bag. "not good" and "good not" are literally identical inputs.
"This phone is not bad" → NEGATIVE at 0.333.
Same machinery, opposite direction. bad and not both pull negative, so the sum is negative. A human reads "not bad" as mild praise.
"Great, another two hour delay. Wonderful service." → POSITIVE at 0.878.
Sarcasm. Every visible signal says happy: great, wonderful, service. The negative content — a two-hour delay — is a fact about the world, and understanding that it is bad requires knowing what delays are.
This one is not a bag-of-words limitation. Sarcasm defeats large modern models too, and human annotators disagree about it constantly. Be honest with anyone who asks: automated sarcasm detection is not solved.
Does adding word pairs fix it?
The standard advice is to add bigrams — pairs of adjacent words — so "not good" becomes its own feature. Test that claim rather than believing it.
bigram = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2)), # single words AND adjacent word pairs
LogisticRegression(C=10, max_iter=2000),
)
bigram.fit(texts, labels)
hard = probe[2:] # the three sentences the word-only model got wrong
print("vocabulary size with bigrams:",
len(bigram.named_steps["tfidfvectorizer"].get_feature_names_out()))
for t, p in zip(hard, bigram.predict_proba(hard)[:, 1]):
print(f"{'POSITIVE' if p > 0.5 else 'NEGATIVE':9} p={p:.3f} {t}")
print()
vocab = set(bigram.named_steps["tfidfvectorizer"].get_feature_names_out())
for pair in ["not good", "not bad", "not recommended", "never ordering"]:
print(f"is {pair!r} in the training vocabulary? {pair in vocab}")vocabulary size with bigrams: 509 POSITIVE p=0.572 The food was not good at all NEGATIVE p=0.394 This phone is not bad POSITIVE p=0.835 Great, another two hour delay. Wonderful service. is 'not good' in the training vocabulary? False is 'not bad' in the training vocabulary? False is 'not recommended' in the training vocabulary? True is 'never ordering' in the training vocabulary? True
It barely moved. 0.636 became 0.572 — still the wrong answer.
The last four lines explain why. A bigram feature only exists if that exact pair appeared in training. "not recommended" is in the vocabulary because I wrote it twice. "not good" is not, because I never wrote it. A feature the model has never seen carries zero weight.
This generalises far beyond this example. n-grams do not teach a model about negation. They memorise specific negated phrases you happened to show it. Real negation handling needs either a model that reads sequences, or a corpus large enough that most common negated phrases appear somewhere in it. In practice, large corpora are why n-gram models worked as well as they did for twenty years.
Common mistakes
Leaving LogisticRegression at its default C=1.0 on small text data. TF-IDF vectors are L2-normalised, so each feature value is small. The default penalty then dominates and every prediction sits near 0.5. This is under-fitting that looks like a broken pipeline. Raise C or use LogisticRegressionCV.
Removing stop words automatically. stop_words="english" deletes "not", "no", "never" and "against". For topic classification that is harmless. For sentiment it deletes the most important words in the sentence.
Fitting the vectorizer before splitting. Test vocabulary leaks into training. Wrap everything in a Pipeline, as this code does, and scikit-learn refits the vectorizer inside each cross-validation fold correctly.
Training on one domain and deploying on another. A model trained on electronics reviews learns that "cold" is negative. Move it to ice-cream reviews and "cold" flips meaning. Sentiment vocabulary is domain-specific, and accuracy drops sharply across domains.
Treating three-star reviews as a coin flip. Most real review data is on a five-point scale, and the middle is genuinely mixed rather than neutral. Either drop the middle, or model three classes and accept a lower headline number.
How to make this genuinely good
- Get real data. The IMDB movie review set (50,000 labelled reviews, about 80 MB) or the SST-2 sentence set (about 70,000 short sentences, roughly 7 MB) are the standard benchmarks. SST-2 is the kinder download on a mobile connection.
- Add character n-grams with
TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5)). These survive spelling mistakes, transliterated Hindi and "gooood". - Predict aspects, not documents. "Great camera, terrible battery" has no single sentiment. Splitting by aspect is what businesses actually want, and it is more useful than a better overall accuracy.
- Move to a transformer when the baseline stops improving. A fine-tuned DistilBERT handles negation properly because it reads the sequence. Be aware this is a download of roughly 250 MB and needs several minutes per epoch on CPU — real, but not free.
- Calibrate and threshold. For a support queue, route only reviews below
p = 0.2to a human. Confidence is more actionable than a hard label. - Measure the human ceiling. Have two people label 200 reviews independently and compute their agreement. That number is your realistic upper bound.
Try it yourself
Add these four reviews to DATA, then re-run the bigram model:
("The food was not good and the service was not great", 0),
("Not happy with this at all, the quality is not good", 0),
("This phone is not bad, actually quite good for the price", 1),
("Not a bad film, I did not regret watching it", 1),Both negation failures flip to the correct answer. The sarcasm example stays wrong.
That is the honest shape of this whole field. You can fix a failure by showing the model examples of it. You cannot fix a failure that requires knowing a two-hour delay is a bad thing.
What to learn next
- Text classification — the general method, beyond sentiment.
- Embeddings — representations that know "reimbursed" and "refund" are related.
- BERT — the sequence model that handles the negation cases this project cannot.
Researcher — Mathematics and papers.
Formulas here are written in plain text, since the site renders no maths typesetting library.
TF-IDF
scikit-learn's TfidfVectorizer computes, for term t in document d across a corpus of N documents:
tf-idf(t, d) = tf(t, d) * idf(t)
idf(t) = ln( (1 + N) / (1 + df(t)) ) + 1 [smooth_idf=True, the default]tf(t, d)is the raw count oftind.df(t)is the number of documents containingt.- The
+1terms implement add-one smoothing, as though one extra document contained every term, soidfnever divides by zero. - The trailing
+1prevents terms appearing in every document from being zeroed out entirely.
Each document vector is then L2-normalised, so || x_d ||_2 = 1. Note that this differs from the textbook definition (Spärck Jones, 1972), where idf = log(N / df(t)). Numbers from scikit-learn will not match a hand calculation from a textbook, which causes a great deal of confusion.
The L2 normalisation interacts directly with regularisation strength: because every feature value is bounded by 1 and typically far smaller, the effective penalty at a given C is much heavier than on raw counts. This is why C=1.0 under-fits on small TF-IDF matrices.
Logistic regression
p(y = 1 | x) = sigma( w . x + b ), sigma(z) = 1 / (1 + exp(-z))Fitted by minimising the L2-penalised negative log-likelihood:
J(w) = C * sum_i [ -y_i log p_i - (1 - y_i) log(1 - p_i) ] + (1/2) || w ||_2^2Cis the inverse regularisation strength. LargeCmeans a weak penalty. Note that scikit-learn scales the loss byCrather than the penalty bylambda; they are equivalent up toC = 1 / lambda.- The objective is convex in
w, so the optimum is global.lbfgsreaches it reliably at this scale.
The decision function is linear in the features, so each coefficient w_t is exactly the change in log-odds contributed per unit of tf-idf(t). That is why sorting clf.coef_[0] gives a directly interpretable word ranking — an interpretability property that does not survive any non-linear model.
Complexity. Vectorising is O(total tokens). Fitting with lbfgs is O(n * nnz) per iteration where nnz is average non-zeros per row. Both are trivial at corpus sizes below a few million documents, which is why this baseline remains the correct first move.
Why bag-of-words cannot represent negation
The representation is a multiset. For any permutation pi of the token sequence, phi(x) = phi(pi(x)). "not good" and "good not" map to the identical vector, so no function of phi(x) can distinguish them.
Adding n-grams enlarges the feature space but does not change the character of the failure. An n-gram model is a bag over spans of length up to n; it can represent negation only for the specific spans observed in training, and the number of possible negated phrases grows combinatorially. Negation scope in English can also extend well beyond a window — "I did not find the acting, the writing, or the direction convincing" negates three constituents at distance.
The classical workaround (Das and Chen, 2001; used by Pang, Lee and Vaithyanathan, 2002) tags every token between a negation cue and the next punctuation with a NOT_ prefix, so good becomes NOT_good as a distinct feature. It measurably helps and is a heuristic, not a solution: it mis-scopes on subordinate clauses and doubles the vocabulary.
Sequence models — LSTMs, and then transformers — represent negation because the hidden state at each position depends on the preceding tokens. Socher et al. (2013) made this concrete with the Stanford Sentiment Treebank, which annotates sentiment at every node of a parse tree specifically so that compositional effects can be trained and measured, rather than assumed.
Benchmarks and their ceilings
- Pang and Lee (2004) polarity dataset: 2,000 movie reviews. Small, and thoroughly over-fitted by the field.
- IMDB (Maas et al., 2011): 50,000 reviews, deliberately polarised — ratings of 5 and 6 are excluded. Real reviews include the middle, so IMDB accuracy overstates deployed performance.
- SST-2 (Socher et al., 2013): 67,000 sentence-level items with phrase-level annotation. Part of GLUE.
- Amazon multi-domain (Blitzer, Dredze and Pereira, 2007): the standard testbed for cross-domain transfer, and the source of the durable finding that sentiment lexicons are domain-specific.
Two ceilings bound all of these. Annotator agreement: inter-annotator agreement on fine-grained sentiment commonly lands in the 0.7 to 0.8 Cohen's kappa range, and a model cannot exceed the reliability of its labels. Benchmark saturation: SST-2 leaderboard scores exceed 97%, above plausible label quality, which means remaining gains partly measure fitting to annotation artefacts.
Sarcasm and irony
SemEval-2018 Task 3 (Van Hee et al.) is the reference evaluation. Top systems reached F1 around 0.7 on binary irony detection in tweets — far below the numbers routinely quoted for sentiment.
The difficulty is structural rather than a modelling deficiency. Irony is signalled by a mismatch between the utterance and a situation the text does not describe. "Great, another two hour delay" requires world knowledge that delays are undesirable, plus the recognition that "great" is incongruent with it. Datasets built from #sarcasm self-labelling also carry a well-documented bias: authors tag sarcasm when it might be missed, so the corpus over-represents unusually explicit cases.
Treat any product claim of reliable sarcasm detection with scepticism, and measure it on your own data before you believe it.
What a production system looks like
Document-level polarity is rarely the deliverable. Aspect-based sentiment analysis (ABSA) extracts (aspect, opinion, polarity) triples, so "great camera, terrible battery" yields two opposite judgements. SemEval-2014 Task 4 (Pontiki et al.) defined the standard formulation, and modern systems cast it as sequence labelling or as generation of structured triples.
A deployed pipeline typically layers:
- Language identification and routing. Indian review streams are heavily code-mixed — Hindi or Tamil written in Latin script, interleaved with English. Monolingual English models degrade badly here; see Khanuja et al. (2020) on the GLUECoS benchmark.
- Spam and incentivised-review filtering, before any sentiment is computed. Otherwise you measure the marketing budget.
- Aspect extraction, then per-aspect polarity.
- Aggregation with uncertainty, reporting intervals rather than a single score.
- Drift monitoring. Product vocabulary and complaint themes shift continuously, so a static model silently degrades. Track the rate of out-of-vocabulary tokens as an early warning.
The unglamorous items — 1, 2 and 5 — determine whether the system is trusted. Model architecture rarely does.
Papers
- Spärck Jones, A Statistical Interpretation of Term Specificity and its Application in Retrieval, Journal of Documentation, 1972.
- Pang, Lee and Vaithyanathan, Thumbs up? Sentiment Classification using Machine Learning Techniques, EMNLP 2002 — the paper that started the field.
- Pang and Lee, A Sentimental Education, ACL 2004.
- Blitzer, Dredze and Pereira, Biographies, Bollywood, Boom-boxes and Blenders: Domain Adaptation for Sentiment Classification, ACL 2007.
- Maas, Daly, Pham, Huang, Ng and Potts, Learning Word Vectors for Sentiment Analysis, ACL 2011 — introduces IMDB.
- Socher, Perelygin, Wu et al., Recursive Deep Models for Semantic Compositionality over a Sentiment Treebank, EMNLP 2013 — introduces SST.
- Wang and Manning, Baselines and Bigrams, ACL 2012 — why this baseline is hard to beat.
- Pontiki et al., SemEval-2014 Task 4: Aspect Based Sentiment Analysis, SemEval 2014.
- Van Hee, Lefever and Hoste, SemEval-2018 Task 3: Irony Detection in English Tweets, SemEval 2018.
- Khanuja, Dandapat, Srinivasan, Sitaram and Choudhury, GLUECoS: An Evaluation Benchmark for Code-Switched NLP, ACL 2020 — arxiv.org/abs/2004.12376
What to learn next
- Text classification — the general method, beyond sentiment.
- Embeddings — representations that know "reimbursed" and "refund" are related.
- BERT — the sequence model that handles the negation cases this project cannot.