Stopwords: when removing them helps and when it breaks things
Removing "the" and "was" shrinks text and sharpens word counts — but the standard stopword lists also contain "not", and deleting it flips meanings.
- 6 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Stopwords are the very common little words — "the", "is", "was", "and" — that some text systems delete because they appear everywhere and distinguish nothing.
Think of sifting flour. The sieve lets the fine, uniform powder fall through and keeps the lumps — the parts with substance. Stopword removal is sifting for sentences: the everywhere-words fall through, the content words stay.
Why it exists
Old-school text systems represent a document by counting words — the bag-of-words approach behind text classification. Counts have a problem: "the" wins every count in every document. Nearly half of typical English text is a few dozen filler words.
For a search engine or a word-count model, those words are dead weight. They match everything, distinguish nothing, and bloat the index. So the field built stopword lists: a fixed set of words to delete before counting. Smaller data, sharper signal — usually.
How it works
"the food was not good"
│ remove stopwords (standard English list)
"food good"Read that example twice — it is the whole lesson. The list removed "the" and "was" harmlessly. It also removed "not", because "not" is extremely common. The review flipped from negative to positive.
Standard lists were built for search engines, where dropping "not" is fine — nobody searches "hotels not near airport" expecting the "not" to work. For meaning-sensitive tasks like sentiment, the same list quietly destroys the label.
A real example you have seen
Search engines are why this idea exists. Early engines ignored words like "the" and "a" in your query entirely. Searching "the taj hotel" and "taj hotel" gave the same results. The index stayed a fraction of the size. Modern engines are subtler, but your library's catalogue search almost certainly still sifts stopwords.
Remember this
- Stopwords = high-frequency filler words deleted before counting.
- Helps: search, topic modelling, word-count classifiers. Shrinks data, sharpens counts.
- Breaks: sentiment and meaning-sensitive tasks — standard lists contain "not" and "no".
What to learn next
- Stemming vs lemmatisation — the next classical reduction, at word-form level.
- Text classification — the bag-of-words pipeline these decisions feed.
- Word embeddings — the representation that replaced counting, where stopwords stay in.
Developer — Code and libraries.
Setup
pip install scikit-learnVerified with scikit-learn 1.7.2, CPU only.
Watch removal flip a review's meaning
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS, CountVectorizer
print("how many stopwords sklearn ships:", len(ENGLISH_STOP_WORDS))
print("a few:", sorted(ENGLISH_STOP_WORDS)[:8])
reviews = ["the food was not good", "the food was good"]
plain = CountVectorizer()
kept = plain.fit_transform(reviews).toarray()
print("\nkeep stopwords:", plain.get_feature_names_out())
print(kept)
filtered = CountVectorizer(stop_words="english")
gone = filtered.fit_transform(reviews).toarray()
print("\nremove stopwords:", filtered.get_feature_names_out())
print(gone)how many stopwords sklearn ships: 318 a few: ['a', 'about', 'above', 'across', 'after', 'afterwards', 'again', 'against'] keep stopwords: ['food' 'good' 'not' 'the' 'was'] [[1 1 1 1 1] [1 1 0 1 1]] remove stopwords: ['food' 'good'] [[1 1] [1 1]]
The walkthrough
The last matrix is the disaster, printed. A negative review and a positive review became identical rows: [1 1] and [1 1]. No classifier on earth can separate identical inputs. The word "not" — on the 318-word list — was the only difference, and it fell through the sieve.
stop_words="english" is one argument, and that convenience is the risk. One innocent-looking flag changed the labels' learnability. Scikit-learn's own documentation warns this list has known problems; it stays for compatibility.
A better default for classifiers: frequency cutoffs. CountVectorizer(max_df=0.9) drops words appearing in over 90% of your documents — a stopword list learned from your corpus, containing "the" but not "not" (in sentiment data, "not" is nowhere near 90% of documents). min_df=2 prunes the rare-typo tail from the other end.
N-grams are the other rescue. ngram_range=(1, 2) keeps "not good" as a single feature, so even with stopwords kept, negation becomes visible to the model.
Common mistakes
Applying stopword removal before transformer models. BERT-family models need function words — their grammar signal lives in exactly those tokens. Stopword removal there reduces accuracy. It belongs to the count-based world only.
Assuming one universal list. NLTK's list has 179 words including "not"; sklearn's has 318; spaCy's has 326. Results change with the list. Name the list in your experiment notes, or your results cannot be reproduced.
English lists on non-English text. Hindi, Tamil or German text passed through an English list loses nothing — and gains nothing. Each language needs its own list, and code-mixed text (Hinglish) fits no list cleanly.
Removing stopwords for embeddings. Sentence embedding models are trained on natural sentences. Feeding them sifted word-soup shifts them away from their training distribution.
Try it yourself
Add max_df=0.9 to the plain vectorizer and check which words survive on a slightly bigger review list you write yourself. Then set ngram_range=(1, 2) with stopwords kept and find the "not good" column in the output.
What to learn next
- Stemming vs lemmatisation — the next classical reduction, at word-form level.
- Text classification — the bag-of-words pipeline these decisions feed.
- Word embeddings — the representation that replaced counting, where stopwords stay in.
Researcher — Mathematics and papers.
The statistical basis
Zipf's law: word frequency is roughly proportional to 1/rank^s with s ≈ 1, so the head of the distribution is extreme — in the Brown corpus, "the" alone is ~7% of tokens and the top ~150 types cover roughly half of all tokens. Removal therefore buys large index and memory reductions for inverted-index retrieval and count matrices, at O(n) filter cost. Under TF-IDF weighting the ranking effect of head words is already suppressed — IDF(w) = log(N/df_w) approaches zero as df_w → N — which is why modern lexical systems (BM25 in Lucene/Elasticsearch) index stopwords and let the weighting handle them, keeping phrase queries ("to be or not to be") answerable.
Information-theoretic view
The removal decision is a claim that I(X_w; Y) ≈ 0 — the word's occurrence carries no mutual information about the label. That claim is corpus- and task-dependent: for topical classification it approximately holds for determiners; for sentiment, negators carry high mutual information; for authorship attribution, function-word frequencies are the primary signal (Mosteller and Wallace, 1964, the Federalist Papers study — stopwords settled the disputed authorship). Data-driven alternatives replace fixed lists with corpus statistics: document-frequency thresholds, or ranking words by pointwise mutual information with labels and pruning the bottom.
Evidence in modern pipelines
For neural models the empirical literature is consistent: stopword removal is neutral-to-harmful for transformer fine-tuning, since pretraining distributions include function words and positional structure. For topic models, Schofield, Magnusson and Mimno (2017), Pulling out the stops, show LDA topic quality is largely insensitive to pre-removal — post-hoc filtering of topic-word lists achieves the same readability, and aggressive pre-removal can distort topic-word distributions. In dense retrieval, queries pass whole into the encoder; the stopword concept survives only in hybrid lexical components and in keyword-extraction heuristics (RAKE explicitly uses stopwords as phrase delimiters — Rose et al., 2010).
Multilingual reality
List quality varies enormously across languages; agglutinative and morphologically rich languages (Tamil, Finnish, Turkish) express function morphemes as affixes, making token-level lists structurally inadequate — the equivalent operation is morphological analysis, next lesson's subject. For code-mixed text (Hinglish), token-level language ID must precede any list application, and published lists barely exist; frequency cutoffs computed on the target corpus remain the defensible default.
What to learn next
- Stemming vs lemmatisation — the next classical reduction, at word-form level.
- Text classification — the bag-of-words pipeline these decisions feed.
- Word embeddings — the representation that replaced counting, where stopwords stay in.