Text Preprocessing

Stemming vs lemmatisation

Both shrink word variants down to one shared form — stemming by blindly chopping endings, lemmatisation by actually knowing the language.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Stemming and lemmatisation both reduce word variants — "studies", "studying", "studied" — to one shared form, so a computer treats them as the same word.

Think of two ways to trim vegetables. Stemming is a fast chop with a big knife: cut the ends off everything the same way, no looking. Lemmatisation is a cook who knows each vegetable: peel this one, de-seed that one, leave the third whole. One is fast and rough; the other is careful and right.

Why it exists

Counting words — the engine behind search and text classification — scatters evidence across grammar. "Study", "studies", "studying" and "studied" get four separate counts for one idea. Search for "studying" and miss the page that says "studies".

Two families of fixes exist:

  • Stemming chops word endings by rules: remove "-ing", remove "-es", and so on. No dictionary. The result is often not a real word. "Studies" becomes "studi". But all four variants land on the same not-a-word, and that is all counting needs.
  • Lemmatisation looks words up and returns the true dictionary form, called the lemma. It knows "was" comes from "be", "mice" from "mouse", "better" from "well". It needs a dictionary, grammar knowledge, and context — so it is heavier and slower.

How it works

word        stemmer (chop rules)     lemmatiser (dictionary + grammar)
studies  →  studi                    study
was      →  wa                       be
mice     →  mice                     mouse
better   →  better                   well

The stemmer never knew "was" and "be" are related — no rule for chopping "-as" exists. The lemmatiser knew, because it has the dictionary. But the stemmer processed the words thousands of times faster.

A real example you have seen

Search a shopping app for "running shoes" and get results titled "Run Shoe – Men". The index stemmed both your query and the product titles to shared roots like "run" and "shoe", so the grammar difference stopped mattering. Library catalogues, job portals and legal search tools do the same.

Remember this

  • Both map word variants to one shared form so evidence pools.
  • Stemming: rule-based chopping. Fast, crude, output may not be a real word.
  • Lemmatisation: dictionary + grammar. Slower, accurate, output is a real word.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install nltk spacy
python -m spacy download en_core_web_sm

The spaCy model is a ~12 MB one-time download. Verified with NLTK 3.10 and spaCy 3.8, CPU only.

Chop versus look up, side by side

stem_vs_lemma.py
import spacy
from nltk.stem import PorterStemmer

words = ["studies", "studying", "better", "was", "mice", "meeting"]

stem = PorterStemmer().stem
print("word      stem       lemma")
nlp = spacy.load("en_core_web_sm")
doc = nlp(" ".join(words))
for w, tok in zip(words, doc):
    print(f"{w:9s} {stem(w):10s} {tok.lemma_}")
Output
word      stem       lemma
studies   studi      study
studying  studi      study
better    better     well
was       wa         be
mice      mice       mouse
meeting   meet       meeting

The walkthrough

Row by row, the trade is visible. "studies" and "studying" both stem to "studi" — ugly, but matching, which is the point. "was" stems to "wa" (the rule stripped a trailing "s") while the lemmatiser returns "be". "mice" defeats the stemmer entirely; irregular forms have no chopping rule.

The last row is the deep one. The stemmer chopped "meeting" to "meet". The lemmatiser kept "meeting" — because in this word list, spaCy's tagger read it as a noun ("a meeting"), whose lemma is "meeting". As a verb ("we are meeting") the lemma would be "meet". Lemmatisation depends on context; stemming cannot even see context.

Why the odd loop structure. spaCy wants full text, not single words — context is its whole advantage. Joining the words gives its tagger at least some context. Note that this also means lemmas can change if you reorder the list.

Choosing in practice:

you are buildingpickwhy
search index over lots of textstemmingspeed; users never see the stems
features users will read (word clouds, topics)lemmatisation"studi" in a report looks broken
transformer model pipelineneithersubword tokenizers replace both

Common mistakes

Stemming text bound for a transformer. BERT-family models bring their own subword handling — see tokenization — and were trained on unstemmed text. Stemming first hurts them.

Mismatched treatment of query and index. Stem the documents but not the search query and "studies" no longer finds "studi". Same rule as all preprocessing: one function, applied everywhere.

Expecting NLTK's WordNetLemmatizer to work like spaCy's. It defaults to treating every word as a noun unless you pass a part-of-speech tag yourself — lemmatize("was") returns "was", not "be". It also needs nltk.download("wordnet") first. spaCy tags automatically, which is why it is shown here.

Over-stemming collisions. Porter maps "university" and "universe" toward the same stem family. Pooled evidence sometimes means wrongly pooled. If precision matters, lemmatise.

Try it yourself

Run the sentence "We are meeting tomorrow to discuss the meeting" through nlp() and print each token's lemma. Watch the same word get two different lemmas from context. Then stem the same sentence and note what stemming does to both occurrences.

What to learn next

Researcher — Mathematics and papers.

Stemming algorithms

The Porter stemmer (Porter, 1980, An algorithm for suffix stripping) applies ~60 rules in five phases, each conditioned on a crude syllable measure m (vowel-consonant alternations) — e.g. (m>0) SSES → SS, (m>1) EMENT → ∅. It is a finite-state transducer: O(1) per word, no lexicon, deterministic. Porter2/Snowball (Porter, 2001) refines the rules and generalises the framework across ~30 languages; the Lancaster stemmer iterates more aggressively (higher collision rate). Formal failure modes: over-stemming (distinct lexemes conflated: "university"/"universe") and under-stemming (inflections of one lexeme kept apart: "mice"/"mouse"); Paice (1994) defines OI/UI indices quantifying both.

Lemmatisation as classification

Modern lemmatisers are not dictionary lookups. spaCy v3's default English lemmatiser is rule-plus-table conditioned on the predicted POS tag (hence its pipeline dependency on the tagger); its neural edit tree lemmatizer (from spaCy 3.3, following Müller et al., 2015, Joint lemmatization and morphological tagging with Lemming) classifies each token over a set of edit trees induced from training pairs (form → lemma), which handles unseen forms by analogy. Contextual accuracy on UD English treebanks runs ~97–98%; morphologically rich languages (Turkish, Finnish, Hindi) benefit far more, since a single lexeme can surface in hundreds of forms — there, lemmatisation (or full morphological analysis) is close to mandatory for sparse models.

Retrieval evidence

Classic IR results: stemming yields small but consistent recall gains for English (Harman, 1991, found near-zero average effect; later TREC analyses show topic-dependent variance), and much larger gains for highly inflected languages — Braschler and Ripplinger (2004) report double-digit improvements for German. Krovetz (1993) built the derivational/inflectional middle ground (the KSTEM lemma-aware stemmer), showing accuracy-sensitive conflation helps precision. Modern BM25 systems still ship Snowball by default; dense retrievers need none of it.

Subword tokenization made both optional

BPE (Sennrich et al., 2016), WordPiece (Schuster and Nakajima, 2012) and SentencePiece/Unigram-LM (Kudo, 2018) solve the sparsity problem statistically: frequent stems and affixes emerge as subword units ("study" + "##ing") without linguistic rules, and unseen forms decompose instead of becoming OOV. This is why the entire stemming/lemmatisation question vanishes inside transformer pipelines and survives at the feature-engineering layer: search indexes, TF-IDF baselines, keyword statistics, and anywhere a human reads the resulting vocabulary. Tokenization covers the mechanics.

What to learn next