Stemming vs lemmatisation
Both shrink word variants down to one shared form — stemming by blindly chopping endings, lemmatisation by actually knowing the language.
- 7 min read
- 3 reading levels
- Published
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Stemming and lemmatisation both reduce word variants — "studies", "studying", "studied" — to one shared form, so a computer treats them as the same word.
Think of two ways to trim vegetables. Stemming is a fast chop with a big knife: cut the ends off everything the same way, no looking. Lemmatisation is a cook who knows each vegetable: peel this one, de-seed that one, leave the third whole. One is fast and rough; the other is careful and right.
Why it exists
Counting words — the engine behind search and text classification — scatters evidence across grammar. "Study", "studies", "studying" and "studied" get four separate counts for one idea. Search for "studying" and miss the page that says "studies".
Two families of fixes exist:
- Stemming chops word endings by rules: remove "-ing", remove "-es", and so on. No dictionary. The result is often not a real word. "Studies" becomes "studi". But all four variants land on the same not-a-word, and that is all counting needs.
- Lemmatisation looks words up and returns the true dictionary form, called the lemma. It knows "was" comes from "be", "mice" from "mouse", "better" from "well". It needs a dictionary, grammar knowledge, and context — so it is heavier and slower.
How it works
word stemmer (chop rules) lemmatiser (dictionary + grammar)
studies → studi study
was → wa be
mice → mice mouse
better → better wellThe stemmer never knew "was" and "be" are related — no rule for chopping "-as" exists. The lemmatiser knew, because it has the dictionary. But the stemmer processed the words thousands of times faster.
A real example you have seen
Search a shopping app for "running shoes" and get results titled "Run Shoe – Men". The index stemmed both your query and the product titles to shared roots like "run" and "shoe", so the grammar difference stopped mattering. Library catalogues, job portals and legal search tools do the same.
Remember this
- Both map word variants to one shared form so evidence pools.
- Stemming: rule-based chopping. Fast, crude, output may not be a real word.
- Lemmatisation: dictionary + grammar. Slower, accurate, output is a real word.
What to learn next
- Splitting text into sentences — the next structural cut, one level above words.
- spaCy pipelines — the machinery that produced those lemmas, opened up.
- Tokenization — the subword approach that made both techniques optional.
Developer — Code and libraries.
Setup
pip install nltk spacy
python -m spacy download en_core_web_smThe spaCy model is a ~12 MB one-time download. Verified with NLTK 3.10 and spaCy 3.8, CPU only.
Chop versus look up, side by side
import spacy
from nltk.stem import PorterStemmer
words = ["studies", "studying", "better", "was", "mice", "meeting"]
stem = PorterStemmer().stem
print("word stem lemma")
nlp = spacy.load("en_core_web_sm")
doc = nlp(" ".join(words))
for w, tok in zip(words, doc):
print(f"{w:9s} {stem(w):10s} {tok.lemma_}")word stem lemma studies studi study studying studi study better better well was wa be mice mice mouse meeting meet meeting
The walkthrough
Row by row, the trade is visible. "studies" and "studying" both stem to "studi" — ugly, but matching, which is the point. "was" stems to "wa" (the rule stripped a trailing "s") while the lemmatiser returns "be". "mice" defeats the stemmer entirely; irregular forms have no chopping rule.
The last row is the deep one. The stemmer chopped "meeting" to "meet". The lemmatiser kept "meeting" — because in this word list, spaCy's tagger read it as a noun ("a meeting"), whose lemma is "meeting". As a verb ("we are meeting") the lemma would be "meet". Lemmatisation depends on context; stemming cannot even see context.
Why the odd loop structure. spaCy wants full text, not single words — context is its whole advantage. Joining the words gives its tagger at least some context. Note that this also means lemmas can change if you reorder the list.
Choosing in practice:
| you are building | pick | why |
|---|---|---|
| search index over lots of text | stemming | speed; users never see the stems |
| features users will read (word clouds, topics) | lemmatisation | "studi" in a report looks broken |
| transformer model pipeline | neither | subword tokenizers replace both |
Common mistakes
Stemming text bound for a transformer. BERT-family models bring their own subword handling — see tokenization — and were trained on unstemmed text. Stemming first hurts them.
Mismatched treatment of query and index. Stem the documents but not the search query and "studies" no longer finds "studi". Same rule as all preprocessing: one function, applied everywhere.
Expecting NLTK's WordNetLemmatizer to work like spaCy's. It defaults to treating every word as a noun unless you pass a part-of-speech tag yourself — lemmatize("was") returns "was", not "be". It also needs nltk.download("wordnet") first. spaCy tags automatically, which is why it is shown here.
Over-stemming collisions. Porter maps "university" and "universe" toward the same stem family. Pooled evidence sometimes means wrongly pooled. If precision matters, lemmatise.
Try it yourself
Run the sentence "We are meeting tomorrow to discuss the meeting" through nlp() and print each token's lemma. Watch the same word get two different lemmas from context. Then stem the same sentence and note what stemming does to both occurrences.
What to learn next
- Splitting text into sentences — the next structural cut, one level above words.
- spaCy pipelines — the machinery that produced those lemmas, opened up.
- Tokenization — the subword approach that made both techniques optional.
Researcher — Mathematics and papers.
Stemming algorithms
The Porter stemmer (Porter, 1980, An algorithm for suffix stripping) applies ~60 rules in five phases, each conditioned on a crude syllable measure m (vowel-consonant alternations) — e.g. (m>0) SSES → SS, (m>1) EMENT → ∅. It is a finite-state transducer: O(1) per word, no lexicon, deterministic. Porter2/Snowball (Porter, 2001) refines the rules and generalises the framework across ~30 languages; the Lancaster stemmer iterates more aggressively (higher collision rate). Formal failure modes: over-stemming (distinct lexemes conflated: "university"/"universe") and under-stemming (inflections of one lexeme kept apart: "mice"/"mouse"); Paice (1994) defines OI/UI indices quantifying both.
Lemmatisation as classification
Modern lemmatisers are not dictionary lookups. spaCy v3's default English lemmatiser is rule-plus-table conditioned on the predicted POS tag (hence its pipeline dependency on the tagger); its neural edit tree lemmatizer (from spaCy 3.3, following Müller et al., 2015, Joint lemmatization and morphological tagging with Lemming) classifies each token over a set of edit trees induced from training pairs (form → lemma), which handles unseen forms by analogy. Contextual accuracy on UD English treebanks runs ~97–98%; morphologically rich languages (Turkish, Finnish, Hindi) benefit far more, since a single lexeme can surface in hundreds of forms — there, lemmatisation (or full morphological analysis) is close to mandatory for sparse models.
Retrieval evidence
Classic IR results: stemming yields small but consistent recall gains for English (Harman, 1991, found near-zero average effect; later TREC analyses show topic-dependent variance), and much larger gains for highly inflected languages — Braschler and Ripplinger (2004) report double-digit improvements for German. Krovetz (1993) built the derivational/inflectional middle ground (the KSTEM lemma-aware stemmer), showing accuracy-sensitive conflation helps precision. Modern BM25 systems still ship Snowball by default; dense retrievers need none of it.
Subword tokenization made both optional
BPE (Sennrich et al., 2016), WordPiece (Schuster and Nakajima, 2012) and SentencePiece/Unigram-LM (Kudo, 2018) solve the sparsity problem statistically: frequent stems and affixes emerge as subword units ("study" + "##ing") without linguistic rules, and unseen forms decompose instead of becoming OOV. This is why the entire stemming/lemmatisation question vanishes inside transformer pipelines and survives at the feature-engineering layer: search indexes, TF-IDF baselines, keyword statistics, and anywhere a human reads the resulting vocabulary. Tokenization covers the mechanics.
What to learn next
- Splitting text into sentences — the next structural cut, one level above words.
- spaCy pipelines — the machinery that produced those lemmas, opened up.
- Tokenization — the subword approach that made both techniques optional.