Text normalisation: case, punctuation and whitespace
Lowercasing, punctuation stripping and whitespace collapsing decide what your model counts as "the same word" — and each step quietly destroys something.
- 6 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Text normalisation makes different spellings of the same thing look identical, so a computer stops treating them as strangers.
Think of washing vegetables before cooking. Two tomatoes from different shops arrive with different dirt, stickers and leaves. After washing, they are interchangeable ingredients. Normalisation is the wash: "GREAT", "great" and " great!! " become one ingredient.
Why it exists
Computers compare text letter by letter. To you, "Delhi", "DELHI" and "delhi " are one city. To a computer they are three unrelated strings, because capital letters, small letters and trailing spaces are different characters.
Left unwashed, your data scatters. A review dataset ends up with "good", "Good", "GOOD", "good!" counted as four separate words, each with a quarter of the evidence. Models learn slower and search boxes miss matches.
Normalisation is the umbrella word for the cleanup passes: lowercasing, trimming spaces, unifying punctuation, standardising look-alike characters.
How it works
" GREAT product!!! "
│ lowercase
" great product!!! "
│ remove punctuation
" great product "
│ collapse whitespace
"great product"Each pass is small and mechanical. The order matters little; which passes you apply matters a lot. Every pass erases a difference — and sometimes that difference carried meaning. Removing punctuation turns "10/10" into "1010". Lowercasing merges "US" the country into "us" the word. Normalisation is always a trade: fewer strangers, but some lost meaning.
A real example you have seen
Search boxes do this constantly. Type "IPHONE 15 " with stray spaces and capitals into a shopping app, and you still get iPhone 15 results. The app normalised your query and its product names into the same washed form before comparing. Without that wash, sloppy typing would return nothing.
Remember this
- Normalisation makes variant spellings identical so evidence stops scattering.
- Every pass deletes information — decide per task, not by habit.
- Standard passes: lowercase, trim and collapse spaces, handle punctuation deliberately.
What to learn next
- Unicode, UTF-8 and mojibake — the character-level chaos NFKC was quietly fixing.
- Stopwords — the next deletion decision, one level up.
- Tokenization — how modern models slice text after (or instead of) your washing.
Developer — Code and libraries.
Setup
No installs — everything here is Python's standard library. Verified on Python 3.10.
The four standard passes
import re
import string
import unicodedata
# \u00a0 is a non-breaking space; web pages are full of them
raw = " GREAT product!!! Very fast\u00a0delivery… 10/10 ❤️ "
step1 = unicodedata.normalize("NFKC", raw) # unify look-alike characters
step2 = step1.casefold() # lowercase, but stronger
step3 = step2.translate(str.maketrans("", "", string.punctuation))
step4 = re.sub(r"\s+", " ", step3).strip() # collapse runs of whitespace
for name, s in [("raw", raw), ("nfkc", step1), ("casefold", step2),
("no punct", step3), ("clean", step4)]:
print(f"{name:9s} {s!r}")raw ' GREAT product!!! Very fast\xa0delivery… 10/10 ❤️ ' nfkc ' GREAT product!!! Very fast delivery... 10/10 ❤️ ' casefold ' great product!!! very fast delivery... 10/10 ❤️ ' no punct ' great product very fast delivery 1010 ❤️ ' clean 'great product very fast delivery 1010 ❤️'
The walkthrough
NFKC earned its keep twice on line one. The invisible \xa0 — a non-breaking space, everywhere in web-scraped text — became an ordinary space. The single … character became three ordinary dots. More on these look-alikes in Unicode and mojibake.
casefold() over lower(). For English they act the same. casefold also handles languages with unusual case rules — German ß becomes ss, so "STRASSE" and "straße" match. Costs nothing; occasionally saves you.
Look at what broke: 10/10 became 1010. The punctuation pass deleted the slash and invented a fake number. string.punctuation covers only ASCII punctuation — it deleted ! but not … (which NFKC had conveniently converted first). Blanket punctuation stripping is the most destructive pass; consider replacing punctuation with spaces instead: re.sub(r"[^\w\s]", " ", text).
The emoji survived. It is neither punctuation nor whitespace. For sentiment tasks, keep it — a heart is signal. For others, strip by Unicode category. Every survivor is a decision.
Common mistakes
Normalising by reflex before every model. Modern transformer tokenizers — see tokenization — expect raw-ish text and handle case themselves. Aggressive washing hurts them. Heavy normalisation belongs to the bag-of-words world: CountVectorizer, TF-IDF, search indexes.
Lowercasing named entities away. "March" the month and "march" the verb, "US" and "us", "Apple" and "apple" merge irreversibly. If your task involves names — NER, for one — keep case, or keep a cased copy.
Deleting punctuation instead of replacing with spaces. "delhi-based" becomes "delhibased" — a brand-new word no dictionary contains. Replace-with-space yields "delhi based", two real words.
Forgetting the query side. Normalise your index but not the search query (or train-time text but not serve-time text) and nothing matches. The wash must be one shared function applied everywhere, always.
Try it yourself
Run the pipeline on "Wi-Fi doesn't work :( 2/5" and inspect each stage. Decide which passes you would keep for a review-sentiment model, then rewrite step 3 as replace-with-space and compare the final tokens.
What to learn next
- Unicode, UTF-8 and mojibake — the character-level chaos NFKC was quietly fixing.
- Stopwords — the next deletion decision, one level up.
- Tokenization — how modern models slice text after (or instead of) your washing.
Researcher — Mathematics and papers.
Normalisation as an equivalence relation
Each pass defines a mapping φ: Σ* → Σ* over strings; the pipeline's composition induces equivalence classes [x] = {x′ : φ(x′) = φ(x)}. Feature extractors then operate on quotient space. The design question is where to place the partition between linguistic variance (case, spacing, encoding artefacts) and semantic signal. Collapsing too far raises Bayes error irrecoverably — information deleted before featurisation cannot be restored downstream. The trade is measurable: vocabulary size and OOV rate fall monotonically with aggressiveness, while task metrics rise then fall; the optimum is task- and model-dependent.
Empirical findings
For sparse lexical models (TF-IDF, BM25), case-folding and whitespace canonicalisation reliably help; wholesale punctuation removal is mixed, and negation/emoticon handling dominates on sentiment corpora. For pretrained transformers the picture inverts: models are trained on lightly normalised text, and mismatched preprocessing is distribution shift. The BERT release itself ships cased and uncased variants precisely because the choice is task-dependent (Devlin et al., 2019): uncased typically edges ahead on classification; cased wins on NER. Subword tokenizers (Sennrich et al., 2016, BPE; Kudo and Richardson, 2018, SentencePiece) absorb spacing and rare-word variance internally — SentencePiece even encodes the space as ▁, making external whitespace collapsing part of the tokenizer's contract rather than yours. Tokenizer internals continues this thread.
Unicode canonical forms
unicodedata.normalize implements UAX #15: NFC/NFD (canonical composition/decomposition) and NFKC/NFKD (adding compatibility mappings: ligatures fi → fi, fullwidth A → A, … → ...). NFKC is lossy by design — superscripts and subscripts flatten (x² → x2), which is why it belongs in matching pipelines, not in archival storage. Case folding is likewise defined by the Unicode standard (full case folding, str.casefold in Python) and is not a bijection; folded text cannot be un-folded. Locale traps are real: Turkish dotless ı/İ breaks naive lowercasing, motivating locale-aware folds (ICU) in multilingual systems.
Cost
All passes are single-scan: O(n) in characters with small constants; str.translate uses a table lookup per character and outruns equivalent regex substitution by a wide margin on large corpora. At corpus scale the practical costs are (1) doing it twice inconsistently across train/serve — a classic skew bug, and (2) doing it in Python row-by-row instead of vectorised (pandas.Series.str, or better, inside the vectorizer's preprocessor hook so it is versioned with the model).
What to learn next
- Unicode, UTF-8 and mojibake — the character-level chaos NFKC was quietly fixing.
- Stopwords — the next deletion decision, one level up.
- Tokenization — how modern models slice text after (or instead of) your washing.