Language and NLP Section 042
Text Preprocessing
Getting raw text into a shape a model can use — the quiet step that decides how well everything after it works.
9 of 9 lessons published Three reading levels on every lesson
Start with “Text normalisation: case, punctuation and whitespace”
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- Text normalisation: case, punctuation and whitespace
- Unicode, UTF-8 and mojibake
- Stopwords: when removing them helps and when it breaks things
- Stemming vs lemmatisation
- Splitting text into sentences
- Regular expressions for text work
- spaCy pipelines
- Getting clean text out of PDFs
- Detecting the language of a text