Skip to main content
Learn AI Create the Future — with Pranay
  • Learn
  • Paths
  • Projects
  • Practice
  • Glossary
  • Errors
  • LeetCode
  • Main site

Main site →
Select language

Translation is unavailable right now. Your browser’s own “Translate page” usually works.

Machine translation by Google

Topics

  • 🧱 Foundations 2 sections
  • 🤖 Machine Learning 11 sections
  • 📊 Statistics and Experiments 4 sections
  • 🧭 Doing the Work 9 sections
  • 🧠 Deep Learning 11 sections
  • 🧰 Libraries and Frameworks 4 sections
  • 💬 Language and NLP 15 sections
    • Natural Language Processing 8 lessons
    • Text Preprocessing 9 lessons
      • Overview
      • Text normalisation: case, punctuation and whitespace
      • Unicode, UTF-8 and mojibake
      • Stopwords: when removing them helps and when it breaks things
      • Stemming vs lemmatisation
      • Splitting text into sentences
      • Regular expressions for text work
      • spaCy pipelines
      • Getting clean text out of PDFs
      • Detecting the language of a text
    • Classical NLP That Still Works 7 lessons
    • The Embedding Family 10 lessons
    • The BERT Family 10 lessons
    • Sequence Labelling and Structure 10 lessons
    • Topic Modelling and Text Clustering 9 lessons
    • Semantic Search and Reranking 11 lessons
    • Chunking and Long Documents 8 lessons
    • Question Answering 10 lessons
    • Summarisation 10 lessons
    • Sentiment, Opinion and Text Mining 9 lessons
    • Multilingual and Indic NLP 12 lessons
    • Messy Real-World Text 10 lessons
    • Evaluating Text Systems 7 lessons
  • 👁️ Computer Vision 13 sections
  • 🎙️ Speech, Forecasting and Recommenders 3 sections
  • ✨ Generative AI and LLMs 4 sections
  • 🔬 Inside a Transformer 11 sections
  • ⚙️ Production and MLOps 16 sections
  • 🌍 AI in the Real World 12 sections
  • ⚖️ Safety, Ethics and Law 2 sections

See all topics

  1. Home
  2. Learn
  3. Text Preprocessing

💬 Language and NLP · Section 042

🧹 Text Preprocessing

Getting raw text into a shape a model can use — the quiet step that decides how well everything after it works.

Every lesson in this section is written by Pranay Mahendrakar.

9 of 9 lessons published · Three reading levels on every lesson

Start with “Text normalisation: case, punctuation and whitespace”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. 01 Text normalisation: case, punctuation and whitespace
  2. 02 Unicode, UTF-8 and mojibake
  3. 03 Stopwords: when removing them helps and when it breaks things
  4. 04 Stemming vs lemmatisation
  5. 05 Splitting text into sentences
  6. 06 Regular expressions for text work
  7. 07 spaCy pipelines
  8. 08 Getting clean text out of PDFs
  9. 09 Detecting the language of a text
Previous Natural Language Processing Next Classical NLP That Still Works

Learn AI

Learn Artificial Intelligence the easy way.

Learn. Build. Ask. Share.

Free forever. No sign-up, no ads, no tracking.

Learn

  • All topics
  • Learning paths
  • Projects
  • Practice playground
  • AI glossary
  • Error database
  • LeetCode solutions

Start here

  • Python for AI
  • Mathematics for AI
  • Machine Learning
  • Classic Algorithms in Depth
  • Linear Models and Regularisation
  • Ensembles and Gradient Boosting

This site

  • Search
  • New lessons feed
  • Sitemap
  • About Pranay Mahendrakar
  • pranaymahendrakar.com

© 2026 Pranay Mahendrakar. Written for people who are starting from zero.

10 new lessons every day. If something here is wrong or confusing, that is worth fixing — say so.