Topic Modelling and Text Clustering

What is topic modelling?

Topic modelling sorts a large pile of text into groups by what each piece is about, using only patterns in the words, with no one reading or labelling anything by hand.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Topic modelling sorts a pile of documents by what they are about — without anyone reading them first.

Picture a shoebox stuffed with old family photographs, dumped on the floor. Nobody wrote captions on the back. You start sorting them into rough piles anyway. One pile holds weddings. One holds birthday parties. One holds holiday trips. You never labelled a single photo. You noticed which ones felt alike, and let the piles form on their own.

That is exactly what topic modelling does to text. Feed it thousands of reviews, articles or support tickets. It forms piles based on which words tend to show up together. No labels required.

Why it exists

Reading every document by hand does not scale. A company with two million reviews cannot read each one by hand. Nobody has time to write "shipping delays" on two million sticky notes.

Topic modelling is software that reads the pile for you. It reports back: "30% of these mention shipping delays, 20% mention pricing, 15% mention a broken app." It does this using nothing but the words themselves — no training labels, no human tagging beforehand. That is why it counts as unsupervised learning: learning structure from data with no answer key.

How it works

2,000 customer reviews
        |
        v
  count which words appear together across documents
        |
        v
  group 1: "battery", "charge", "phone"      -> looks like "battery life"
  group 2: "delivery", "late", "courier"     -> looks like "shipping"
  group 3: "refund", "support", "waiting"    -> looks like "customer service"
        |
        v
  each review gets a rough share of each group
  (a review can be 70% "battery" and 30% "customer service")

Notice the last line. A document is rarely 100% one topic. A review complaining about a dead battery and a slow refund belongs partly to both piles. Topic models embrace that mix instead of forcing one label per document.

A real example you have seen

Open the Play Store and scroll to the "what people are saying" summary under a popular app. Someone built that by running topic modelling over thousands of raw reviews. It pulled out the recurring themes automatically — "battery drain," "great customer support," "crashes on startup." News aggregators do the same thing. They cluster hundreds of articles about one event into a single story card.

Remember this

  • Topic modelling groups documents by shared word patterns, with no one labelling anything first.
  • It is unsupervised: there is no answer key telling it what the topics should be called.
  • A document usually belongs to several topics at once, each in a different proportion.

What to learn next

Developer — Code and libraries.

Two families of algorithms dominate classic topic modelling: LDA (Latent Dirichlet Allocation, a probabilistic model) and NMF (Non-negative Matrix Factorization, a matrix-algebra model). Both get their own lesson later in this section. Here, run a first end-to-end example so the shape of the problem is concrete before the theory arrives.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

Twelve tiny documents, three real topics

first_topic_model.py
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import NMF

docs = [
    "the batsman hit a six to win the match in the final over",
    "the bowler took three wickets as the batsman walked back after the match",
    "the team captain praised his batsman after the thrilling match",
    "the umpire signalled a boundary as the batsman completed the run",
    "add chopped onion to hot oil then simmer the curry with salt",
    "the recipe needs turmeric, salt and slow simmering of the curry",
    "fry the spices in oil before you add the curry paste for dinner",
    "the chef seasoned the curry with salt and a spoon of turmeric",
    "the candidate promised new roads and hospitals before the election",
    "voters lined up outside the polling booth to vote in the election",
    "the election result was announced after votes were counted all night",
    "the losing candidate conceded the election after the final count",
]

vectorizer = TfidfVectorizer(stop_words="english")
X = vectorizer.fit_transform(docs)          # each document -> a row of word weights

nmf = NMF(n_components=3, random_state=0, max_iter=500)
doc_topics = nmf.fit_transform(X)           # each document -> its share of each topic

words = vectorizer.get_feature_names_out()
for i, topic in enumerate(nmf.components_):
    top = [words[j] for j in topic.argsort()[-5:][::-1]]
    print(f"topic {i}: {', '.join(top)}")

print()
for doc, row in zip(docs, doc_topics):
    winner = row.argmax()
    print(f"topic {winner}  {doc[:40]}...")
Output
topic 0: curry, salt, turmeric, oil, add
topic 1: batsman, match, hit, win, final
topic 2: election, candidate, count, conceded, losing

topic 1  the batsman hit a six to win the match i...
topic 1  the bowler took three wickets as the bat...
topic 1  the team captain praised his batsman aft...
topic 1  the umpire signalled a boundary as the b...
topic 0  add chopped onion to hot oil then simmer...
topic 0  the recipe needs turmeric, salt and slow...
topic 0  fry the spices in oil before you add the...
topic 0  the chef seasoned the curry with salt an...
topic 2  the candidate promised new roads and hos...
topic 2  voters lined up outside the polling boot...
topic 2  the election result was announced after ...
topic 2  the losing candidate conceded the electi...

No one told the code "cricket," "cooking" or "elections." It found three groups of words that repeat together, and every one of the twelve documents landed in the group a human would put it in too.

The walkthrough

TfidfVectorizer turns text into numbers first. Every topic model needs a matrix, not raw sentences. Each row here is a document; each column is a word; each value says how distinctive that word is for that document. This is the same idea covered in embeddings, reduced to word counts rather than learned vectors.

n_components=3 is a guess you make up front. You told the model there are three topics. Real corpora don't come with that number attached — the whole next lesson in this section, choosing how many topics, is dedicated to that guess.

nmf.components_ holds the topics themselves. Each row is one topic's "recipe" — a weight for every word in the vocabulary. Sorting a row and reading the top values is how you read what a topic is "about."

doc_topics holds each document's mixture. argmax here only picks the single strongest topic to print a clean list — in reality every document carries a small weight in the other two topics as well.

Common mistakes

Forgetting to remove stop words. Without stop_words="english", "the" and "and" dominate every topic, because they appear in every document. Every topic ends up looking identical.

Expecting clean topics from real-world data on the first try. This toy example works this well because the three groups barely share any vocabulary. Real corpora — support tickets, product reviews — overlap far more, and topics come out fuzzier. That is normal, not a bug in your code.

Treating the topic count as ground truth. n_components=3 here happens to match how the data was written. On a real corpus you almost never know the "true" number, because there usually isn't one — themes blend into each other.

Try it yourself

Add a thirteenth document that mixes two themes on purpose, such as "the winning candidate celebrated with a cricket match against his campaign team", and re-run. Watch its topic weights — it should no longer land cleanly on one topic.

What to learn next

Researcher — Mathematics and papers.

The formal problem

Given a corpus of n documents over a vocabulary of V words, topic modelling seeks two things simultaneously:

  • A set of k topics, each a distribution or weighting over the V words.
  • For each document, a mixture describing how much of each topic it draws on.

This is a form of dimensionality reduction and soft clustering performed jointly: instead of assigning each document one hard cluster label, every document gets a k-dimensional weight vector, and every topic gets a V-dimensional word vector. The two families in this section differ mainly in what mathematical object they factor and what constraints they place on it.

LDA (next lesson) treats the whole thing as a generative probabilistic process over a Dirichlet-distributed mixture, fit by variational inference or Gibbs sampling.

NMF (two lessons ahead) treats the document-term matrix X (shape n x V) as a product of two non-negative matrices, X ~ W H, with W shape n x k and H shape k x V, fit by minimising reconstruction error directly.

Both approaches predate large language models by a decade or more and remain in production because they are cheap, fully interpretable — every topic is a literal list of words — and require no GPU.

Where this sits relative to modern methods

Since roughly 2019, embedding-based pipelines (BERTopic, covered later in this section) have added a third family: embed each document with a transformer, cluster the embeddings, then describe each cluster with its most distinctive words. This captures meaning beyond exact word overlap — "physician" and "doctor" pull documents into the same cluster even though they share no letters — at the cost of needing a neural network and losing the clean probabilistic guarantees LDA offers.

Key references

  • Blei, D., Ng, A. & Jordan, M. (2003). Latent Dirichlet Allocation. Journal of Machine Learning Research 3.
  • Lee, D. & Seung, H. (1999). Learning the parts of objects by non-negative matrix factorization. Nature 401.
  • Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794

What to learn next

What to learn next

These follow on from what you just read.

  • Topic Modelling and Text Clustering

    Latent Dirichlet Allocation

    LDA imagines every document as a random mix of topics and every topic as a random mix of words, then works backward from the finished text to guess both mixes.

  • Topic Modelling and Text Clustering

    NMF for topics

    NMF finds topics by breaking a document-word table into two smaller tables that are never allowed to go negative, avoiding a lot of the guesswork LDA needs.

  • Topic Modelling and Text Clustering

    BERTopic

    BERTopic finds topics by embedding documents with a language model first and clustering the embeddings, so it groups by meaning instead of only shared exact words.