Topic Modelling and Text Clustering
What is topic modelling?
Topic modelling sorts a large pile of text into groups by what each piece is about, using only patterns in the words, with no one reading or labelling anything by hand.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Topic modelling sorts a pile of documents by what they are about — without anyone reading them first.
Picture a shoebox stuffed with old family photographs, dumped on the floor. Nobody wrote captions on the back. You start sorting them into rough piles anyway. One pile holds weddings. One holds birthday parties. One holds holiday trips. You never labelled a single photo. You noticed which ones felt alike, and let the piles form on their own.
That is exactly what topic modelling does to text. Feed it thousands of reviews, articles or support tickets. It forms piles based on which words tend to show up together. No labels required.
Why it exists
Reading every document by hand does not scale. A company with two million reviews cannot read each one by hand. Nobody has time to write "shipping delays" on two million sticky notes.
Topic modelling is software that reads the pile for you. It reports back: "30% of these mention shipping delays, 20% mention pricing, 15% mention a broken app." It does this using nothing but the words themselves — no training labels, no human tagging beforehand. That is why it counts as unsupervised learning: learning structure from data with no answer key.
How it works
2,000 customer reviews
|
v
count which words appear together across documents
|
v
group 1: "battery", "charge", "phone" -> looks like "battery life"
group 2: "delivery", "late", "courier" -> looks like "shipping"
group 3: "refund", "support", "waiting" -> looks like "customer service"
|
v
each review gets a rough share of each group
(a review can be 70% "battery" and 30% "customer service")Notice the last line. A document is rarely 100% one topic. A review complaining about a dead battery and a slow refund belongs partly to both piles. Topic models embrace that mix instead of forcing one label per document.
A real example you have seen
Open the Play Store and scroll to the "what people are saying" summary under a popular app. Someone built that by running topic modelling over thousands of raw reviews. It pulled out the recurring themes automatically — "battery drain," "great customer support," "crashes on startup." News aggregators do the same thing. They cluster hundreds of articles about one event into a single story card.
Remember this
- Topic modelling groups documents by shared word patterns, with no one labelling anything first.
- It is unsupervised: there is no answer key telling it what the topics should be called.
- A document usually belongs to several topics at once, each in a different proportion.
What to learn next
- Latent Dirichlet Allocation — the classic probabilistic topic model, in depth.
- NMF for topics — the matrix-factorization alternative, run on this exact corpus.
- Clustering — the wider family of algorithms topic modelling belongs to.
Developer — Code and libraries.
Two families of algorithms dominate classic topic modelling: LDA (Latent Dirichlet Allocation, a probabilistic model) and NMF (Non-negative Matrix Factorization, a matrix-algebra model). Both get their own lesson later in this section. Here, run a first end-to-end example so the shape of the problem is concrete before the theory arrives.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU.
Twelve tiny documents, three real topics
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import NMF
docs = [
"the batsman hit a six to win the match in the final over",
"the bowler took three wickets as the batsman walked back after the match",
"the team captain praised his batsman after the thrilling match",
"the umpire signalled a boundary as the batsman completed the run",
"add chopped onion to hot oil then simmer the curry with salt",
"the recipe needs turmeric, salt and slow simmering of the curry",
"fry the spices in oil before you add the curry paste for dinner",
"the chef seasoned the curry with salt and a spoon of turmeric",
"the candidate promised new roads and hospitals before the election",
"voters lined up outside the polling booth to vote in the election",
"the election result was announced after votes were counted all night",
"the losing candidate conceded the election after the final count",
]
vectorizer = TfidfVectorizer(stop_words="english")
X = vectorizer.fit_transform(docs) # each document -> a row of word weights
nmf = NMF(n_components=3, random_state=0, max_iter=500)
doc_topics = nmf.fit_transform(X) # each document -> its share of each topic
words = vectorizer.get_feature_names_out()
for i, topic in enumerate(nmf.components_):
top = [words[j] for j in topic.argsort()[-5:][::-1]]
print(f"topic {i}: {', '.join(top)}")
print()
for doc, row in zip(docs, doc_topics):
winner = row.argmax()
print(f"topic {winner} {doc[:40]}...")topic 0: curry, salt, turmeric, oil, add topic 1: batsman, match, hit, win, final topic 2: election, candidate, count, conceded, losing topic 1 the batsman hit a six to win the match i... topic 1 the bowler took three wickets as the bat... topic 1 the team captain praised his batsman aft... topic 1 the umpire signalled a boundary as the b... topic 0 add chopped onion to hot oil then simmer... topic 0 the recipe needs turmeric, salt and slow... topic 0 fry the spices in oil before you add the... topic 0 the chef seasoned the curry with salt an... topic 2 the candidate promised new roads and hos... topic 2 voters lined up outside the polling boot... topic 2 the election result was announced after ... topic 2 the losing candidate conceded the electi...
No one told the code "cricket," "cooking" or "elections." It found three groups of words that repeat together, and every one of the twelve documents landed in the group a human would put it in too.
The walkthrough
TfidfVectorizer turns text into numbers first. Every topic model needs a matrix, not raw sentences. Each row here is a document; each column is a word; each value says how distinctive that word is for that document. This is the same idea covered in embeddings, reduced to word counts rather than learned vectors.
n_components=3 is a guess you make up front. You told the model there are three topics. Real corpora don't come with that number attached — the whole next lesson in this section, choosing how many topics, is dedicated to that guess.
nmf.components_ holds the topics themselves. Each row is one topic's "recipe" — a weight for every word in the vocabulary. Sorting a row and reading the top values is how you read what a topic is "about."
doc_topics holds each document's mixture. argmax here only picks the single strongest topic to print a clean list — in reality every document carries a small weight in the other two topics as well.
Common mistakes
Forgetting to remove stop words. Without stop_words="english", "the" and "and" dominate every topic, because they appear in every document. Every topic ends up looking identical.
Expecting clean topics from real-world data on the first try. This toy example works this well because the three groups barely share any vocabulary. Real corpora — support tickets, product reviews — overlap far more, and topics come out fuzzier. That is normal, not a bug in your code.
Treating the topic count as ground truth. n_components=3 here happens to match how the data was written. On a real corpus you almost never know the "true" number, because there usually isn't one — themes blend into each other.
Try it yourself
Add a thirteenth document that mixes two themes on purpose, such as "the winning candidate celebrated with a cricket match against his campaign team", and re-run. Watch its topic weights — it should no longer land cleanly on one topic.
What to learn next
- Latent Dirichlet Allocation — the classic probabilistic topic model, in depth.
- NMF for topics — the matrix-factorization alternative, run on this exact corpus.
- Clustering — the wider family of algorithms topic modelling belongs to.
Researcher — Mathematics and papers.
The formal problem
Given a corpus of n documents over a vocabulary of V words, topic modelling seeks two things simultaneously:
- A set of
ktopics, each a distribution or weighting over theVwords. - For each document, a mixture describing how much of each topic it draws on.
This is a form of dimensionality reduction and soft clustering performed jointly: instead of assigning each document one hard cluster label, every document gets a k-dimensional weight vector, and every topic gets a V-dimensional word vector. The two families in this section differ mainly in what mathematical object they factor and what constraints they place on it.
LDA (next lesson) treats the whole thing as a generative probabilistic process over a Dirichlet-distributed mixture, fit by variational inference or Gibbs sampling.
NMF (two lessons ahead) treats the document-term matrix X (shape n x V) as a product of two non-negative matrices, X ~ W H, with W shape n x k and H shape k x V, fit by minimising reconstruction error directly.
Both approaches predate large language models by a decade or more and remain in production because they are cheap, fully interpretable — every topic is a literal list of words — and require no GPU.
Where this sits relative to modern methods
Since roughly 2019, embedding-based pipelines (BERTopic, covered later in this section) have added a third family: embed each document with a transformer, cluster the embeddings, then describe each cluster with its most distinctive words. This captures meaning beyond exact word overlap — "physician" and "doctor" pull documents into the same cluster even though they share no letters — at the cost of needing a neural network and losing the clean probabilistic guarantees LDA offers.
Key references
- Blei, D., Ng, A. & Jordan, M. (2003). Latent Dirichlet Allocation. Journal of Machine Learning Research 3.
- Lee, D. & Seung, H. (1999). Learning the parts of objects by non-negative matrix factorization. Nature 401.
- Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794
What to learn next
- Latent Dirichlet Allocation — the classic probabilistic topic model, in depth.
- NMF for topics — the matrix-factorization alternative, run on this exact corpus.
- Clustering — the wider family of algorithms topic modelling belongs to.