Topic Modelling and Text Clustering

Latent Dirichlet Allocation

LDA imagines every document as a random mix of topics and every topic as a random mix of words, then works backward from the finished text to guess both mixes.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

LDA guesses a document's mix of topics by imagining it was cooked up from several topics, blended together.

Picture a thali — a plate with small portions of rice, dal, a vegetable, and a sweet. Nobody made you one dish. Someone spooned a bit from several separate pots onto one plate, in different amounts.

LDA — short for Latent Dirichlet Allocation — imagines every document was assembled the same way. A news article is not "about" one single thing. It is a bit of politics, a bit of economics, maybe a pinch of sport. Each piece is spooned in from a separate "pot" of vocabulary. LDA's job is to work backward. Given the finished plate, it figures out which pots it was spooned from, and in what amounts.

Why it exists

Before LDA, the closest tool was pure clustering: force every document into exactly one bucket. That breaks the moment a document genuinely covers more than one thing, which is most of the time. A film review might spend two paragraphs on the plot and one on ticket prices. It is not purely a "plot" document, and not purely a "pricing" document either.

Blei, Ng and Jordan published LDA in 2003 to fix this with a proper mixture. Every document gets its own blend of topics. Every topic is itself a blend of words. Both are learned jointly, from the whole corpus at once.

How it works

   Topic A (cricket-ish)        Topic B (cooking-ish)       Topic C (election-ish)
   batsman, wicket, match  ...  curry, salt, turmeric  ...  candidate, vote, booth  ...

                    LDA assumes each document was written by:
                    1. picking a blend of topics for that document
                       (say: 80% cricket, 15% cooking, 5% election)
                    2. for every word position, picking a topic from that blend,
                       then picking a word from that topic's word list

   Given ONLY the finished document, LDA runs this backward and reports
   its best guess at step 1 and step 2 for every document in the corpus.

LDA never sees "topic A is about cricket." It only ever sees which words keep appearing next to which other words, across thousands of documents. It reverse-engineers a story that explains that pattern well.

A real example you have seen

Academic search engines often show "related research areas" next to a paper. Many are running LDA-style models over millions of abstracts. A single machine learning paper often gets tagged as a blend of "computer vision," "optimisation" and "healthcare." That is the mixed-membership idea LDA was built for, instead of one rigid category.

Remember this

  • LDA treats a document as a mix of topics, and a topic as a mix of words — never a single hard label.
  • It works backward from finished text to guess both mixes at once.
  • Nobody tells it what the topics mean. It only finds words that repeat together; naming the topic is left to you (or to the next lesson on automatic labelling).

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

Twenty-four documents, three real topics, and an honest result

This uses a bigger version of the cricket / curry / election corpus from the previous lesson — eight documents per topic instead of four, because LDA is more data-hungry than the NMF example you saw there.

lda_demo.py
from collections import Counter
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.decomposition import LatentDirichletAllocation

cricket = [
    "the batsman hit a six to win the match in the final over",
    "the bowler took three wickets as the batsman walked back after the match",
    "the team captain praised his batsman after the thrilling match",
    "the umpire signalled a boundary as the batsman completed the run",
    "the batsman and bowler shook hands after a hard fought match",
    "a six from the batsman sealed the match in the final over",
    "the bowler ran in fast and beat the batsman for a wicket",
    "the crowd cheered as the batsman brought up his century in the match",
]
curry = [
    "add chopped onion to hot oil then simmer the curry with salt",
    "the recipe needs turmeric, salt and slow simmering of the curry",
    "fry the spices in oil before you add the curry paste for dinner",
    "the chef seasoned the curry with salt and a spoon of turmeric",
    "simmer the curry slowly so the turmeric and salt blend into the oil",
    "the curry recipe calls for onion, turmeric and a pinch of salt",
    "heat the oil, add turmeric, then simmer the curry until thick",
    "the chef added salt and turmeric to the simmering curry pot",
]
election = [
    "the candidate promised new roads and hospitals before the election",
    "voters lined up outside the polling booth to vote in the election",
    "the election result was announced after votes were counted all night",
    "the losing candidate conceded the election after the final count",
    "the candidate campaigned for votes across the election constituency",
    "polling booths across the city saw voters queue for the election",
    "the winning candidate thanked voters after the election result",
    "election officials counted votes at the polling booth past midnight",
]
docs = cricket + curry + election
true_label = ["cricket"] * 8 + ["curry"] * 8 + ["election"] * 8

vectorizer = CountVectorizer(stop_words="english")  # LDA works on raw counts, not TF-IDF
counts = vectorizer.fit_transform(docs)

lda = LatentDirichletAllocation(n_components=3, random_state=0, max_iter=200, learning_method="batch")
doc_topics = lda.fit_transform(counts)

words = vectorizer.get_feature_names_out()
for i, topic in enumerate(lda.components_):
    top = [words[j] for j in topic.argsort()[-6:][::-1]]
    print(f"topic {i}: {', '.join(top)}")

winners = doc_topics.argmax(axis=1)
topic_label = {}
for t in range(3):
    labels_here = [true_label[i] for i in range(len(docs)) if winners[i] == t]
    topic_label[t] = Counter(labels_here).most_common(1)[0][0]

print()
right = 0
for doc, label, row, t in zip(docs, true_label, doc_topics, winners):
    predicted = topic_label[t]
    ok = predicted == label
    right += ok
    if not ok:
        print(f"MISS  true={label:9s} got={predicted:9s} {row.round(2)}  {doc[:42]}...")

print(f"\n{right}/{len(docs)} documents landed in the topic matching their true theme")
Output
topic 0: election, candidate, polling, voters, votes, booth
topic 1: curry, oil, salt, turmeric, add, simmer
topic 2: batsman, match, bowler, turmeric, salt, curry

MISS  true=cricket   got=election  [0.61 0.05 0.34]  the team captain praised his batsman after...
MISS  true=curry     got=cricket   [0.04 0.05 0.91]  the recipe needs turmeric, salt and slow s...
MISS  true=curry     got=cricket   [0.04 0.05 0.9 ]  the curry recipe calls for onion, turmeric...
MISS  true=curry     got=cricket   [0.04 0.05 0.9 ]  the chef added salt and turmeric to the si...
MISS  true=election  got=cricket   [0.36 0.05 0.59]  the candidate promised new roads and hospi...
MISS  true=election  got=curry     [0.34 0.61 0.05]  the losing candidate conceded the election...

18/24 documents landed in the topic matching their true theme

The part worth reading twice

18 out of 24 correct is not a bug. It is LDA behaving exactly as designed, on a corpus far smaller than it wants. Look at topic 2's word list: batsman, match, bowler, turmeric, salt, curry. It leaked two curry words into the cricket topic, because with only 24 tiny documents and a shared vocabulary of function words, the model has very little evidence to separate them cleanly.

This is a genuinely common experience the first time someone runs LDA. Read the miss list again: every wrong document is a real sentence a human would classify instantly, yet the model hesitated. That gap between "obvious to a person" and "hard to a word-counting algorithm" is the single most useful thing to understand about LDA before using it on real data — it needs far more than 24 documents to find its footing, and even then, some ambiguity is normal.

The walkthrough

CountVectorizer, not TfidfVectorizer. LDA's generative story is built on raw word counts (it models drawing whole words), not weighted scores, so it always takes a count matrix, unlike NMF in the next lesson which prefers TF-IDF.

learning_method="batch" tells sklearn to use the full corpus for every update instead of small random chunks. For a corpus this size, batch is both faster and more stable.

The confused rows share vocabulary with the wrong topic. "the recipe needs turmeric, salt and slow simmering of the curry" landed in the cricket-leaning topic because that topic's word list had quietly absorbed turmeric, salt and curry. Fixing this is exactly what the choosing how many topics and NMF lessons address from two different angles.

Common mistakes

Feeding it TF-IDF scores. LDA's maths assumes integer word counts. TF-IDF weights are fractional and break the model's probabilistic assumptions, even though sklearn will not raise an error — it will still run, and quietly under-perform.

Judging LDA on a tiny corpus and giving up. As shown above, LDA needs volume to separate topics cleanly. Production LDA runs on thousands to millions of documents, not dozens.

Not fixing random_state. LDA's inference is stochastic. Two runs on the same data with different seeds can produce topics in a different order, or occasionally group words slightly differently. Pin the seed for anything you plan to compare across runs.

Try it yourself

Change random_state=0 to a few other integers and re-run. The exact miss list will shift, but the overall accuracy stays in a similar range — a useful way to feel how much of the "wrongness" here is genuine ambiguity in the tiny corpus rather than a one-off unlucky seed.

What to learn next

Researcher — Mathematics and papers.

The generative model

For a corpus of D documents, vocabulary size V, and K topics, LDA assumes each document d was generated as:

text
theta_d ~ Dirichlet(alpha)              # document d's topic mixture
phi_k   ~ Dirichlet(beta)   for k = 1..K  # topic k's word distribution

z_{d,n} ~ Categorical(theta_d)          # topic assigned to word position n in doc d
w_{d,n} ~ Categorical(phi_{z_{d,n}})    # the observed word at that position

Where:

  • theta_d — document d's topic-mixture, a K-length probability vector.
  • phi_k — topic k's word-distribution, a V-length probability vector.
  • alpha, beta — Dirichlet concentration hyperparameters (sklearn's doc_topic_prior, topic_word_prior).
  • z_{d,n} — the topic assigned to the n-th word position in document d.
  • w_{d,n} — the observed word at that position.

Small alpha pulls each document's mixture toward a few dominant topics (sparse membership). Small beta pulls each topic's word list toward a few dominant words (peaky, distinctive topics). Both defaults in most libraries hover near 1/K.

Inference

The generative story is easy to write down and intractable to invert exactly — the posterior over (theta, phi, z) given only the observed words has no closed form. Two families of approximation dominate:

Collapsed Gibbs sampling (Griffiths & Steyvers, 2004) integrates out theta and phi analytically and repeatedly resamples each z_{d,n} conditioned on every other assignment, converging to samples from the true posterior. Accurate, but sequential and slow on large corpora.

Variational inference (Blei, Ng & Jordan, 2003, original batch version; Hoffman, Blei & Bach, 2010, the online/stochastic version scikit-learn implements) replaces the true posterior with a simpler factorised approximation and optimises a lower bound (ELBO) on the log-likelihood via coordinate ascent. Faster and parallelisable, at the cost of a systematic approximation bias.

Complexity

Per iteration, cost scales as O(D * n_avg * K), where n_avg is the average document length in tokens. Online variational LDA processes the corpus in mini-batches, giving roughly constant memory regardless of corpus size — the reason it remains usable on document collections that will not fit in RAM as a dense matrix.

Choosing K and the perplexity trap

Held-out perplexity — exponentiated negative log-likelihood on unseen documents — is the textbook model-selection metric, and LatentDirichletAllocation.perplexity() computes it directly. Chang, Boyd-Graber, Wang, Gerrig & Blei (2009), Reading Tea Leaves: How Humans Interpret Topic Models, found that models with better held-out likelihood were sometimes judged by human raters to have less interpretable topics. Likelihood and human-perceived topic quality are correlated but not the same thing, which is why coherence-based metrics (covered in the next-but-one lesson) largely replaced perplexity for practical topic-count selection.

Key references

  • Blei, D., Ng, A. & Jordan, M. (2003). Latent Dirichlet Allocation. Journal of Machine Learning Research 3.
  • Griffiths, T. & Steyvers, M. (2004). Finding scientific topics. PNAS 101(suppl 1).
  • Hoffman, M., Blei, D. & Bach, F. (2010). Online Learning for Latent Dirichlet Allocation. NeurIPS.
  • Chang, J., Boyd-Graber, J., Wang, C., Gerrig, R. & Blei, D. (2009). Reading Tea Leaves: How Humans Interpret Topic Models. NeurIPS.

Current state and variants

LDA remains widely deployed for its speed and full interpretability — every topic is a literal, inspectable list of word probabilities, with no neural network to explain. Correlated Topic Models (Blei & Lafferty, 2006) relax the Dirichlet's assumption that topics are independent, allowing "sports" and "health" to co-occur more than chance predicts. Structural Topic Models (Roberts, Stewart & Tingley) let document metadata (date, author, source) influence the topic mixture directly. The embedding-based approaches later in this section (BERTopic) largely supersede LDA on short, informal text, where LDA's count-based assumptions struggle most.

What to learn next

What to learn next

These follow on from what you just read.

  • Topic Modelling and Text Clustering

    NMF for topics

    NMF finds topics by breaking a document-word table into two smaller tables that are never allowed to go negative, avoiding a lot of the guesswork LDA needs.

  • Topic Modelling and Text Clustering

    BERTopic

    BERTopic finds topics by embedding documents with a language model first and clustering the embeddings, so it groups by meaning instead of only shared exact words.

  • Topic Modelling and Text Clustering

    Choosing how many topics

    Picking the topic count is a judgment call, not a lookup table — held-out likelihood, reconstruction error and a human reading the topics often disagree with each other.