Topic Modelling and Text Clustering

Choosing how many topics

Picking the topic count is a judgment call, not a lookup table — held-out likelihood, reconstruction error and a human reading the topics often disagree with each other.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

There is no formula that tells you the right number of topics. You try a few counts and judge the results.

Think about tidying a cluttered cupboard. Sort everything into two boxes and each box is a jumble of unrelated things. Sort everything into fifty boxes instead, and each box holds one lonely item. That is not sorting at all. It is only spreading the mess out thinner. Somewhere in between sits a number of boxes that actually helps you find things later. You find that number by trying a few and looking inside.

Choosing the number of topics for LDA or NMF is the same problem. Too few topics and unrelated documents get crammed together. Too many and the model starts inventing hair-splitting distinctions nobody cares about.

Why it exists

Every topic model needs the topic count fixed before it starts. It is not something the algorithm figures out for you along the way. Someone has to supply a number. That number changes the entire result.

The honest answer is that there is usually no single "true" number of topics hiding in real-world text. A pile of product reviews does not arrive pre-labelled with "this corpus secretly has 7 topics." Choosing a count is closer to choosing a zoom level on a map. It is nothing like measuring a fixed physical quantity.

How it works

try k = 3            try k = 6            try k = 12

"sports"              "cricket"             "cricket - batting"
"food"          vs    "football"     vs     "cricket - bowling"
"politics"            "food"                "football"
                       "politics - local"    "food"
                       "politics - national"  ... (getting fragmented)

At each k, look at:
  - do the topics still make sense to a person reading them?
  - does a held-out slice of documents fit this model well?
  - do topics stop changing much as k grows further?

Nobody runs this once and reads off an answer. You run it at several values of k, using both a metric and your own eyes. You stop where adding more topics stops teaching you anything new.

A real example you have seen

A news app that lets you follow "Cricket" or "Elections" chose that granularity on purpose. It is fine enough to be useful. It is coarse enough that each topic still has a steady stream of stories. Had they chosen "Cricket - domestic T20 - bowling analysis" instead, most days would have no story at all.

Remember this

  • There is no correct topic count waiting to be discovered — only a count that is useful for what you are trying to do.
  • Metrics like held-out likelihood can disagree with what a human judges as sensible topics.
  • Try several values of k, read the topics at each, and watch how much benefit each extra topic actually buys you.

What to learn next

Developer — Code and libraries.

Two different signals are shown here on the same 24-document cricket / curry / election corpus from earlier lessons: held-out perplexity for LDA, and reconstruction error for NMF. Neither hands you a clean answer — that is the point, and it is discussed honestly below.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

NMF's reconstruction error across several values of k

choosing_k.py
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import NMF

cricket = [
    "the batsman hit a six to win the match in the final over",
    "the bowler took three wickets as the batsman walked back after the match",
    "the team captain praised his batsman after the thrilling match",
    "the umpire signalled a boundary as the batsman completed the run",
    "the batsman and bowler shook hands after a hard fought match",
    "a six from the batsman sealed the match in the final over",
    "the bowler ran in fast and beat the batsman for a wicket",
    "the crowd cheered as the batsman brought up his century in the match",
]
curry = [
    "add chopped onion to hot oil then simmer the curry with salt",
    "the recipe needs turmeric, salt and slow simmering of the curry",
    "fry the spices in oil before you add the curry paste for dinner",
    "the chef seasoned the curry with salt and a spoon of turmeric",
    "simmer the curry slowly so the turmeric and salt blend into the oil",
    "the curry recipe calls for onion, turmeric and a pinch of salt",
    "heat the oil, add turmeric, then simmer the curry until thick",
    "the chef added salt and turmeric to the simmering curry pot",
]
election = [
    "the candidate promised new roads and hospitals before the election",
    "voters lined up outside the polling booth to vote in the election",
    "the election result was announced after votes were counted all night",
    "the losing candidate conceded the election after the final count",
    "the candidate campaigned for votes across the election constituency",
    "polling booths across the city saw voters queue for the election",
    "the winning candidate thanked voters after the election result",
    "election officials counted votes at the polling booth past midnight",
]
docs = cricket + curry + election

vectorizer = TfidfVectorizer(stop_words="english")
X = vectorizer.fit_transform(docs)

errors = []
for k in [2, 3, 4, 5, 6, 7, 8]:
    nmf = NMF(n_components=k, random_state=0, max_iter=500)
    nmf.fit(X)
    errors.append(nmf.reconstruction_err_)

prev = None
for k, err in zip([2, 3, 4, 5, 6, 7, 8], errors):
    drop = "" if prev is None else f"   drop from previous k: {prev - err:.2f}"
    print(f"k={k}  reconstruction error={err:.2f}{drop}")
    prev = err
Output
k=2  reconstruction error=4.35
k=3  reconstruction error=4.11   drop from previous k: 0.25
k=4  reconstruction error=3.94   drop from previous k: 0.17
k=5  reconstruction error=3.79   drop from previous k: 0.15
k=6  reconstruction error=3.64   drop from previous k: 0.15
k=7  reconstruction error=3.50   drop from previous k: 0.14
k=8  reconstruction error=3.36   drop from previous k: 0.14

The part worth reading twice

Reconstruction error goes down at every single value of k, without exception. Mathematically, it has no choice — more topics can only ever fit the data at least as well, never worse. There is no "bottom of the curve" to find here, because there isn't one.

Look at the drop column instead. It falls sharply from k=2 to k=3 (0.25), then settles into a much gentler, fairly steady decline (0.14–0.17) from k=3 onward. That flattening — not a dramatic dip, only the rate of improvement slowing down — is the "elbow" people mean when they say "look for the elbow." It is subtle here, and it is often this subtle on real data too. The true number of topics in this corpus is 3, and the elbow, while real, is not dramatic. Do not expect a sharp knee in the curve; expect exactly this kind of gentle levelling-off, and treat it as one signal among several, not a verdict.

The walkthrough

Why NMF's error can never increase with more topics. Adding a topic gives the model one more building block to reconstruct the data with. In the worst case it ignores the extra block; it can never be forced to do worse. This is exactly why raw reconstruction error, alone, cannot select k — it will always vote for "more."

Perplexity has the identical problem for LDA, from the opposite direction. More topics generally lets LDA fit the training data better too. This is why any principled comparison must be done on a held-out split — data the model never trained on — using LatentDirichletAllocation.perplexity() on a test_counts matrix the model has not seen. Even then, as the researcher block covers, held-out likelihood and human-judged topic quality can point in different directions.

Reading the actual topics still matters more than either number. After picking a few candidate values of k using the drop-off pattern, print the top words for each topic at each k — exactly as in the LDA and NMF lessons — and judge for yourself whether the extra topics at higher k are genuinely distinct or only splitting one real theme in two.

Common mistakes

Picking the k with the lowest error or perplexity, full stop. As shown above, that number keeps moving as k grows, on both metrics, for structural reasons that have nothing to do with topic quality.

Only ever trying one value of k. The entire technique here is comparative — a single reconstruction error number, on its own, tells you nothing. It is only informative next to the number one step below and one step above it.

Ignoring how the choice depends on what you'll do with the topics. A dashboard meant for a five-minute daily glance wants fewer, broader topics than a research project cataloguing every fine distinction in a corpus. The "right" k is partly a product decision, not only a statistical one.

Try it yourself

Re-run the loop up to k=15 and watch whether the drop ever levels off completely to near zero, or keeps a small steady trickle going. On this tiny 24-document corpus, past k=8 you are well beyond the three real themes — the extra "topics" are mostly overfitting to individual documents.

What to learn next

Researcher — Mathematics and papers.

Why perplexity and human judgment can diverge

Held-out perplexity for LDA is:

text
perplexity(D_test) = exp( -sum over d in D_test of log p(w_d | Theta)
                            -----------------------------------------
                                  sum over d in D_test of N_d       )

Where p(w_d | Theta) is the model's likelihood of document d's observed words under the fitted parameters Theta, and N_d is its length in tokens. Lower perplexity means the model assigns higher probability to unseen text — a legitimate measure of statistical fit.

Chang, Boyd-Graber, Wang, Gerrig & Blei (2009), Reading Tea Leaves: How Humans Interpret Topic Models, ran a word-intrusion task: show people a topic's top words plus one word from a different, low-probability topic slipped in, and see if they can spot the intruder. Models tuned to maximise held-out likelihood sometimes produced topics where the intruder was harder to spot — meaning the topics were less semantically coherent to a human, even while fitting the data distribution better. This finding is the reason the field largely moved toward coherence metrics for topic-count selection.

Coherence metrics

UMass coherence (Mimno, Wallach, Talley, Leenders & McCallum, 2011) scores a topic's top-M words using only co-occurrence counts within the training corpus itself:

text
C_UMass = sum over i=2..M of sum over j=1..i-1 of log( (D(w_i, w_j) + epsilon) / D(w_j) )

Where D(w_j) is the document frequency of word w_j, D(w_i, w_j) is the co-document frequency of the pair, and epsilon is a small smoothing constant.

C_v coherence (Röder, Both & Hinneburg, 2015), the most widely used in practice via gensim's CoherenceModel, instead builds word vectors from a sliding-window co-occurrence matrix over a reference corpus, computes pointwise mutual information between top words, and aggregates via cosine similarity — a more expensive but empirically better-correlated-with-human-judgment measure, validated against six human-annotated datasets in the original paper.

A practical workflow: fit models across a range of k, compute C_v for each, and pick the k at or near the coherence peak — while still reading the actual topics, since coherence metrics carry their own biases (they tend to reward topics with tightly co-occurring but sometimes redundant words).

Stability as a third signal

Beyond fit-to-data and coherence, topic stability across random restarts is a practical third check: refit the same k with several different seeds and measure how much the resulting topics agree (for example, via word-overlap or Hungarian-matched cosine similarity between topic-word vectors). A k that produces wildly different topics from run to run is likely asking the model to draw boundaries the data does not actually support.

Key references

  • Chang, J., Boyd-Graber, J., Wang, C., Gerrig, R. & Blei, D. (2009). Reading Tea Leaves: How Humans Interpret Topic Models. NeurIPS.
  • Mimno, D., Wallach, H., Talley, E., Leenders, M. & McCallum, A. (2011). Optimizing Semantic Coherence in Topic Models. EMNLP.
  • Röder, M., Both, A. & Hinneburg, A. (2015). Exploring the Space of Topic Coherence Measures. WSDM.

Current state

No single automatic metric is treated as authoritative in current practice. Production pipelines commonly compute C_v coherence across a k sweep as a first filter, then have a human read the resulting topics at the two or three strongest candidates before committing — treating the metric as a shortlist generator, not a final judge.

What to learn next