Topic Modelling and Text Clustering

Naming your topics automatically

A topic model hands you a ranked list of words per group. Turning that list into a short, readable name is its own well-defined problem, and it has real solutions.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Auto-labelling picks the handful of words that best summarise a topic, so a human doesn't have to.

Picture packing for a house move. You fill a box with kitchen things and seal it. You scrawl a label on the side. That way, the box can be found later without opening it: "pots, pans, spices." You did not list everything inside. You picked the two or three items that best represent what the box is about.

Every topic model from earlier in this section hands back a topic as a long ranked list of words. That list is accurate but not friendly to read at a glance. Auto-labelling shrinks that list into a short box label instead: "cooking / curry recipes."

Why it exists

A ranked word list works for a data scientist staring at a table. It does not work on a dashboard someone glances at for three seconds. It does not work in a report someone else has to read cold. Someone still has to decide which words earn a place on the label. That decision follows a real pattern, not guesswork.

How it works

topic's word list (from LDA / NMF / BERTopic), ranked by weight:

  curry (0.31)  salt (0.24)  turmeric (0.19)  add (0.09)  oil (0.08)  the (0.02)  ...

The trick: don't only take the highest-weighted words in THIS topic.
Check whether each word ALSO shows up strongly in the OTHER topics.

  "curry"    -> high here, near-zero everywhere else   -> KEEP, distinctive
  "add"      -> high here, but also high in "cooking-adjacent" topics -> weaker signal
  "the"      -> high everywhere                        -> DROP, useless as a label

label = the top surviving distinctive words, joined together:  "curry-salt-turmeric"

This is the same idea as TF-IDF, applied one level up. Instead of scoring a word's importance to one document, it scores importance to one topic, against every other topic.

A real example you have seen

App store summaries often show short tags like "battery life," "crashes often," or "great value" under an app. Nobody typed those in by reading every review. A model clustered the reviews first. Then an auto-labelling step picked the words most distinctive to each cluster. Those words became that short tag.

Remember this

  • Labelling a topic means finding words that are common inside it but rare everywhere else — not only the topic's single most frequent words.
  • This is TF-IDF's idea, applied to whole topics instead of single documents.
  • A good auto-label is a starting point for a human to refine, not always the final answer.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

Building class-based TF-IDF (c-TF-IDF) from scratch

This is the exact mechanism BERTopic uses under the hood — implemented here in plain NumPy so nothing is hidden.

ctfidf_labels.py
import numpy as np
from sklearn.feature_extraction.text import CountVectorizer

cricket = [
    "the batsman hit a six to win the match in the final over",
    "the bowler took three wickets as the batsman walked back after the match",
    "the team captain praised his batsman after the thrilling match",
    "the umpire signalled a boundary as the batsman completed the run",
]
curry = [
    "add chopped onion to hot oil then simmer the curry with salt",
    "the recipe needs turmeric, salt and slow simmering of the curry",
    "fry the spices in oil before you add the curry paste for dinner",
    "the chef seasoned the curry with salt and a spoon of turmeric",
]
election = [
    "the candidate promised new roads and hospitals before the election",
    "voters lined up outside the polling booth to vote in the election",
    "the election result was announced after votes were counted all night",
    "the losing candidate conceded the election after the final count",
]
clusters = {"cluster_0": cricket, "cluster_1": curry, "cluster_2": election}

# collapse each cluster into one "mega-document" -- the whole trick behind c-TF-IDF
vectorizer = CountVectorizer(stop_words="english")
mega_docs = [" ".join(docs) for docs in clusters.values()]
counts = vectorizer.fit_transform(mega_docs).toarray()   # shape: (n_clusters, n_terms)
words = vectorizer.get_feature_names_out()

tf = counts / counts.sum(axis=1, keepdims=True)             # how common a word is INSIDE this cluster
avg_words_per_class = counts.sum(axis=1).mean()
term_freq_across_classes = counts.sum(axis=0)                # how common a word is OVERALL
idf = np.log(1 + avg_words_per_class / term_freq_across_classes)

c_tf_idf = tf * idf

for name, row in zip(clusters.keys(), c_tf_idf):
    top_idx = row.argsort()[-4:][::-1]
    label_words = [words[i] for i in top_idx]
    print(f"{name}: auto-label = '{'-'.join(label_words)}'")
Output
cluster_0: auto-label = 'batsman-match-team-wickets'
cluster_1: auto-label = 'curry-salt-oil-turmeric'
cluster_2: auto-label = 'election-candidate-counted-conceded'

Three tidy labels, built from nothing but word counts and one formula — no human read a single document.

The walkthrough

" ".join(docs) per cluster is the key move. Ordinary TF-IDF compares a word's importance across individual documents. c-TF-IDF first merges every document in a cluster into one giant "document," so the comparison happens across clusters instead. Everything downstream is standard TF-IDF maths applied to these merged texts.

tf measures local importance. A word's share of all words inside this one cluster's merged text — high for words that dominate the cluster's vocabulary.

idf measures how unusual that word is elsewhere. avg_words_per_class / term_freq_across_classes grows large when a word is rare across the other clusters, and the log(1 + ...) compresses that ratio into a sane range while keeping it always positive.

Multiplying tf * idf rewards words that are both common here and rare elsewhere — exactly the property a good label needs. A word that is common everywhere (like a leftover stop word) gets a small idf and drops out of the top-4 automatically, with no manual stop-word list required at this stage.

Common mistakes

Labelling with raw word frequency instead of c-TF-IDF. The single most frequent word in a topic is very often a generic word that happens to be common in that domain overall ("said," "today," "app") rather than the word that actually distinguishes this topic from the others.

Trusting the auto-label without glancing at a few real documents from the cluster. "curry-salt-oil-turmeric" is a fair label here because the underlying documents are genuinely about cooking. On messier real data, a label can look plausible and still miss the point of the cluster — a quick sanity read of two or three source documents catches this cheaply.

Using too few or too many words in the label. Two words is often ambiguous; eight words stops reading like a label and starts reading like the same ranked list you were trying to summarise. Three to five words is a reasonable working default.

Try it yourself

Add a fourth, deliberately overlapping cluster — reuse a few cricket sentences mixed with a few curry sentences — and recompute. Watch the idf term pull previously "distinctive" words like curry down, since it no longer belongs to only one cluster.

What to learn next

Researcher — Mathematics and papers.

The c-TF-IDF formula

Following Grootendorst (2022), for term t and class (cluster) c:

text
W(t, c) = tf(t, c) * log(1 + A / f(t))

Where:

  • tf(t, c) — the raw frequency of term t within the concatenated documents of class c.
  • A — the average number of words per class, across all classes.
  • f(t) — the total frequency of term t summed over every class.

This differs from ordinary TF-IDF's inverse-document-frequency term, log(N / df(t)), in exactly one respect: the denominator counts term occurrences across classes rather than the number of documents a term appears in. That substitution is what makes the formula rank a topic's distinguishing words rather than a single document's.

Beyond ranked word lists

MMR (Maximal Marginal Relevance) re-ranks a topic's top-N c-TF-IDF words to reduce redundancy — without it, a topic can surface near-synonyms ("curry," "curries," "curried") that all say the same thing; MMR trades off relevance against diversity when selecting the final label words.

KeyBERT-style extraction scores candidate label phrases by their embedding-space cosine similarity to the cluster's centroid embedding, rather than by word-count statistics — better at surfacing multi-word phrases ("customer service," not only "customer" and "service" separately) at the cost of needing the embedding model at labelling time, not only at clustering time.

LLM-written labels are the most recent addition to library pipelines (including current BERTopic releases): the top c-TF-IDF words plus a handful of representative documents are placed in a prompt, and a language model is asked to write a short human-readable name. This produces the most natural-sounding labels but adds an external API dependency and cost per topic, and needs the same honesty check as any LLM output — verify a sample against the source documents rather than trusting it by default.

Key references

  • Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794
  • Carbonell, J. & Goldstein, J. (1998). The Use of MMR, Diversity-Based Reranking for Reordering Documents and Producing Summaries. SIGIR.

Current state

c-TF-IDF remains the default representation step across embedding-based topic pipelines because it is fast, needs no extra model beyond the counts already available, and is fully auditable — every word in a label traces back to an exact count in the source text, unlike an LLM-generated label, which requires separate verification against the underlying documents.

What to learn next