Topic Modelling and Text Clustering
BERTopic
BERTopic finds topics by embedding documents with a language model first and clustering the embeddings, so it groups by meaning instead of only shared exact words.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
BERTopic groups documents by meaning first, then describes each group with its most distinctive words.
Imagine two people sorting the same box of old photographs. Your grandmother checks the back of each photo for a matching sticker. Only then does she put two photos in the same pile. Your friend was actually at most of these events. She glances at each photo and instantly knows where it belongs. A wedding photo joins the other wedding photos, even from a strange angle, even with different people in frame.
LDA and NMF sort like your grandmother. They need matching words to notice two documents belong together. BERTopic sorts like your friend. It converts every document into an embedding first — a list of numbers capturing meaning. Then it groups documents whose meanings sit close together, even if they share almost no actual words.
Why it exists
"Doctor" and "physician" mean the same thing and share zero letters. A word-counting topic model treats them as completely unrelated. One review says "the doctor was rude." Another says "the physician ignored me." They might land in different topics purely because the words differ. A human reads them as the same complaint anyway.
BERTopic, introduced by Maarten Grootendorst in 2022, fixes this. It moves the entire problem into embedding space before any grouping happens. Meaning-based similarity replaces word-based similarity.
How it works
documents
|
v
embed each one with a language model ("doctor was rude" and "physician ignored
| me" land close together in this space)
v
shrink the embeddings down to a
handful of dimensions (UMAP)
|
v
cluster the shrunk embeddings (HDBSCAN) -- finds groups of any shape, any count,
| and leaves outliers unclustered
v
for each cluster, find the words that
appear unusually often INSIDE it -- this becomes the topic's readable labelFour stages, each swappable. This is BERTopic's real design. Not one clever trick — a pipeline of well-understood pieces, stitched together. An embedding model. A dimensionality reduction step. A clustering step. A labelling step.
A real example you have seen
Some dashboards show "trending topics in customer feedback this week," with no hand-written keyword list at all. Most are running an embedding-and-cluster pipeline like this one. That is why the dashboard can surface "people are confused about the new pricing tiers." Not one customer needed to use the word "pricing." Some said "cost." Others said "how much it charges." The embeddings pulled them together anyway.
Remember this
- BERTopic embeds documents by meaning first, then clusters the embeddings — synonyms and rewordings land in the same group automatically.
- It is a pipeline of separate, swappable steps: embed, shrink dimensions, cluster, then label.
- Unlike LDA or NMF, it can leave some documents unclustered as noise rather than forcing every document into a topic.
What to learn next
- Clustering documents with embeddings — the same embed-then-cluster idea, used for organising documents rather than naming topics.
- How HNSW and IVF actually work — the search structures that make embedding pipelines fast at real scale.
- UMAP — a deeper look at the dimensionality-reduction step BERTopic relies on.
Developer — Code and libraries.
Setup
pip install bertopic sentence-transformersThis pulls in umap-learn, hdbscan and sentence-transformers as dependencies — expect several hundred MB downloaded and a few minutes on first install. The embedding model, all-MiniLM-L6-v2, is a further ~90 MB, downloaded once and cached. Everything below runs on CPU. Outputs verified with bertopic 0.17.4 and sentence-transformers 5.4.1.
The same 24 cricket / curry / election documents
from bertopic import BERTopic
from sklearn.feature_extraction.text import CountVectorizer
from umap import UMAP
from hdbscan import HDBSCAN
cricket = [
"the batsman hit a six to win the match in the final over",
"the bowler took three wickets as the batsman walked back after the match",
"the team captain praised his batsman after the thrilling match",
"the umpire signalled a boundary as the batsman completed the run",
"the batsman and bowler shook hands after a hard fought match",
"a six from the batsman sealed the match in the final over",
"the bowler ran in fast and beat the batsman for a wicket",
"the crowd cheered as the batsman brought up his century in the match",
]
curry = [
"add chopped onion to hot oil then simmer the curry with salt",
"the recipe needs turmeric, salt and slow simmering of the curry",
"fry the spices in oil before you add the curry paste for dinner",
"the chef seasoned the curry with salt and a spoon of turmeric",
"simmer the curry slowly so the turmeric and salt blend into the oil",
"the curry recipe calls for onion, turmeric and a pinch of salt",
"heat the oil, add turmeric, then simmer the curry until thick",
"the chef added salt and turmeric to the simmering curry pot",
]
election = [
"the candidate promised new roads and hospitals before the election",
"voters lined up outside the polling booth to vote in the election",
"the election result was announced after votes were counted all night",
"the losing candidate conceded the election after the final count",
"the candidate campaigned for votes across the election constituency",
"polling booths across the city saw voters queue for the election",
"the winning candidate thanked voters after the election result",
"election officials counted votes at the polling booth past midnight",
]
docs = cricket + curry + election
# small corpora need smaller neighbourhood/cluster sizes than BERTopic's defaults
umap_model = UMAP(n_neighbors=5, n_components=5, min_dist=0.0, metric="cosine", random_state=42)
hdbscan_model = HDBSCAN(min_cluster_size=5, metric="euclidean", cluster_selection_method="eom", prediction_data=True)
topic_model = BERTopic(
embedding_model="all-MiniLM-L6-v2",
umap_model=umap_model,
hdbscan_model=hdbscan_model,
vectorizer_model=CountVectorizer(stop_words="english"),
calculate_probabilities=False,
)
topics, _ = topic_model.fit_transform(docs)
for t in sorted(set(topics)):
if t == -1:
continue
words = [w for w, _ in topic_model.get_topic(t)[:6]]
print(f"topic {t}: {', '.join(words)}")
print("\nper-document topic ids:", topics)topic 0: batsman, match, bowler, final, wicket, walked topic 1: curry, turmeric, salt, oil, add, simmer topic 2: election, candidate, polling, votes, voters, booth per-document topic ids: [0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2]
A perfect three-way split, with no document left as noise — every one of the 24 documents landed exactly where a human would put it. You may also see harmless log lines from TensorFlow's oneDNN backend printed above this output during import; they are safe to ignore.
The walkthrough
embedding_model="all-MiniLM-L6-v2" is doing the meaning-based lifting. It turns each sentence into a 384-number vector where semantically similar sentences land close together, regardless of exact wording. See embeddings for how that training works.
Why umap_model and hdbscan_model are set explicitly. BERTopic's defaults assume corpora of at least a few hundred documents. On 24 tiny documents, the default neighbourhood and cluster-size settings would refuse to find any structure. n_neighbors=5 and min_cluster_size=5 are scaled down to match this toy size — on real data with thousands of documents, the library's defaults work well unmodified.
Topic word lists come from vectorizer_model, not the embedding model. BERTopic clusters using embeddings, but describes each cluster afterward with class-based TF-IDF over the raw words — the same idea built by hand in the topic-labelling lesson. Embeddings decide the grouping; word counts decide the label.
No -1 topic appeared here. In HDBSCAN's vocabulary, -1 means "noise — didn't fit any cluster confidently." On this clean, well-separated toy corpus nothing gets discarded. On real, messier text, expect a meaningful fraction of documents to land in -1, and treat that as useful information, not failure — see Clustering documents with embeddings for more on reading noise honestly.
Common mistakes
Running BERTopic's defaults on a tiny corpus and concluding it "doesn't work." As shown above, the neighbourhood and cluster-size parameters need to shrink to match small data. This is the single most common first-time BERTopic complaint online, and it is almost always a parameter mismatch, not a broken library.
Not fixing random_state in the UMAP step. UMAP is stochastic. Without a seed, cluster boundaries — and sometimes cluster counts — shift slightly between runs on the same data.
Forgetting the embedding model download on a fresh machine. The first run downloads all-MiniLM-L6-v2; a locked-down or offline CI environment needs that model cached ahead of time.
Try it yourself
Lower hdbscan_model's min_cluster_size to 2 and re-run. With such a small threshold on this well-separated data, you likely still get three clean clusters — but on noisier real text, a min_cluster_size this low tends to fragment one real topic into several spurious ones. Try raising it to 8 as well and see clusters start merging.
What to learn next
- Clustering documents with embeddings — the same embed-then-cluster idea, used for organising documents rather than naming topics.
- How HNSW and IVF actually work — the search structures that make embedding pipelines fast at real scale.
- UMAP — a deeper look at the dimensionality-reduction step BERTopic relies on.
Researcher — Mathematics and papers.
The pipeline, formally
BERTopic (Grootendorst, 2022) is explicitly modular, with four default stages:
- Embed. Any sentence-embedding model maps each document to a vector
vof dimensiond(typicallydbetween 384 and 768). - Reduce. UMAP (McInnes, Healy & Melville, 2018) projects to a lower dimension, typically 5–10, before clustering — high-dimensional density estimation is unreliable, and UMAP's non-linear manifold assumption tends to preserve local cluster structure better than PCA for this purpose.
- Cluster. HDBSCAN finds density-based clusters of arbitrary shape and count, explicitly modelling low-density points as noise rather than forcing membership.
- Represent. Class-based TF-IDF (c-TF-IDF) treats every cluster as one concatenated "document" and scores terms by:
W(t, c) = tf(t, c) * log(1 + A / f(t))Where:
tf(t, c)— frequency of termtwithin the concatenated documents of clusterc.A— the average number of words per cluster, across all clusters.f(t)— the total frequency of termtacross all clusters combined.
This is ordinary TF-IDF with "class" substituted for "document" — a term is a good label for a cluster if it is common inside that cluster and rare in the others.
Why this displaced classic topic models for short, informal text
LDA and NMF's word-count assumptions degrade badly on short documents (tweets, chat messages, reviews under a sentence) because there is not enough co-occurrence signal within any single document — this is examined directly in Topic modelling on tweets and short text. Because BERTopic clusters whole-document embeddings rather than per-document word co-occurrence, document length matters far less: a five-word review and a five-paragraph review can both be embedded and clustered by the same mechanism, provided the embedding model was trained on comparable text.
Complexity and scale
UMAP fitting is roughly O(n log n) via approximate nearest-neighbour graphs (typically pynndescent); HDBSCAN clustering is near O(n log n) with a dual-tree Boruvka MST as covered in the HDBSCAN lesson. The embedding step dominates wall-clock time at scale, since it requires one forward pass per document through a transformer — batchable and GPU-accelerable, but the clear bottleneck past a few hundred thousand documents.
Key references
- Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794
- McInnes, L., Healy, J. & Melville, J. (2018). UMAP: Uniform Manifold Approximation and Projection. arXiv:1802.03426
- Campello, R., Moulavi, D. & Sander, J. (2013). Density-Based Clustering Based on Hierarchical Density Estimates. PAKDD.
Current state
Later BERTopic releases add optional representation-tuning steps on top of c-TF-IDF — MMR for less redundant word lists, KeyBERT-inspired re-ranking, and hooks to have an LLM write a short natural-language topic name from the top words and a sample of documents. The core embed-reduce-cluster-represent pipeline, and its dependence on embedding quality, has stayed stable since the original paper.
What to learn next
- Clustering documents with embeddings — the same embed-then-cluster idea, used for organising documents rather than naming topics.
- How HNSW and IVF actually work — the search structures that make embedding pipelines fast at real scale.
- UMAP — a deeper look at the dimensionality-reduction step BERTopic relies on.