Topic Modelling and Text Clustering

Clustering documents with embeddings

Turn each document into an embedding, then hand those vectors to an ordinary clustering algorithm — a pattern useful for organising documents even when you never wanted "topics" in the first place.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Document clustering groups similar documents together using their meaning, without any topic labels attached.

Imagine two people organising a bookshelf. One sorts strictly by title, alphabetically. It is fast and mechanical. It puts a cookbook called "Angela's Kitchen" right next to a novel called "Angela's Ashes," for no good reason. The other actually skims each book's back cover. She groups by what it is really about — cooking together, memoirs together, thrillers together.

That second approach is document clustering with embeddings. Every document becomes a numeric embedding that captures what it means, not only what words it contains. An ordinary clustering algorithm then groups the embeddings that sit close together.

Why it exists

Earlier lessons in this section used clustering to find "topics" specifically. But grouping similar documents is useful even when you never wanted a topic model at all. Deduplicating near-identical support tickets. Organising a folder of research papers. Spotting that fifty complaints this week are really the same issue, phrased fifty different ways.

Document clustering is the general-purpose tool. Topic modelling, in the earlier lessons, is one particular way of using it.

How it works

documents  ->  embed each one (a list of numbers capturing meaning)
                        |
                        v
           hand the embeddings to a clustering algorithm
           (K-means if you know roughly how many groups you expect,
            HDBSCAN if you don't, and want honest "doesn't fit anywhere" labels too)
                        |
                        v
           inspect each group -- what do these documents actually have in common?

The clustering step itself does not care that the input happens to be text. It is the same clustering machinery used on any numeric data. Embeddings are what make it work for documents.

A real example you have seen

Some support inboxes automatically group today's incoming tickets into "likely the same issue" clusters. An agent can then answer twenty similar tickets with one reply instead of twenty. That inbox is running exactly this pipeline. Embed each ticket, cluster the embeddings, and hand the clusters to a human.

Remember this

  • Document clustering with embeddings groups by meaning, not by shared exact words.
  • It is useful well beyond "finding topics" — deduplication, organisation and issue-spotting all use the same pattern.
  • A good clustering algorithm should be allowed to say "this document doesn't fit any group" rather than forced to place it somewhere wrong.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install sentence-transformers scikit-learn

all-MiniLM-L6-v2 downloads once, about 90 MB. Outputs verified with sentence-transformers 5.4.1 and scikit-learn 1.7.2 on CPU.

Twenty-six documents, three real themes, and two that belong nowhere

document_clustering.py
from sentence_transformers import SentenceTransformer
from sklearn.cluster import HDBSCAN

cricket = [
    "the batsman hit a six to win the match in the final over",
    "the bowler took three wickets as the batsman walked back after the match",
    "the team captain praised his batsman after the thrilling match",
    "the umpire signalled a boundary as the batsman completed the run",
    "the batsman and bowler shook hands after a hard fought match",
    "a six from the batsman sealed the match in the final over",
    "the bowler ran in fast and beat the batsman for a wicket",
    "the crowd cheered as the batsman brought up his century in the match",
]
curry = [
    "add chopped onion to hot oil then simmer the curry with salt",
    "the recipe needs turmeric, salt and slow simmering of the curry",
    "fry the spices in oil before you add the curry paste for dinner",
    "the chef seasoned the curry with salt and a spoon of turmeric",
    "simmer the curry slowly so the turmeric and salt blend into the oil",
    "the curry recipe calls for onion, turmeric and a pinch of salt",
    "heat the oil, add turmeric, then simmer the curry until thick",
    "the chef added salt and turmeric to the simmering curry pot",
]
election = [
    "the candidate promised new roads and hospitals before the election",
    "voters lined up outside the polling booth to vote in the election",
    "the election result was announced after votes were counted all night",
    "the losing candidate conceded the election after the final count",
    "the candidate campaigned for votes across the election constituency",
    "polling booths across the city saw voters queue for the election",
    "the winning candidate thanked voters after the election result",
    "election officials counted votes at the polling booth past midnight",
]
odd_ones_out = [
    "the new phone has a great camera and long battery life",
    "I can't decide what to wear to the party tonight",
]
docs = cricket + curry + election + odd_ones_out

model = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = model.encode(docs, normalize_embeddings=True)

# clustering directly on the embeddings -- no dimensionality reduction step
labels = HDBSCAN(min_cluster_size=3, metric="euclidean").fit_predict(embeddings)

for c in sorted(set(labels)):
    members = [docs[i] for i in range(len(docs)) if labels[i] == c]
    tag = "noise" if c == -1 else f"cluster {c}"
    print(f"{tag} ({len(members)} docs)")
    for m in members[:2]:
        print(f"   - {m[:55]}...")
Output
noise (2 docs)
   - the new phone has a great camera and long battery life...
   - I can't decide what to wear to the party tonight...
cluster 0 (8 docs)
   - add chopped onion to hot oil then simmer the curry with...
   - the recipe needs turmeric, salt and slow simmering of t...
cluster 1 (8 docs)
   - the batsman hit a six to win the match in the final ove...
   - the bowler took three wickets as the batsman walked bac...
cluster 2 (8 docs)
   - the candidate promised new roads and hospitals before t...
   - voters lined up outside the polling booth to vote in th...

sorted(set(labels)) lists -1 (noise) before any real cluster id, since -1 sorts lowest — that is why noise prints first here. Which numeric id (0, 1, 2) lands on which topic is an artifact of HDBSCAN's internal processing order, not something to rely on; only the grouping itself is the meaningful result.

All three real themes recovered cleanly, and the two deliberately unrelated sentences — about a phone and a party outfit — landed in -1, correctly flagged as noise rather than crammed into whichever cluster happened to be nearest.

The walkthrough

No UMAP step here, on purpose. BERTopic's default pipeline reduces dimensions with UMAP before clustering, and that is the right call on real corpora with thousands of documents. On a corpus this tiny, adding UMAP first actually pulled the two unrelated sentences into the election cluster instead of flagging them as noise, because UMAP's neighbourhood graph has too little data to build a reliable low-dimensional map from only 26 points. Clustering the raw 384-dimensional embeddings directly worked better here — worth trying both on your own data rather than assuming the "standard" pipeline is always the right size.

min_cluster_size=3 sets the smallest group HDBSCAN will call a real cluster, as covered in the HDBSCAN lesson. With only two odd-ones-out, no value at or above 3 lets them form their own cluster — they can only be absorbed into a real cluster or left as noise, and here HDBSCAN chose noise, correctly.

-1 is information, not failure. Read the noise group before discarding it. It is often where genuinely novel or off-topic documents surface, which can matter more than the tidy clusters do.

Common mistakes

Always reaching for UMAP before clustering out of habit. As shown above, dimensionality reduction is a real tool with a real cost on small data, not a mandatory step. Try clustering the raw embeddings first, especially below a few hundred documents.

Choosing K-means when you don't actually know the cluster count. K-means needs a fixed n_clusters decided in advance and will force every document into one of them, noise included. HDBSCAN's -1 label is what to reach for when "how many groups exist" is itself part of the question — see how many clusters? for that decision in general.

Skipping normalisation before Euclidean-based clustering. normalize_embeddings=True keeps every vector the same length, so Euclidean distance between them behaves like cosine similarity — the metric these sentence embeddings were actually trained to respect.

Try it yourself

Add a UMAP reduction step back in (from umap import UMAP, reduce to 5 dimensions, then cluster the result) and compare where the two odd-ones-out land this time. Try it again with more neighbours or a fixed random_state and see how sensitive the outcome is on a corpus this small.

What to learn next

Researcher — Mathematics and papers.

Why embeddings, not raw term-document matrices, generalise clustering to meaning

A term-document matrix (as used by LDA and NMF) places two documents close together only if they share literal vocabulary. A sentence-embedding model, trained with a contrastive objective over paired or labelled sentence data (Reimers & Gurevych, 2019, Sentence-BERT), instead places documents close together based on a learned notion of semantic equivalence — which generalises across synonyms, paraphrase and, to a bounded extent, cross-lingual meaning (see cross-lingual retrieval for the honest limits of that last point).

Formally, clustering on embeddings is clustering in a learned metric space rather than the raw lexical space, and the quality of the clustering is bounded by how well that learned space actually reflects the notion of "similar" relevant to the task — an embedding model trained on general web text will not automatically know that two legal clauses are "similar" in the way a lawyer means it.

The role of dimensionality reduction before clustering, precisely

Most distance-based and density-based clustering algorithms degrade in high dimensions because of the curse of dimensionality: in high-dimensional space, the ratio between the nearest and farthest neighbour distances tends toward 1, eroding the very distance contrasts clustering depends on (covered directly in The curse of dimensionality). UMAP is the standard mitigation in embedding pipelines because it explicitly optimises to preserve local neighbourhood structure during the dimensionality reduction, unlike PCA, which optimises for global variance and can distort the local neighbourhoods density-based clustering relies on.

The developer block's finding — that skipping UMAP worked better on 26 points — is consistent with UMAP's own stated assumptions: its neighbourhood-graph construction (n_neighbors) needs enough local density to estimate a manifold reliably, which a few dozen points in 384 dimensions cannot reliably supply.

Key references

  • Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv:1908.10084
  • McInnes, L., Healy, J. & Melville, J. (2018). UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv:1802.03426
  • Campello, R., Moulavi, D. & Sander, J. (2013). Density-Based Clustering Based on Hierarchical Density Estimates. PAKDD.

Current state

Embedding-based document clustering is now the default first step in most practical deduplication, organisation and exploratory-analysis pipelines over text, largely displacing purely lexical clustering except where full auditability of "why did these documents cluster together" in terms of literal shared words is itself a requirement — a case where LDA or NMF's transparent word-count basis remains an advantage over an opaque embedding space.

What to learn next