Topic Modelling and Text Clustering
Topic modelling on tweets and short text
Classic topic models lean on which words repeat together inside one document, and a tweet is barely long enough to give them that signal — here is what breaks, and what to use instead.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Short messages don't give word-counting topic models enough to work with, so a different approach is needed.
Try guessing a friend's mood from a single text message: "ugh." versus "ugh, this rain again, and now my umbrella broke on the way to work." The one-word message leaves you guessing wildly. The longer one hands you real clues: rain, a broken umbrella, a bad start to the day. Your guess gets far more confident.
Topic models built on word counts, like LDA and NMF, work the same way. They learn which words repeat together within a document. A tweet, a chat message or a one-line review is the "ugh." — too short to show the model much of a pattern.
Why it exists
Customer support chat logs, tweets, app store reviews and search queries are often five to fifteen words long. That is most of the short text a modern product actually generates. Running a word-counting topic model on a pile of these produces oddly fragmented, unstable topics. The underlying themes are not unclear to a person. Each individual document is too short to carry enough co-occurring words for the model to learn from.
How it works
long document (200 words): plenty of word pairs repeat WITHIN this one document
-> LDA/NMF can learn "these words travel together" from this document alone
short document (8 words): almost every word appears once, no repeats to learn from
-> LDA/NMF has to lean on patterns from OTHER short documents instead,
and with so little text per document, that signal is thin and noisy
fix: don't count words within one tiny document at all --
turn each short text into a MEANING vector (an embedding) first,
then group by how close those meanings sit to each otherThe fix mirrors BERTopic's core idea from two lessons back. It sidesteps sparse word co-occurrence entirely by comparing meanings instead of counting shared words.
A real example you have seen
Social media "trending topics" panels group thousands of short posts into clusters like "match highlights" or "election results." This happens within minutes of an event. Classic word-counting topic models struggle badly at that speed and text length. The pipelines behind these panels lean on embedding-based grouping, for exactly the reason covered in this lesson.
Remember this
- Word-counting topic models need repeated word co-occurrence within a document, and short text barely has any.
- The practical fix is to stop counting words per document and start comparing document meanings (embeddings) instead.
- This is the same shift BERTopic makes, and it helps most exactly where document length is shortest.
What to learn next
- Clustering documents with embeddings — the general-purpose version of the technique used here.
- BERTopic — the packaged pipeline built on this exact embed-then-cluster idea.
- Sentence-transformers — how models like
all-MiniLM-L6-v2are trained to place meaning close together.
Developer — Code and libraries.
Setup
pip install scikit-learn sentence-transformerssentence-transformers downloads the all-MiniLM-L6-v2 model, about 90 MB, on first run. Outputs verified with scikit-learn 1.7.2 and sentence-transformers 5.4.1 on CPU.
Twelve tweet-length messages, three real themes
from collections import Counter
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import NMF
tweets = [
"what a six!! match won in the last ball",
"bowler on fire today, 3 wickets already",
"best innings I've seen this season",
"captain's knock saved the whole match",
"this curry needs more salt honestly",
"simmering turmeric smell filling the kitchen",
"added too much oil to the fry today oops",
"grandma's curry recipe never fails",
"who else voted this morning?",
"queue at the booth was so long",
"results tonight, nervous for my candidate",
"election ads everywhere on my phone",
]
true_label = ["cricket"] * 4 + ["curry"] * 4 + ["election"] * 4
vectorizer = TfidfVectorizer(stop_words="english")
X = vectorizer.fit_transform(tweets)
print("average non-empty words per tweet:", round(X.getnnz(axis=1).mean(), 2))
nmf = NMF(n_components=3, random_state=0, max_iter=500)
doc_topics = nmf.fit_transform(X)
winners = doc_topics.argmax(axis=1)
topic_label = {}
for t in range(3):
labels_here = [true_label[i] for i in range(len(tweets)) if winners[i] == t]
topic_label[t] = Counter(labels_here).most_common(1)[0][0] if labels_here else "empty"
right = sum(topic_label[t] == label for t, label in zip(winners, true_label))
print(f"NMF on raw short text: {right}/{len(tweets)} correct")average non-empty words per tweet: 3.75 NMF on raw short text: 7/12 correct
Under four content words per tweet, on average, after stop words are removed. NMF, which was a clean 24/24 on the longer documents in the NMF lesson, drops to 7 out of 12 here — worse than a coin flip would give you wrong on average, and not learning the themes reliably either.
Now the same twelve tweets, through embeddings and clustering instead:
from collections import Counter
from sentence_transformers import SentenceTransformer
from sklearn.cluster import KMeans
tweets = [
"what a six!! match won in the last ball",
"bowler on fire today, 3 wickets already",
"best innings I've seen this season",
"captain's knock saved the whole match",
"this curry needs more salt honestly",
"simmering turmeric smell filling the kitchen",
"added too much oil to the fry today oops",
"grandma's curry recipe never fails",
"who else voted this morning?",
"queue at the booth was so long",
"results tonight, nervous for my candidate",
"election ads everywhere on my phone",
]
true_label = ["cricket"] * 4 + ["curry"] * 4 + ["election"] * 4
model = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = model.encode(tweets, normalize_embeddings=True)
kmeans = KMeans(n_clusters=3, random_state=0, n_init=10).fit(embeddings)
winners = kmeans.labels_
topic_label = {}
for t in range(3):
labels_here = [true_label[i] for i in range(len(tweets)) if winners[i] == t]
topic_label[t] = Counter(labels_here).most_common(1)[0][0] if labels_here else "empty"
right = sum(topic_label[t] == label for t, label in zip(winners, true_label))
print(f"Embeddings + KMeans on the same short text: {right}/{len(tweets)} correct")Embeddings + KMeans on the same short text: 12/12 correct
Same twelve tweets, same three themes, a perfect result this time. The only thing that changed is comparing meaning instead of counting shared words.
The walkthrough
X.getnnz(axis=1).mean() measures the actual sparsity problem. Fewer than four non-stopword tokens per tweet, on average, is not enough repeated vocabulary for a word-counting model to find reliable co-occurrence patterns across only twelve documents.
The embedding model has already "read" far more text than this corpus contains. all-MiniLM-L6-v2 was trained on hundreds of millions of sentence pairs before you ever ran this script. It already knows "wickets" and "innings" belong to the same general space as "match" and "captain," so it needs almost no co-occurrence evidence from this tiny corpus to place tweets correctly — the model brought its own prior knowledge to the sparsity problem.
normalize_embeddings=True matters for KMeans here. Normalising each embedding to unit length turns Euclidean distance (what KMeans minimises) into a close proxy for cosine similarity, which is what these sentence embeddings were trained to respect.
Common mistakes
Concluding "topic modelling doesn't work on short text" from an NMF or LDA result like the 7/12 above. The algorithm family was the wrong choice for this data length, not topic modelling as a concept.
Pooling too aggressively to fix sparsity. A traditional workaround, sometimes called document pooling, concatenates many short texts that share an author, hashtag or timestamp into one longer pseudo-document before running LDA. It can help, but it also blurs together genuinely different topics if the grouping key is loose — a user who tweets about both cricket and food gets one merged pseudo-document that muddies both themes.
Forgetting that embeddings need a model appropriate for the text's register. all-MiniLM-L6-v2 handles everyday English well. Heavy slang, code-mixed text (see Hinglish and code-mixed text), or a specialised domain vocabulary can still trip up a general-purpose embedding model — check a few examples by hand before trusting the clusters blindly.
Try it yourself
Shrink the tweets further — cut each one down to three or four words — and re-run both approaches. Watch the NMF accuracy fall further while the embedding approach holds up noticeably better.
What to learn next
- Clustering documents with embeddings — the general-purpose version of the technique used here.
- BERTopic — the packaged pipeline built on this exact embed-then-cluster idea.
- Sentence-transformers — how models like
all-MiniLM-L6-v2are trained to place meaning close together.
Researcher — Mathematics and papers.
Why sparsity breaks co-occurrence-based models specifically
LDA and NMF both rely, in different mathematical forms, on within-document word co-occurrence as their primary learning signal. For a document of length N, the number of observed word pairs is N*(N-1)/2 — quadratic in length. A 200-word document offers roughly 20,000 pairs of evidence; an eight-word tweet offers 28. Averaged across a corpus, this collapses the effective sample size the model has to separate topics, independent of how many documents you collect, because each individual document still contributes almost nothing.
Purpose-built short-text models
Biterm Topic Model (Yan, Guo, Lan & Cheng, 2013) addresses this directly by modelling co-occurrence at the corpus level rather than the document level: it aggregates all word pairs ("biterms") that appear together anywhere across the whole corpus, then fits topics over that pooled biterm distribution instead of per-document word counts. This sidesteps document-level sparsity by design, at the cost of discarding document boundaries during training — topic assignment for a specific document is inferred afterward from its biterms' topic probabilities.
GSDMM (Yin & Wang, 2014), the Dirichlet Multinomial Mixture model for short text, makes a stronger simplifying assumption: each document, not each word, is drawn from a single topic (the "one topic per document" assumption is far more defensible for a tweet than for a news article). Inference proceeds by a collapsed Gibbs sampling procedure resembling a Chinese Restaurant Process, where documents can also self-organise into a number of clusters not fixed in advance, given an upper bound.
Why embeddings sidestep the problem entirely
A pretrained sentence embedding model, as used in the developer block, was fit on a training objective (typically contrastive, following Reimers & Gurevych, 2019's Sentence-BERT) over hundreds of millions of sentence pairs before it ever sees the target corpus. The within-corpus sparsity problem specific to LDA and NMF does not apply here, because the model is not trying to learn word co-occurrence patterns from this corpus at all — it is applying patterns learned once, at scale, elsewhere. The trade-off is a dependency on the embedding model's own training distribution matching the target text's register and language, which does not hold uniformly across domains or languages — see the honest coverage gaps discussed in Multilingual sentence embeddings.
Key references
- Yan, X., Guo, J., Lan, Y. & Cheng, X. (2013). A Biterm Topic Model for Short Texts. WWW.
- Yin, J. & Wang, J. (2014). A Dirichlet Multinomial Mixture Model-based Approach for Short Text Clustering. KDD.
- Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv:1908.10084
Current state
Embedding-and-cluster pipelines (BERTopic and its relatives) have largely displaced GSDMM and BTM in production short-text topic modelling, for the same reason they outperformed classic models in the developer block above: they do not depend on within-corpus word co-occurrence to begin with. GSDMM and BTM remain relevant where a fully self-contained, no-external-model solution is required, or where research explicitly wants a probabilistic, generative account of short-text topics rather than a geometric clustering of embeddings.
What to learn next
- Clustering documents with embeddings — the general-purpose version of the technique used here.
- BERTopic — the packaged pipeline built on this exact embed-then-cluster idea.
- Sentence-transformers — how models like
all-MiniLM-L6-v2are trained to place meaning close together.