Topic Modelling and Text Clustering
Tracking topics over time
Fit one topic model across your whole corpus, then split the results by time period afterward, to watch which topics are rising and which are fading without retraining a separate model for every period.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Tracking topics over time watches how much of the conversation each topic takes up, month by month.
Picture flipping through a year of old newspapers, one month at a time. In April, the front pages are full of cricket. By October, cricket has quietly shrunk to the back page. Election coverage has taken over the front instead. Nothing about "news" itself changed. The mix of what people were talking about shifted, gradually. You can watch it happen by comparing month to month.
That is what tracking topics over time does to a document collection. The topics themselves usually stay the same handful of themes. What changes is how much of the conversation each one occupies. That share is what you are actually measuring, as time passes.
Why it exists
A single topic model, run once, gives you a snapshot: "here is what this corpus is about, overall." That hides a genuinely useful signal. Is a topic growing? Shrinking? Staying flat? A support team cares far more about "complaints about the new pricing tier tripled this week." A static list of topics has no sense of direction at all.
How it works
all documents, across every time period, fit ONE shared topic model
(so "topic 2" means the same thing in January as it does in December)
|
v
tag each document with its date, and its topic weights from that shared model
|
v
group documents by time period, and average each topic's weight within each period
|
v
Jan: topic A = 0.50 topic B = 0.33 topic C = 0.17
Feb: topic A = 0.33 topic B = 0.33 topic C = 0.33
Mar: topic A = 0.17 topic B = 0.33 topic C = 0.50
^
topic C is on the riseThe key decision is fitting one topic model on everything, not a separate model per time slice. A separate model per period would relabel and reshuffle topics independently each time. "Topic 2" would mean something different in every period — nothing to actually compare.
A real example you have seen
A news aggregator's "trending this week" graph often shows a topic's share of coverage climbing steadily before an election. It then falls away with the same steadiness afterward. That graph is built exactly this way: one shared topic model, sliced by publication date after the fact.
Remember this
- Fit one topic model across the whole corpus first. Comparing topics over time only makes sense if "topic 2" means the same thing in every period.
- Track each topic's average share of the conversation per time period, not only whether it appears at all.
- A topic rising or falling over time is often more actionable information than a static snapshot of what topics exist.
What to learn next
- Clustering documents with embeddings — the grouping step this lesson builds a time dimension on top of.
- Choosing how many topics — the upstream decision that shapes what "one shared topic model" actually finds.
- Moving averages — smoothing a noisy trend line like the one built in this lesson.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU.
Four weeks, one shared model, a clear trend
The same cricket / curry / election sentences from earlier lessons, this time split across four weeks in the run-up to an election — cricket chatter fading, election chatter growing, curry talk staying flat as a control group.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import NMF
cricket = [
"the batsman hit a six to win the match in the final over",
"the bowler took three wickets as the batsman walked back after the match",
"the team captain praised his batsman after the thrilling match",
"the umpire signalled a boundary as the batsman completed the run",
"the batsman and bowler shook hands after a hard fought match",
"a six from the batsman sealed the match in the final over",
"the bowler ran in fast and beat the batsman for a wicket",
"the crowd cheered as the batsman brought up his century in the match",
]
curry = [
"add chopped onion to hot oil then simmer the curry with salt",
"the recipe needs turmeric, salt and slow simmering of the curry",
"fry the spices in oil before you add the curry paste for dinner",
"the chef seasoned the curry with salt and a spoon of turmeric",
"simmer the curry slowly so the turmeric and salt blend into the oil",
"the curry recipe calls for onion, turmeric and a pinch of salt",
"heat the oil, add turmeric, then simmer the curry until thick",
"the chef added salt and turmeric to the simmering curry pot",
]
election = [
"the candidate promised new roads and hospitals before the election",
"voters lined up outside the polling booth to vote in the election",
"the election result was announced after votes were counted all night",
"the losing candidate conceded the election after the final count",
"the candidate campaigned for votes across the election constituency",
"polling booths across the city saw voters queue for the election",
"the winning candidate thanked voters after the election result",
"election officials counted votes at the polling booth past midnight",
]
# week 1 -> week 4, run-up to election day: election chatter grows, cricket chatter shrinks
weeks = [
cricket[0:3] + curry[0:2] + election[0:1],
cricket[3:5] + curry[2:4] + election[1:3],
cricket[5:7] + curry[4:6] + election[3:5],
cricket[7:8] + curry[6:8] + election[5:8],
]
docs = [d for week in weeks for d in week]
week_id = [i for i, week in enumerate(weeks, start=1) for _ in week]
vectorizer = TfidfVectorizer(stop_words="english")
X = vectorizer.fit_transform(docs)
nmf = NMF(n_components=3, random_state=0, max_iter=500)
doc_topics = nmf.fit_transform(X) # fit ONCE, across every week
words = vectorizer.get_feature_names_out()
topic_names = [words[topic.argsort()[-1]] for topic in nmf.components_]
print("topics found:", topic_names)
doc_topics_norm = doc_topics / doc_topics.sum(axis=1, keepdims=True)
print(f"\n{'week':>6} " + " ".join(f"{n:>10}" for n in topic_names))
for w in [1, 2, 3, 4]:
idx = [i for i, x in enumerate(week_id) if x == w]
avg = doc_topics_norm[idx].mean(axis=0)
print(f"{w:>6} " + " ".join(f"{v:>10.2f}" for v in avg))topics found: ['curry', 'batsman', 'election']
week curry batsman election
1 0.33 0.50 0.17
2 0.33 0.33 0.33
3 0.33 0.37 0.29
4 0.33 0.17 0.50Cricket ("batsman") started at half the conversation and fell to a sixth. Election chatter did the exact opposite. Curry, the control group, sat almost exactly at 0.33 every single week, because its share of the conversation genuinely did not change — only the mix of documents fed in each week did.
The walkthrough
nmf.fit_transform(docs) runs once, over every week combined. This is the whole trick. Because the model never sees "week" as an input, topic 0 refers to the same thing in every week's row — the model was never given the chance to relabel topics between periods.
week_id is metadata bolted on after fitting, not used during fitting. The model has no idea which document came from which week. All the "time" reasoning happens afterward, by grouping already-computed topic weights.
Averaging normalised weights, not counting documents. doc_topics_norm[idx].mean(axis=0) averages each document's share of each topic within a week, which is more informative than only counting how many documents "belong" to each topic — a document that is 60% election and 40% cricket contributes partial credit to both trends, matching how real documents actually blend themes.
Common mistakes
Fitting a separate topic model per time period. This is the single most common mistake here, and it looks reasonable at first. It produces topics that cannot be compared, because "topic 0" in January's model and "topic 0" in February's model were fit independently and have no guaranteed relationship to each other.
Reading noise in small time bins as a real trend. With only a handful of documents in a period, a topic's average share swings around a lot from ordinary sampling variation alone. Wider bins, or a smoothed trend line across bins, are more trustworthy than reading too much into one week's number.
Ignoring that vocabulary itself can drift over a long enough time span. A topic model fit once on data spanning several years may find that the words people use for a stable topic have changed enough that the fixed word list stops matching well by the end of the period. True Dynamic Topic Models, covered in the researcher block, address this directly; the fit-once-and-slice approach here assumes vocabulary stays reasonably stable across the tracked window.
Try it yourself
Add a fifth week where election chatter drops back down (the election is over) and curry chatter spikes (a festival, say). Re-run and confirm the trend table reflects both shifts — this is the same mechanism a real "topics over time" dashboard runs on a schedule.
What to learn next
- Clustering documents with embeddings — the grouping step this lesson builds a time dimension on top of.
- Choosing how many topics — the upstream decision that shapes what "one shared topic model" actually finds.
- Moving averages — smoothing a noisy trend line like the one built in this lesson.
Researcher — Mathematics and papers.
True Dynamic Topic Models
The approach in the developer block — fit once, slice by time afterward — is a practical approximation. The original Dynamic Topic Model (Blei & Lafferty, 2006) instead lets the topics themselves evolve continuously across time slices. Each topic's word distribution at time slice t is linked to its distribution at t-1 through a state-space (Gaussian random walk) model on the natural parameters:
beta(k, t) | beta(k, t-1) ~ Normal( beta(k, t-1), sigma^2 * I )Where beta(k, t) is topic k's (unnormalised, logistic-normal) word-distribution parameters at time slice t, and sigma^2 controls how much drift is allowed between adjacent slices. Inference uses a variational Kalman filter across slices, jointly smoothing topic-word distributions through time rather than holding them fixed. This lets a topic's vocabulary, not only its prevalence, evolve — "election" coverage might shift from words like "campaign" and "rally" early on to "results" and "recount" later, tracked as the same topic drifting smoothly.
Topics-over-time as an alternative formulation
Topics over Time (Wang & McCallum, 2006) takes a different route: it keeps topics fixed (as in ordinary LDA) but adds a per-topic Beta distribution over a continuous timestamp, learned jointly with the topic-word distributions. A document's observed timestamp becomes an additional variable the model explains, rather than external metadata bolted on afterward — this is a genuinely joint model of content and time, at the cost of assuming a single unimodal Beta shape for each topic's prevalence curve, which real-world bursty trends (a sudden news event) do not always follow well.
Trade-offs against the fit-once approach
The post-hoc slicing method from the developer block scales to any topic-fitting algorithm (LDA, NMF, or an embedding-cluster pipeline such as BERTopic, which ships its own topics_over_time() utility using this exact pattern) and requires no specialised inference machinery. Its limitation is the vocabulary-drift assumption noted above. True DTM handles drifting vocabulary correctly but requires custom inference code, scales less gracefully to very large corpora, and is far less commonly available in mainstream libraries than standard LDA or NMF.
Key references
- Blei, D. & Lafferty, J. (2006). Dynamic Topic Models. ICML.
- Wang, X. & McCallum, A. (2006). Topics over Time: A Non-Markov Continuous-Time Model of Topical Trends. KDD.
- Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794
Current state
In production monitoring and trend-dashboard settings, the fit-once, slice-afterward pattern from the developer block dominates in practice, largely because it composes cleanly with any topic-fitting method already in use and needs no specialised state-space inference code. True Dynamic Topic Models remain more common in research settings studying long-running corpora — scientific literature, historical archives — where tracking genuine vocabulary drift over years or decades is itself the object of study, not only a nuisance to work around.
What to learn next
- Clustering documents with embeddings — the grouping step this lesson builds a time dimension on top of.
- Choosing how many topics — the upstream decision that shapes what "one shared topic model" actually finds.
- Moving averages — smoothing a noisy trend line like the one built in this lesson.