Content-based filtering
Content-based filtering recommends items that look like the ones you already liked, using the item's own description instead of other people's behaviour.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Content-based filtering recommends items that resemble the ones you already liked, judged by what the item is made of.
The analogy you have already lived
You pick up a packet of biscuits in a shop and turn it over to read the ingredients. Elaichi, cashew, no palm oil. You liked the last packet with those words on it, so you buy this one too.
You did not ask anybody. You did not check what the rest of the queue was buying. You read the label and matched it against what you already know you enjoy.
That is content-based filtering, exactly. The system reads the label of every item and finds the ones whose labels look like the labels of things you liked.
Why it exists
Collaborative filtering has a hole in it. It knows nothing about an item until a crowd has interacted with that item.
Think about what that means for a news website. A story published four minutes ago has no audience yet. By the time enough people have clicked it for collaborative filtering to work, the story is old.
Same problem for a shop adding 500 products a week, or a music service where an unknown artist uploads a track that nobody has ever played.
Content-based filtering does not care. A brand-new item has a title, a description, a category and a picture on the day it is created. That is enough to place it next to similar items immediately.
How it works, in one picture
each item -> a list of words about it -> a row of numbers
Dangal -> drama sports wrestling -> [0, 0.4, 0.5, 0, 0.3, ...]
Lagaan -> drama sports cricket -> [0, 0.4, 0, 0, 0.3, ...]
Tumbbad -> horror fantasy mythology -> [0.6, 0, 0, 0.5, 0, ...]
your profile = the average of the rows you liked
score every item by how close its row is to your profile rowThe important move is turning words into numbers. Once a film is a row of numbers, "how alike are these two films" becomes arithmetic instead of an opinion.
The simple version counts words and downweights the common ones. The word "drama" appears on half the catalogue, so it says little about any one film. The word "wrestling" appears twice, so it says a lot. This weighting has a name: TF-IDF. It stands for term frequency times inverse document frequency. In plain words, that is how often a word appears here, divided by how ordinary it is everywhere.
The modern version replaces word counting with embeddings, which capture meaning rather than spelling. That change matters, and we come back to it below.
Where you have already seen it
- News apps showing you a story published minutes ago, related to what you read yesterday.
- A shopping site suggesting a case for the exact phone you looked at.
- Spotify placing a track by an artist with 40 listeners next to one with 40 million.
- A job board matching your saved searches against a posting made this morning.
The honest part
Content-based filtering has a boring, permanent weakness: it can only ever give you more of the same.
Read the picture above again. Your profile is built from what you liked. Every recommendation is scored by closeness to that profile. Nothing in the method can propose something unlike your history, because unlike-your-history is exactly what it scores as bad.
This is called over-specialisation. Read only cricket articles for a week and the app becomes a cricket-only app. It never suggests the astronomy piece you would have loved, because you have never read one.
Collaborative filtering has the opposite personality. It surprises you, because it discovers connections nobody wrote down. Content-based filtering never surprises you, and it never fails on a new item.
You need both. Combining them is a whole lesson: hybrid recommenders.
There is a second honest point. This method is only as good as your item descriptions. If the metadata is thin, wrong or copy-pasted by a supplier, no amount of clever maths rescues it. A lot of "the recommender is bad" turns out to be "the product catalogue is bad".
Remember this
- Content-based filtering compares items to items, using the item's own description.
- It handles brand-new items on day one, which collaborative filtering cannot.
- It cannot surprise you, and it is only as good as your metadata.
What to learn next
- Hybrid recommenders — using this and collaborative filtering together.
- The cold-start problem — the problem this method exists to solve.
- Embeddings — the modern replacement for word counting.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnNo dataset download. Six films with one line of tags each, so the whole thing runs instantly and you can check every number by hand.
A complete content-based recommender
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
# Everything the system is allowed to know about a film. No ratings, no audience.
FILMS = {
"Sholay": "action revenge friendship bandit village classic",
"Lagaan": "drama sports cricket village underdog colonial",
"3 Idiots": "drama comedy college friendship exams",
"Dangal": "drama sports wrestling family underdog biopic",
"Tumbbad": "horror fantasy mythology greed period",
"Stree": "horror comedy folklore ghost town",
}
TITLES = list(FILMS)
vec = TfidfVectorizer()
X = np.asarray(vec.fit_transform(FILMS.values()).toarray()) # rows are already unit length
def profile(liked):
"""A user is the average of the item vectors they liked, scaled back to unit length."""
p = X[[TITLES.index(t) for t in liked]].mean(axis=0)
return p / np.linalg.norm(p)
def recommend(liked, k=3):
scores = X @ profile(liked) # a dot product of unit vectors is cosine similarity
order = sorted(range(len(TITLES)), key=lambda i: (-scores[i], TITLES[i]))
return [(TITLES[i], float(scores[i])) for i in order if TITLES[i] not in liked][:k]
print("vocabulary size:", len(vec.get_feature_names_out()))
print("\nfilm against film, from the words alone:")
for a, b in [("Lagaan", "Dangal"), ("Tumbbad", "Stree"),
("Lagaan", "Tumbbad"), ("3 Idiots", "Stree")]:
print(f" {a:<9} vs {b:<9} {float(X[TITLES.index(a)] @ X[TITLES.index(b)]):.3f}")
for liked in (["Lagaan", "Dangal"], ["Tumbbad"], ["Sholay"]):
print(f"\nliked {liked} ->")
for title, score in recommend(liked):
print(f" {score:.3f} {title}")vocabulary size: 25 film against film, from the words alone: Lagaan vs Dangal 0.392 Tumbbad vs Stree 0.149 Lagaan vs Tumbbad 0.000 3 Idiots vs Stree 0.165 liked ['Lagaan', 'Dangal'] -> 0.136 3 Idiots 0.082 Sholay 0.000 Stree liked ['Tumbbad'] -> 0.149 Stree 0.000 3 Idiots 0.000 Dangal liked ['Sholay'] -> 0.149 3 Idiots 0.137 Lagaan 0.000 Dangal
Four things in that output are worth stopping on
Lagaan vs Dangal is 0.392, the highest pair. They share drama, sports and underdog. No user ever told the system these two go together.
Lagaan vs Tumbbad is exactly 0.000. They share no word at all, so cosine similarity is zero. Read that carefully: zero does not mean "unrelated". It means "no shared token". Two films could be about the same thing in different words and still score zero. Word matching has no idea that colonial and period are neighbours in meaning.
3 Idiots vs Stree is 0.165, and both are Hindi comedies. One is a college film, one is a horror comedy. A shared tag is a shallow kind of similarity, and this is what shallow looks like.
Look at the Tumbbad block. After Stree at 0.149, everything scores 0.000. Those are not third and fourth place. Those are ties at nothing, put in alphabetical order by the tie-break in sorted. A ranked list will hand you items with zero evidence and present them like recommendations. Put a floor under it:
picks = [(t, s) for t, s in recommend(liked) if s > 0.05]Line by line, for the parts that are not obvious
TfidfVectorizer() does three jobs at once: it splits text into tokens, counts them, and weights each count by how rare that word is across the whole catalogue. It also L2-normalises each row, which is why the dot product below is already a cosine.
X[[TITLES.index(t) for t in liked]].mean(axis=0) builds the user profile by averaging the item vectors. This is the Rocchio method from 1971 information retrieval, and it is still a reasonable default. Averaging assumes your taste has one centre. A person who likes horror and sports gets a profile that sits between the two and matches neither.
p / np.linalg.norm(p) — normalising the profile keeps scores comparable between a user who liked two items and one who liked two hundred. Skip it and heavy users get systematically larger scores.
X @ profile(liked) — one matrix-vector product scores the whole catalogue. Shape (6, 25) @ (25,) gives (6,). At real scale this becomes an approximate nearest-neighbour lookup instead; see vector databases.
Common mistakes
A word that appears on every item. Add hindi to all six descriptions above and re-run. Lagaan vs Tumbbad moves from 0.000 to 0.041. Scikit-learn's default smoothed IDF gives a term present in every document a weight of 1, not 0, so a constant field like language or country quietly inflates every similarity. Strip constant fields, or pass them in stop_words.
Averaging a split personality. One profile per user breaks for people with two unrelated tastes. Cluster the user's liked items first and keep one profile per cluster, then merge the ranked lists. The clustering lesson covers the first half of that.
Recommending near-duplicates. Content similarity is very good at finding the sequel, the re-release and the same product in a different colour. Users read a page of near-identical items as a broken system. De-duplicate on a content key before display, and see diversity and filter bubbles.
Assuming TF-IDF is the modern answer. It is a strong, cheap baseline and it is still widely deployed. It cannot match paraphrases. Swapping TfidfVectorizer for a sentence embedding model closes the colonial/period gap and costs you a model download and more compute per item.
Letting suppliers write the metadata. On a marketplace, item descriptions are written by sellers who are optimising for search, not for accuracy. Content-based quality is bounded by that text, and no model fixes it.
Try it yourself
Replace TfidfVectorizer with a sentence embedding model and re-run the same six films:
# pip install sentence-transformers (about 90 MB of model download, CPU is fine)
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
X = model.encode(list(FILMS.values()), normalize_embeddings=True)Everything below that line works unchanged, because both approaches produce unit-length rows.
No output block here on purpose. The exact numbers depend on the model version, so an invented sample would teach you to expect something that will not happen. What you should check is the ordering: Lagaan vs Tumbbad should stop being exactly zero, because the model knows those words are related even though they never co-occur.
What to learn next
- Hybrid recommenders — blending this with collaborative signals.
- Embeddings — meaning-based item vectors.
- Vector databases — running this over millions of items.
Researcher — Mathematics and papers.
The representation
Classic content-based recommenders represent item $i$ as a weighted term vector over a vocabulary $V$:
$$ w_{t,i} = \underbrace{\mathrm{tf}(t, i)}{\text{count in item } i} \times \underbrace{\log \frac{|I|}{\mathrm{df}(t)}}{\text{rarity across the catalogue}} $$
Where $t \in V$ is a term, $\mathrm{tf}(t,i)$ its frequency in item $i$'s document, $\mathrm{df}(t)$ the number of items containing $t$, and $|I|$ the catalogue size. Scikit-learn's default uses the smoothed form $\log\frac{1+|I|}{1+\mathrm{df}(t)} + 1$, whose additive $+1$ is precisely why a term occurring in every item retains non-zero weight.
Vectors are then L2-normalised and compared with cosine similarity:
$$ \mathrm{sim}(i, j) = \frac{\mathbf{w}_i \cdot \mathbf{w}_j}{|\mathbf{w}_i| \, |\mathbf{w}_j|} $$
The user profile
The Rocchio formulation (Rocchio, 1971) builds a profile from positive and negative feedback:
$$ \mathbf{p}_u = \alpha \mathbf{p}_u^{(0)} + \frac{\beta}{|D_u^+|}\sum_{i \in D_u^+} \mathbf{w}_i - \frac{\gamma}{|D_u^-|}\sum_{i \in D_u^-} \mathbf{w}_i $$
Where $D_u^+$ and $D_u^-$ are the sets of items $u$ liked and disliked, $\mathbf{p}_u^{(0)}$ any prior profile, and $\alpha, \beta, \gamma \ge 0$ weight the three terms. Typical settings weight positives far above negatives ($\beta \gg \gamma$), which reflects an empirical asymmetry: in recommendation, a positive signal is specific and a negative signal is diffuse.
The profile-as-centroid assumption is the method's real limitation. For a multimodal preference distribution the centroid lies in a low-density region and matches nothing. Modelling $\mathbf{p}u$ as a mixture, or keeping a per-item nearest-neighbour score $\max{j \in D_u^+} \mathrm{sim}(i,j)$ instead of a centroid, both address this and both are cheap.
Where the modern version diverges
Two changes account for most of the improvement since 2015.
Learned item representations. Rather than hand-built term vectors, encode item content with a trained model: a sentence encoder for text, a CNN or ViT for images, CLIP-style joint text-image encoders for both. Van den Oord, Dieleman and Schrauwen (2013), Deep content-based music recommendation, is the canonical demonstration: they trained a CNN on audio spectrograms to predict the latent factors of a collaborative filtering model, giving usable factors for tracks with zero plays.
Learned mapping instead of cosine. Cosine assumes item space and user space are the same space. Two-tower architectures learn separate encoders $f(u)$ and $g(i)$ into a shared space with a trained objective. Content features enter through $g$, so cold items get a position without any interaction history. This is now the default construction for cold-start-tolerant retrieval, and it is covered in two-tower retrieval models.
Over-specialisation, formally
Let $\mathcal{S}_u = { i : \mathrm{sim}(i, \mathbf{p}_u) > \tau }$ be the reachable set at threshold $\tau$. Since $\mathbf{p}_u$ is a convex combination of vectors in $D_u^+$, and cosine similarity is continuous, $\mathcal{S}_u$ is contained in a neighbourhood of the convex hull of the consumed items. Any item whose representation lies outside that neighbourhood is unreachable, regardless of the user's true utility for it.
Collaborative signals are the standard escape, because co-occurrence is not constrained by content geometry. This is the formal reason hybrids are not a convenience but a requirement. Adomavicius and Tuzhilin (2005) state the argument in its original form.
Evaluation caveat
Content-based methods look strong on offline metrics computed against logged data, and part of that is circular. Logged interactions were generated by a system that already ranked similar items highly, so items resembling the user's history are over-represented in the test set. Held-out accuracy therefore partly measures agreement with the previous policy. Report catalogue coverage and intra-list diversity alongside accuracy, or the number will flatter the method.
Papers
- Rocchio (1971), Relevance Feedback in Information Retrieval, in The SMART Retrieval System.
- Salton and Buckley (1988), Term-weighting approaches in automatic text retrieval, Information Processing and Management.
- Adomavicius and Tuzhilin (2005), Toward the Next Generation of Recommender Systems, IEEE TKDE — the survey that named the limitations precisely.
- Van den Oord, Dieleman and Schrauwen (2013), Deep content-based music recommendation, NIPS.
- Lops, de Gemmis and Semeraro (2011), Content-based Recommender Systems: State of the Art and Trends, in the Recommender Systems Handbook.
What to learn next
- Two-tower retrieval models — learned encoders replacing cosine on hand-built vectors.
- Hybrid recommenders — the formal escape from over-specialisation.
- Embeddings — how the item representation is learned.