The cold-start problem
A recommender learns from history, so it has nothing to say about a brand-new user or a brand-new item — and every real system carries a second, simpler system to cover that gap.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
The cold-start problem is what happens when a recommender meets somebody or something it has no history for.
The analogy you have already lived
Remember your first day at a new school. Nobody knew your name. You did not know who sat where, or who would become a friend.
For those first few days everybody treated you the same way they treat any new person — politely, generically, with no idea what you actually like. It took weeks of small interactions before anyone could say "you would enjoy this".
A recommender is in that position with every new user, every single day. And it is in the opposite position with every new item: a film nobody has watched is the quiet new student nobody has spoken to yet.
The three cold starts
A new user. They signed up ninety seconds ago. They have clicked nothing. Every method in this section needs history, and there is none.
A new item. A song uploaded this morning, a product listed today, a news story from four minutes ago. Collaborative filtering works from who-watched-what, so an item with no audience has no neighbours and cannot be recommended. It stays invisible, which means it gets no audience, which keeps it invisible.
A new system. You launched the product last week. There is no history for anyone or anything. This one is the hardest, and it is why brand-new apps ship with editorially chosen lists.
Why it exists
There is no clever fix hiding here. The cause is arithmetic, not engineering.
Every method in this section learns patterns from past behaviour. When there is no past behaviour, there is nothing to learn from. A model cannot extract a signal from an empty row.
So the honest framing is this: cold start is a data problem wearing a modelling costume. The solutions are all about finding a different kind of data to lean on until the usual kind arrives.
What actually works
Ask. Show a new user twelve well-chosen items and ask them to pick three. Thirty seconds of their time replaces weeks of guessing. Spotify, Netflix and news apps all do this at sign-up.
Use what the item is made of. A new film has a title, a genre, a cast and a poster on day one. Content-based filtering needs no audience, so it can place a new item next to similar old ones immediately.
Use what you already know about the person. Their city, their phone, the time of day, where they arrived from. None of it is taste, and all of it is better than nothing.
Show it to a few people on purpose. Deliberately put the new item in front of a small, well-chosen audience to find out how it performs. That costs you some quality now to buy information for later. It is called exploration, and it is the reason a system that only ever shows its safest bets slowly gets worse.
Fall back to popularity. Not personal, and not useless. The most-watched item is a reasonable guess for a person you know nothing about.
How the pieces fit together
does this person have history?
|
no ---> ask them to pick 3 things, or show popular items
|
yes
|
does this item have an audience?
|
no ---> place it by its tags, and show it to a small test group
|
yes ---> the normal recommenderEvery serious production system has this ladder in it, and the fallback rungs get far less attention than they deserve.
The honest part
Cold start does not get solved. It gets covered up, and the covering has a cost.
The fallback is usually popularity, and popularity is exactly the thing that makes big items bigger. So the mechanism protecting new users quietly pushes attention towards items that already have plenty.
There is a fairness problem sitting inside this, and it is not abstract. A new creator on a platform, a new seller in a marketplace, a small restaurant that opened last month — all of them start invisible. The system needs data about them to show them, and it needs to show them to get data. Being new is a real disadvantage, imposed by the software rather than by the audience.
Read diversity and filter bubbles next if this bothers you. It should.
Remember this
- Cold start hits new users, new items and new systems, and it is a shortage of data.
- The fixes are: ask, use item content, use context, explore, fall back to popularity.
- Every fallback has a cost, and the usual one makes popular items more popular.
What to learn next
- Hybrid recommenders — how the fallback ladder is built properly.
- Content-based filtering — the tool that covers new items.
- Diversity and filter bubbles — what the popularity fallback costs.
Developer — Code and libraries.
Setup
python3 --versionStandard library only. The point of this file is to make the failure visible, not to build a big model.
Watch collaborative filtering fail, then cover it
from collections import Counter
# Tags exist for every film from the day it is added to the catalogue.
TAGS = {
"Sholay": {"action", "revenge", "friendship", "classic"},
"Lagaan": {"drama", "sports", "underdog", "village"},
"3 Idiots": {"drama", "comedy", "college", "friendship"},
"Dangal": {"drama", "sports", "underdog", "family"},
"Tumbbad": {"horror", "fantasy", "mythology"},
"Stree": {"horror", "comedy", "folklore"},
"Naya-Film": {"drama", "sports", "underdog"}, # added today, nobody has seen it
}
# Watch history. Note that nobody has watched Naya-Film.
WATCHED = {
"Aarav": {"Sholay", "Lagaan", "3 Idiots"},
"Bhavna": {"Lagaan", "3 Idiots", "Dangal"},
"Chetan": {"Sholay", "Lagaan", "Dangal"},
"Divya": {"Tumbbad", "Stree"},
"Esha": {"Tumbbad", "Stree", "3 Idiots"},
"Gauri": set(), # signed up two minutes ago
}
CATALOGUE = sorted(TAGS)
POPULARITY = Counter(f for films in WATCHED.values() for f in films)
def collaborative(person, k=3):
"""Item-item co-watching. Needs history from this person AND an audience per film."""
scores = Counter()
for film in WATCHED[person]:
for other, films in WATCHED.items():
if other != person and film in films:
for candidate in films - WATCHED[person]:
scores[candidate] += 1
ranked = sorted(scores.items(), key=lambda p: (-p[1], p[0]))
return [f for f, s in ranked[:k] if s > 0]
def content(seed_films, person, k=3):
"""Tag overlap. Works on any film that has tags, watched or not."""
profile = Counter(t for f in seed_films for t in TAGS[f])
scored = []
for film in CATALOGUE:
if film in WATCHED[person] or film in seed_films:
continue
overlap = sum(profile[t] for t in TAGS[film])
scored.append((overlap, film))
scored.sort(key=lambda p: (-p[0], p[1]))
return [f for s, f in scored[:k] if s > 0]
def popular(person, k=3):
unseen = [f for f in CATALOGUE if f not in WATCHED[person]]
unseen.sort(key=lambda f: (-POPULARITY[f], f))
return unseen[:k]
print("can collaborative filtering serve this person?")
for person in WATCHED:
picks = collaborative(person)
print(f" {person:<7} history {len(WATCHED[person])} films -> {picks if picks else 'NOTHING'}")
print("\nhow often does each film get recommended by collaborative filtering?")
shown = Counter(f for p in WATCHED for f in collaborative(p))
for film in CATALOGUE:
print(f" {film:<10} {shown[film]} (watched by {POPULARITY[film]} people)")
print("\nfallback for the new user Gauri:")
print(" popularity ->", popular("Gauri"))
print(" after she picks Dangal ->", content({"Dangal"}, "Gauri"))can collaborative filtering serve this person? Aarav history 3 films -> ['Dangal', 'Stree', 'Tumbbad'] Bhavna history 3 films -> ['Sholay', 'Stree', 'Tumbbad'] Chetan history 3 films -> ['3 Idiots'] Divya history 2 films -> ['3 Idiots'] Esha history 3 films -> ['Lagaan', 'Dangal', 'Sholay'] Gauri history 0 films -> NOTHING how often does each film get recommended by collaborative filtering? 3 Idiots 2 (watched by 3 people) Dangal 2 (watched by 2 people) Lagaan 1 (watched by 3 people) Naya-Film 0 (watched by 0 people) Sholay 2 (watched by 2 people) Stree 2 (watched by 2 people) Tumbbad 2 (watched by 2 people) fallback for the new user Gauri: popularity -> ['3 Idiots', 'Lagaan', 'Dangal'] after she picks Dangal -> ['Lagaan', 'Naya-Film', '3 Idiots']
Three things in that output, in order of importance
Gauri gets NOTHING. No error, no warning, no exception. The function returns an empty list and a careless front end renders an empty row. This is what cold start looks like in a real codebase: a silently blank shelf, not a crash.
Naya-Film is recommended 0 times. Look at the middle block. Every other film gets recommended once or twice. The new film is recommended to nobody, by anybody, ever. It has no audience, so it has no co-watchers, so it can never enter a score. It is not ranked last. It is not ranked at all.
One question fixes both. After Gauri picks a single film during sign-up, the content fallback returns ['Lagaan', 'Naya-Film', '3 Idiots']. Naya-Film reaches second place, purely on its tags, on the day it was added.
That is the whole lesson in three lines of output. Thirty seconds of onboarding plus item metadata beats any amount of clever modelling on an empty row.
Line by line, for the parts that are not obvious
if picks else 'NOTHING' — an explicit branch for the empty case. Write this branch on purpose. A recommender that returns [] and a recommender that returns three bad items look identical to a monitoring dashboard that only counts errors.
Counter(f for p in WATCHED for f in collaborative(p)) — this is a coverage measurement: what fraction of the catalogue ever gets recommended to anybody. Almost nobody computes it, and it exposes cold start faster than any accuracy metric. An offline accuracy score cannot see Naya-Film at all, because Naya-Film is in nobody's test set either.
profile = Counter(t for f in seed_films for t in TAGS[f]) — the onboarding profile, built from tags rather than from behaviour. Notice it needs one film, not a hundred.
overlap = sum(profile[t] for t in TAGS[film]) — a plain tag-count match, kept deliberately crude. The production version weights rare tags higher, which is TF-IDF again; see content-based filtering.
Common mistakes
Letting an empty list reach the page. Every recommender call needs a documented fallback, and the fallback needs a fallback. Popularity is the usual last rung because it always returns something.
Onboarding with popular items. Showing new users the twelve most popular films to choose from tells you almost nothing, because most people like popular things. Pick onboarding items that split the audience — the ones where opinion divides. Those questions carry information.
Not logging which fallback fired. Tag every response with the strategy that produced it. Without that, your metrics average a personalised system and a popularity list together, and you cannot tell which is doing the work.
Assuming cold start is rare. On many consumer products, a large share of daily sessions come from users with almost no history. The cold path is not an edge case. It is a main path with a low-effort implementation.
Forgetting cold items age into warm ones. An item needs impressions to accumulate history. If the ranker never shows it, it never warms up. Reserve a small fraction of slots for under-exposed items, on purpose, and accept the short-term cost.
Try it yourself
Add a second brand-new film with horror tags and re-run. Confirm it also gets recommended 0 times by the collaborative path.
Then change Gauri's onboarding pick from Dangal to Tumbbad. The content fallback now returns a single film, ['Stree'], because that is the only other item sharing a tag with it.
Sit with that result. One onboarding answer swung the whole system, and it also exposed how thin the tags are on the horror side of this catalogue. Cold start and metadata quality are the same problem viewed from two directions.
What to learn next
- Hybrid recommenders — routing between these strategies cleanly.
- Two-tower retrieval models — putting content features inside the model so cold items get real vectors.
- A/B testing a recommender — measuring what exploration actually costs.
Researcher — Mathematics and papers.
Why the failure is structural
In a pure collaborative model the item representation is a free parameter $\mathbf{q}_i \in \mathbb{R}^k$ estimated only from observations involving $i$. For a new item, ${(u,i) \in \mathcal{K}} = \varnothing$, so the objective is independent of $\mathbf{q}_i$ and the regularised solution is $\mathbf{q}_i = \mathbf{0}$. The score $\mathbf{p}_u^\top \mathbf{q}_i$ is then identically zero for every user.
The item is not badly ranked. It is unidentifiable. No amount of optimisation changes that, because the gradient with respect to $\mathbf{q}_i$ has no data flowing into it.
The fix is to make the representation a function of features rather than a free parameter:
$$ \mathbf{q}i = g\theta(\mathbf{c}_i) $$
Where $\mathbf{c}i$ are content features (text, category, image, price, creator) and $g\theta$ is shared across all items. Now a new item inherits $\theta$, which was estimated from the whole catalogue, and gets a usable vector from features alone. This is the single most important structural idea in cold-start work, and it is why two-tower models with content features are the default answer today.
Approaches, with their trade-offs
Feature-based factorisation. Factorisation Machines (Rendle, 2010) put user, item and side features into one model with pairwise interactions, so a cold item is represented by its feature interactions. Cost: the feature engineering is now the model.
Hybrid with a learned gate. Train both a collaborative and a content model, and learn a mixing weight conditioned on interaction count. Cost: two models to maintain, and the gate needs its own validation.
DropoutNet (Volkovs, Yu and Poutanen, 2017) trains a network that consumes both collaborative and content input, and randomly zeroes the collaborative half during training. The model is therefore forced to perform when the preference input is missing, which is exactly the cold condition. The elegance is that no separate cold model exists — the warm model degrades gracefully by construction.
Meta-learning. MeLU (Lee et al., 2019) frames per-user adaptation as MAML: learn an initialisation from which a few gradient steps on a handful of interactions produce a good user model. Cost: substantially more complex training, and the reported gains are sensitive to the evaluation protocol.
Explicit exploration. Treat the new item as an arm of a bandit and allocate impressions to reduce uncertainty about its value. LinUCB (Li et al., 2010) does this with a linear reward model over context features and a confidence bound, and was evaluated on Yahoo front-page news, which is a maximally cold domain.
The cost of exploration, quantified
Exploration is not free, and the trade is explicit. For a contextual bandit over $T$ rounds with $d$-dimensional context, LinUCB-style algorithms achieve regret
$$ R(T) = O!\left( d \sqrt{T \log T} \right) $$
Where $R(T)$ is the difference in cumulative reward between the algorithm and an oracle that always picks the best arm. Regret growing as $\sqrt{T}$ means average per-round loss decays like $1/\sqrt{T}$ — exploration cost is real, bounded, and amortises.
The awkward part in recommendation is that the "arms" are items, $|I|$ is enormous, and item value is non-stationary. Practical deployments therefore explore over item clusters or over a small reserved traffic slice rather than over the raw catalogue.
Measure the right things
Accuracy metrics computed on logged data are blind to cold start, because a cold item appears in nobody's held-out set. Report alongside them:
- Catalogue coverage — the fraction of items appearing in any user's top-$N$.
- Cold-item exposure — impressions to items with age below some threshold, as a share of total impressions.
- Time-to-first-impression — the distribution of the delay between an item entering the catalogue and its first serve. A long tail here is a supply-side fairness problem.
- Gini coefficient over item exposure — inequality of attention, discussed in diversity and filter bubbles.
Papers
- Schein et al. (2002), Methods and Metrics for Cold-Start Recommendations, SIGIR — the paper that named and framed the problem.
- Rendle (2010), Factorization Machines, ICDM.
- Li et al. (2010), A Contextual-Bandit Approach to Personalized News Article Recommendation (LinUCB), WWW — arxiv.org/abs/1003.0146
- Van den Oord, Dieleman and Schrauwen (2013), Deep content-based music recommendation, NIPS.
- Volkovs, Yu and Poutanen (2017), DropoutNet: Addressing Cold Start in Recommender Systems, NeurIPS.
- Lee et al. (2019), MeLU: Meta-Learned User Preference Estimator for Cold-Start Recommendation, KDD — arxiv.org/abs/1908.00413
What to learn next
- Two-tower retrieval models — feature-conditioned item towers as the standard fix.
- Hybrid recommenders — combination strategies and their formal taxonomy.
- Diversity and filter bubbles — exposure inequality, measured.