Recommender Systems

Diversity and filter bubbles

Optimising a recommender for accuracy alone concentrates attention on a few items and narrows what each person sees, and the fixes cost measurable engagement, so this is a trade-off rather than a bug to remove.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Two different problems, often confused
  4. Why the machine drifts this way
  5. The part that is genuinely contested
  6. What actually helps
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A recommender trained only to be accurate will show you a narrower and narrower slice of the catalogue, and will bury most of it from everybody.

The analogy you have already lived

Imagine a canteen that watched what you ate for a week and decided you like rajma chawal. From then on, that is what lands on your plate. Every single day.

It is not wrong. You do like rajma chawal. You will keep eating it, and the canteen's records will keep confirming it was a good choice.

But you will never find out that you would have loved the fish curry. Nobody ever put it in front of you. And the cook who makes the fish curry eventually stops making it, because nobody orders it.

That is the whole lesson. The system is right about you today, and it is quietly making tomorrow smaller.

Two different problems, often confused

Popularity bias is about items. A few items get most of the attention, and most of the catalogue is shown to nobody. This one is measured easily and it is not in dispute.

The filter bubble is about people. The claim is that each person ends up in a narrowed world of similar content and similar opinions. This one is harder to measure and, as we will see, harder to prove than most articles about it suggest.

They come from the same mechanism, and they need to be discussed separately, because the evidence for them is not equally strong.

Why the machine drifts this way

Nothing malicious is happening. Popularity bias falls out of the arithmetic.

   an item is shown  ->  some people interact with it
          ^                          |
          |                          v
   the system's estimate  <-  it looks more popular
   of that item goes up

Look at that loop. It has no entrance for an item that was never shown. An item that gets a slow start gets fewer impressions, so it gathers less evidence, so it stays low forever.

The item's actual quality never enters the loop. This is worth reading twice, because it is the thing people find hardest to believe: an item can be buried permanently without ever having been bad. The developer block measures exactly that, and finds the buried items are the same average quality as the shown ones.

The part that is genuinely contested

Here is where careful honesty matters more than a good story.

Well supported: recommenders concentrate exposure on a small part of the catalogue. Users who follow recommendations see a narrower range of content over time than users who do not. New creators and new items face a real structural disadvantage.

Not well supported: the strong claim that recommendation feeds are a main cause of political polarisation.

Large field experiments have tested that stronger claim directly. In 2023, researchers working with Meta switched tens of thousands of consenting users to a plain reverse-chronological feed for three months during a US election campaign. Time on the platform dropped and the content mix changed, but measured political attitudes and polarisation did not move significantly. A companion study that reduced exposure to like-minded content found the same absence of effect on attitudes.

Those studies have real limitations, and researchers argue about them. Three months is short. The people who agreed to take part are not everybody. Turning the algorithm off for some users does not change the wider information environment they live in.

So the correct summary is uncomfortable and true: the exposure effects are measurable, and the strong attitude effects have not been demonstrated. Anyone who tells you either "filter bubbles are a myth" or "the algorithm radicalises people" is going further than the evidence goes.

What actually helps

Re-rank for variety. After the ranker scores everything, build the list one slot at a time, penalising each candidate for being similar to what you already picked. This is the standard method and the developer block implements it.

Match the user's own mix. Say somebody watches 70 percent comedy and 30 percent documentary. A list that is 100 percent comedy distorts their taste rather than reflecting it. Matching the proportions is called calibration.

Reserve slots. Keep a small, fixed share of impressions for items that have not had a fair chance yet. Not as a reward, as a measurement.

Measure the right things. Add catalogue coverage and exposure inequality to your dashboard beside accuracy. What is not measured is not managed, and almost nobody measures these.

The honest part

Every one of those fixes costs you something, immediately and measurably.

Show a less-relevant item and fewer people click it. That lands in this week's numbers as a loss.

The benefit arrives much later, if at all, and is far harder to attribute. A healthier catalogue. More creators who keep publishing. Users who have not got bored by month six.

That asymmetry is the real reason diversity work is hard. It is not that engineers do not care. It is that the cost is immediate and legible, and the benefit is delayed and arguable.

If you take one thing from this lesson, take this: the diversity setting is a business decision wearing a technical costume. Someone has to decide how much short-term engagement the organisation is willing to trade. That someone should not be the person who happened to pick a default value.

Remember this

  • Popularity bias is measurable, mechanical, and independent of item quality.
  • The strong filter-bubble claim about polarisation is not well supported by the best experiments.
  • Every fix costs measurable engagement now for benefits that arrive later.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Two programs. The first measures what an accuracy-only ranker does to a catalogue. The second fixes a single list and shows the price.

Program one: watch the catalogue collapse

A crude simulation on purpose. One global popularity number, no personalisation, no clever model. If the effect appears here it is mechanical, not a property of any particular algorithm.

feedback_loop.py
import numpy as np

N_USERS, N_ITEMS, ROUNDS = 200, 50, 15


def gini(counts):
    """0 = every item shown equally often, 1 = one item takes every impression."""
    x = np.sort(np.asarray(counts, dtype=float))
    n = len(x)
    if x.sum() == 0:
        return 0.0
    return float((2 * np.arange(1, n + 1) - n - 1) @ x / (n * x.sum()))


def simulate(epsilon, seed=0):
    rng = np.random.default_rng(seed)
    user_genre = rng.integers(0, 2, N_USERS)
    item_genre = rng.integers(0, 2, N_ITEMS)
    quality = rng.uniform(0.3, 0.9, N_ITEMS)     # fixed, real, and unknown to the system
    shown = np.zeros(N_ITEMS)
    belief = np.ones(N_ITEMS)                    # what the system thinks, starts flat
    history = [set() for _ in range(N_USERS)]

    for _ in range(ROUNDS):
        for u in range(N_USERS):
            unseen = [i for i in range(N_ITEMS) if i not in history[u]]
            if not unseen:
                continue
            if rng.random() < epsilon:
                pick = int(rng.choice(unseen))            # deliberate exploration
            else:
                pick = max(unseen, key=lambda i: belief[i])
            shown[pick] += 1
            history[u].add(pick)
            match = 1.0 if item_genre[pick] == user_genre[u] else 0.25
            if rng.random() < quality[pick] * match:      # the user enjoyed it
                belief[pick] += 1                         # and the belief grows
    return shown, quality


print("15 rounds, 200 people, a catalogue of 50\n")
print(f"{'exploration':>12}{'gini':>8}{'never shown':>13}{'mean quality shown':>21}{'not shown':>11}")
for eps in (0.0, 0.05, 0.1, 0.3):
    shown, quality = simulate(eps)
    seen = shown > 0
    q_seen = quality[seen].mean()
    q_unseen = f"{quality[~seen].mean():.3f}" if (~seen).any() else "-"
    print(f"{eps:>11.0%}{gini(shown):>8.3f}{int((~seen).sum()):>9}/{N_ITEMS:<3}"
          f"{q_seen:>21.3f}{q_unseen:>11}")
Output
15 rounds, 200 people, a catalogue of 50

 exploration    gini  never shown   mean quality shown  not shown
         0%   0.700       35/50                 0.610      0.613
         5%   0.684        3/50                 0.612      0.615
        10%   0.648        0/50                 0.612          -
        30%   0.533        0/50                 0.612          -

These numbers come from a fixed seed. NumPy does not promise an identical random stream across major versions, so if your third decimal differs, the pattern is the result, not the digits.

Read the last two columns before anything else

With no exploration, 35 of the 50 items were never shown to a single person, across three thousand impressions.

Now the column that matters. The mean quality of those 35 buried items is 0.613. The mean quality of the 15 items that got every impression is 0.610.

The buried items are not worse. They are very slightly better, and the difference is noise. They lost a race that was decided in the first few rounds by which items happened to get an early positive response.

Nothing in the loop ever asked about quality. Quality only enters through interactions, and interactions only happen after an impression, and impressions only go to items that already look good. An item that starts slowly can never recover, because it is never given the chance that would let it recover.

Five percent exploration cuts the never-shown count from 35 to 3. That is a cheap fix with a real price, and the price is that 5 percent of your impressions go to a randomly chosen item.

Program two: buy diversity in one list, and see the bill

Maximal marginal relevance builds the list one slot at a time. Each candidate is scored on its own relevance, minus how similar it is to what is already in the list.

mmr.py
import numpy as np

ITEMS = ["Dangal", "Lagaan", "Chak-De", "Sultan", "3 Idiots", "Piku", "Tumbbad", "Stree"]
FEATURES = ["drama", "sports", "comedy", "horror", "family"]
TAGS = np.array([
    [1, 1, 0, 0, 1],   # Dangal
    [1, 1, 0, 0, 0],   # Lagaan
    [1, 1, 0, 0, 0],   # Chak-De
    [1, 1, 0, 0, 1],   # Sultan
    [1, 0, 1, 0, 0],   # 3 Idiots
    [1, 0, 1, 0, 1],   # Piku
    [0, 0, 0, 1, 0],   # Tumbbad
    [0, 0, 1, 1, 0],   # Stree
], dtype=float)
REL = np.array([0.92, 0.90, 0.88, 0.86, 0.62, 0.58, 0.40, 0.35])   # the ranker's scores

T = TAGS / np.linalg.norm(TAGS, axis=1, keepdims=True)
SIM = T @ T.T


def mmr(lmbda, k=4):
    chosen = []
    while len(chosen) < k:
        best, best_score = None, -1e9
        for i in range(len(ITEMS)):
            if i in chosen:
                continue
            penalty = max((SIM[i, j] for j in chosen), default=0.0)
            score = lmbda * REL[i] - (1 - lmbda) * penalty
            if score > best_score + 1e-12:
                best, best_score = i, score
        chosen.append(best)
    return chosen


def ild(chosen):
    """Intra-list diversity: the mean distance between every pair in the list."""
    return float(np.mean([1 - SIM[a, b] for x, a in enumerate(chosen) for b in chosen[x + 1:]]))


print(f"{'lambda':>7}  {'the four films shown':<44}{'sum rel':>9}{'diversity':>11}")
for l in (1.0, 0.9, 0.8, 0.7, 0.65, 0.6, 0.5):
    c = mmr(l)
    print(f"{l:>7.2f}  {', '.join(ITEMS[i] for i in c):<44}{REL[c].sum():>9.2f}{ild(c):>11.3f}")
Output
 lambda  the four films shown                          sum rel  diversity
   1.00  Dangal, Lagaan, Chak-De, Sultan                  3.56      0.122
   0.90  Dangal, Lagaan, Chak-De, Sultan                  3.56      0.122
   0.80  Dangal, Lagaan, Chak-De, Sultan                  3.56      0.122
   0.70  Dangal, Lagaan, Chak-De, Sultan                  3.56      0.122
   0.65  Dangal, Lagaan, Tumbbad, 3 Idiots                2.84      0.713
   0.60  Dangal, Tumbbad, Lagaan, 3 Idiots                2.84      0.713
   0.50  Dangal, Tumbbad, 3 Idiots, Lagaan                2.84      0.713

What the two ends of that table look like to a user

At lambda = 1.0, pure relevance, the page shows four sports dramas in a row. Every one of them is a good pick on its own. Together they read as a broken system, because the fourth one tells the user nothing the first one did not.

At lambda = 0.65, the page shows a sports drama, another sports drama, a horror film and a comedy. Total relevance dropped from 3.56 to 2.84, a loss of 20 percent. Diversity went from 0.122 to 0.713.

That is the bill, printed. Twenty percent of your relevance score, for a list a person might actually browse.

Two honest notes about this table. The change is a step rather than a slope, because eight items in two clear clusters give MMR nothing in between; on a real catalogue with thousands of items at intermediate similarity, the curve is much smoother. And sum rel is not user happiness — it is the old model's opinion, which is exactly the thing you suspect of being too narrow.

Line by line, for the parts that are not obvious

belief = np.ones(N_ITEMS) — the system starts with a flat, equal prior. There is no built-in favouritism anywhere in the simulation. The inequality is produced entirely by the loop.

max((SIM[i, j] for j in chosen), default=0.0) — MMR penalises by the maximum similarity to anything already chosen, not by the average. Using the average lets a candidate that duplicates one item hide behind being unlike the other three.

score > best_score + 1e-12 — an explicit epsilon so that exact ties resolve to the lower index deterministically. Floating-point ties are common when items share tags, and without this the list can change between runs.

gini — the standard sorted-array form. It is a single number for "how unequally is attention distributed". Report it per week and watch the direction, not the absolute value.

Common mistakes

Reporting only accuracy. A ranker that shows the same twelve items to everybody can score beautifully on nDCG. Coverage and Gini are the metrics that catch it, and they take an hour to add.

Adding diversity by shuffling. Random noise in the list lowers relevance and raises nothing that matters. MMR removes redundancy, which is a different operation from adding randomness.

Diversifying on the wrong feature. Variety in genre is meaningful. Variety in thumbnail colour is not. The similarity function inside MMR is the actual definition of "diverse" that your product is using, and it deserves more thought than the lambda.

Confusing diversity with novelty and serendipity. Diversity is within one list. Novelty is about items the user has not met. Serendipity is a relevant surprise. A list can be perfectly diverse and contain nothing new.

Shipping a diversity change without an experiment. It will cost short-term engagement. Measure it, name the cost, and let somebody decide whether the organisation accepts it — see A/B testing a recommender.

Try it yourself

In feedback_loop.py, change belief[pick] += 1 to also decay: belief *= 0.999 at the end of each round. Re-run and see whether forgetting slowly is enough on its own to keep the catalogue open.

Then add an impression floor instead: force the least-shown item into one slot every round. Compare the Gini and the never-shown count against the exploration rows. You are now comparing two real production strategies on the same simulation, which is more than a lot of shipped systems ever do.

What to learn next

Researcher — Mathematics and papers.

Beyond-accuracy measurements

Let $L_u$ be the top-$N$ list for user $u$, $d(i,j)$ a distance in item feature space, and $p(i)$ the popularity share of item $i$.

Intra-list diversity (Ziegler et al., 2005):

$$ \mathrm{ILD}(L_u) = \frac{2}{|L_u|(|L_u|-1)} \sum_{i \in L_u} \sum_{j \in L_u,\, j > i} d(i,j) $$

Catalogue coverage: $\left| \bigcup_u L_u \right| / |I|$.

Novelty as mean self-information:

$$ \mathrm{Nov}(L_u) = \frac{1}{|L_u|} \sum_{i \in L_u} -\log_2 p(i) $$

Gini coefficient over exposure counts $x_{(1)} \le \dots \le x_{(n)}$:

$$ G = \frac{\sum_{r=1}^{n} (2r - n - 1)\, x_{(r)}}{n \sum_{r=1}^{n} x_{(r)}} $$

Gini and coverage answer different questions. Coverage asks whether an item is ever shown; Gini asks how unequally impressions are split among those that are. A system can have full coverage and a Gini of 0.9.

Re-ranking methods

MMR (Carbonell and Goldstein, 1998), greedy and $O(N k)$ per list:

$$ \mathrm{MMR} = \operatorname*{arg\,max}{i \in R \setminus S} \left[ \lambda\, \mathrm{rel}(i) - (1-\lambda) \max{j \in S} \mathrm{sim}(i, j) \right] $$

Determinantal point processes (Chen, Zhang and Zhou, 2018) model the whole slate at once. With a kernel $L_{ij} = \mathrm{rel}(i)\,\mathrm{rel}(j)\,\mathrm{sim}(i,j)$, the probability of selecting set $S$ is proportional to $\det(L_S)$, so overlapping items shrink the determinant geometrically. Exact MAP inference is NP-hard; the greedy algorithm is submodular-optimal to within $1 - 1/e$, and the fast greedy variant runs in $O(N k^2)$, which is practical for slate re-ranking.

Calibration (Steck, 2018) matches the genre distribution of the recommendations $q(g \mid u)$ to the user's own history distribution $p(g \mid u)$ by penalising Kullback–Leibler divergence:

$$ \max_{S} \; \sum_{i \in S} \mathrm{rel}(i) \; - \; \gamma \, D_{\mathrm{KL}}!\left( p(\cdot \mid u) \,|\, \tilde{q}(\cdot \mid S) \right) $$

The framing is worth noting. Calibration is not a diversity injection; it is a bias correction. An accuracy-maximising ranker systematically over-represents a user's majority interest relative to their own revealed proportions, and calibration removes that distortion. It is easier to defend to a product team than "we made it less accurate on purpose".

Popularity bias and feedback loops

Abdollahpouri, Burke and Mobasher (2017, 2019) document that collaborative methods amplify popularity beyond the base rate in the data, and that the amplification is unequal across user groups: users with niche taste are served worse by the same model.

Chaney, Stewart and Engelhardt (2018) simulate the closed loop and show homogenisation: user utility falls and inter-user recommendation similarity rises as the loop iterates, even when each individual round is locally optimal. Jiang et al. (2019), Degenerate Feedback Loops in Recommender Systems, give conditions under which user interest degenerates towards an extreme, and show that the degeneracy speed depends on model capacity and on the amount of exploration.

The mechanism is the missing-not-at-random structure discussed in implicit feedback. It is a property of the loop, not of the model class.

The empirical state of the filter-bubble claim

This literature is frequently misreported, so state it precisely.

Content narrowing. Nguyen et al. (2014), Exploring the filter bubble, tracked MovieLens users and found the content diversity of consumed items narrowed over time, with a smaller narrowing for users who followed recommendations more closely than for those who did not. This is a measured within-platform narrowing, on a films dataset, and it is the kind of claim the evidence supports well.

Political effects. Bakshy, Messing and Adamic (2015, Science) found on Facebook that individual choice accounted for more of the reduction in cross-cutting exposure than ranking did, on a sample of users who self-declare political affiliation.

The 2020 US Facebook and Instagram Election Study ran randomised interventions with tens of thousands of consenting users. Guess et al. (2023, Science) replaced the ranked feed with reverse-chronological for three months: it reduced time on platform and changed the content mix, and produced no detectable change in political attitudes or polarisation. Nyhan et al. (2023, Nature) reduced exposure to like-minded content and likewise found no detectable attitude effect. Gonzalez-Bailon et al. (2023, Science) documented substantial ideological segregation in news exposure on the platform, with the strongest segregation on the audience side rather than the algorithmic one.

The limitations are real and are argued about in print: three months is short relative to attitude formation, consenting participants are a selected sample, the interventions cannot alter the wider information environment, and null results on attitudes do not license claims about other outcomes.

The defensible summary: exposure concentration and content narrowing are measurable and reproducible. The causal claim that feed ranking drives political polarisation is not supported by the strongest available experiments, and those experiments do not settle every version of the question.

Report your own system's exposure statistics. They are measurable and specific, and they are a stronger basis for a design decision than any general claim about the internet.

Papers

  • Carbonell and Goldstein (1998), The Use of MMR, Diversity-Based Reranking, SIGIR.
  • Ziegler et al. (2005), Improving Recommendation Lists Through Topic Diversification, WWW.
  • Nguyen et al. (2014), Exploring the Filter Bubble, WWW.
  • Bakshy, Messing and Adamic (2015), Exposure to ideologically diverse news and opinion on Facebook, Science.
  • Abdollahpouri, Burke and Mobasher (2017), Controlling Popularity Bias in Learning-to-Rank Recommendation, RecSys.
  • Chaney, Stewart and Engelhardt (2018), How Algorithmic Confounding Increases Homogeneity — arxiv.org/abs/1710.11214
  • Steck (2018), Calibrated Recommendations, RecSys.
  • Chen, Zhang and Zhou (2018), Fast Greedy MAP Inference for DPP, NeurIPS — arxiv.org/abs/1709.05135
  • Jiang et al. (2019), Degenerate Feedback Loops in Recommender Systems, AIES — arxiv.org/abs/1902.10730
  • Guess et al. (2023), How do social media feed algorithms affect attitudes and behavior in an election campaign?, Science.
  • Nyhan et al. (2023), Like-minded sources on Facebook are prevalent but not polarizing, Nature.

What to learn next