AI Safety and Ethics

Bias in datasets

Dataset bias is a measurable property — who is in the data, what the labels reward, and which columns quietly stand in for identity.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Where the bias actually enters
  4. The picture
  5. Why deleting the column does not work
  6. Where you have already seen it
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A dataset is biased when the examples in it do not match the world the model will be used in.

The analogy you have already lived

Imagine cooking for a hostel where half the residents are vegetarian. You taste-test your menu on ten friends, and all ten eat meat. Every dish passes. Opening night, half the hall goes hungry.

Nothing was wrong with your tasting. The problem was who you tasted with.

That is dataset bias. The measurement was careful, the sample was not.

Where the bias actually enters

People imagine one villain writing prejudiced code. Real bias enters through three ordinary doors, and all three are measurable.

Door one: who is in the data. If a face dataset is built from photos scraped off English-language news sites, some faces are common and others are rare. Rare means fewer examples to learn from, which means worse accuracy.

Door two: what the labels reward. A hiring model is not trained on "who is good at the job". It is trained on "who got hired before". If past hiring was skewed, the model is being taught to copy the skew, faithfully.

Door three: the stand-in column. You delete the column naming someone's gender, caste or region, and you feel safer. But the pin code stayed. The college name stayed. The phone model stayed.

Those columns still carry the same information. A proxy is a harmless-looking column that quietly stands in for a sensitive one.

The picture

      the world                    your dataset
   ┌──────────────┐             ┌──────────────┐
   │  everybody   │  ──────►    │  whoever was │
   │              │  sampling   │  easy to     │
   │              │             │  collect     │
   └──────────────┘             └──────────────┘
                                       │
                                labels written by
                                past human decisions
                                       │
                                       ▼
                                 ┌──────────┐
                                 │  model   │  ──► repeats the pattern,
                                 └──────────┘      faster and at scale

The model is the last box, and it is the least interesting one. Nearly all of the damage was done before training started.

Why deleting the column does not work

This is the part that surprises careful, well-meaning engineers, so read it twice.

You remove the gender column from a CV dataset. The model still finds gender. It reads it from the college, the sport played, the gap in employment, the wording of the summary.

Removing a label does not remove the information. It removes your ability to measure what the model is doing with it.

To check for a bias, you have to keep the column that lets you check. That feels backwards, and it is the single most useful thing on this page.

Where you have already seen it

  • Speech systems that handle one accent well and another badly.
  • Camera auto-focus that locks onto some faces faster than others.
  • Translation that turns a gender-neutral sentence into a stereotyped one.
  • Credit scoring that keeps declining a whole neighbourhood.

What is honestly hard here

There is no unbiased dataset. Every dataset was collected by somebody, somewhere, with a budget and a deadline.

The goal is not purity. The goal is knowing your data's shape well enough to say who this model should not be used on. That sentence belongs in your documentation, and the lesson on model cards shows where to put it.

Remember this

  • Bias arrives through sampling, labels, and proxy columns — not through malice.
  • Deleting a sensitive column hides the problem instead of fixing it.
  • Keep the sensitive column for measurement, even when you never feed it to the model.

What to learn next

Developer — Code and libraries.

Make it a number, not an argument

Dataset bias is four measurements. Run them before you train anything, every time, as a fixed part of your pipeline.

  1. Representation — how many records per group.
  2. Label base rate — the fraction of positive labels within each group.
  3. Proxy leakage — can the group be predicted from the remaining columns?
  4. Downstream selection rate — how often the trained model says yes, per group.

Setup

bash
pip install numpy scikit-learn

Both are small, CPU-only, and download in seconds.

An audit you can run before training

The data below is synthetic, and built with one deliberate property: true ability is identical in both groups. Anything the audit finds is therefore a data artefact, not merit.

bias_audit.py
import numpy as np
from sklearn.linear_model import LogisticRegression

rng = np.random.default_rng(7)
N = 2000

# group: 0 = applicants from the college the company already hires from,
#         1 = applicants from everywhere else. The company's history skews the sample.
group = (rng.random(N) < 0.25).astype(int)          # only 25% of records are group 1

skill = rng.normal(0.0, 1.0, N)                     # true ability, identical in both groups

# "years at a partner company" is a PROXY: it tracks the group, not the skill
proxy = rng.normal(0.0, 1.0, N) + 1.4 * (1 - group)

# The historical label is the biased part. Past managers leaned on the proxy,
# so the group gap walked into the labels without anyone naming the group.
score = 1.2 * skill + 1.1 * proxy - 1.5 + rng.normal(0.0, 0.5, N)
hired = (score > 0).astype(int)

print("REPRESENTATION")
for g in (0, 1):
    print(f"  group {g}: {int((group == g).sum()):5d} records "
          f"({(group == g).mean():.1%} of data)")

print("TRUE ABILITY  (built to be identical, so any gap below is not merit)")
for g in (0, 1):
    print(f"  group {g}: mean skill {skill[group == g].mean():+.3f}")

print("LABEL BASE RATE  (fraction labelled 'hired' in the historical data)")
for g in (0, 1):
    print(f"  group {g}: {hired[group == g].mean():.3f}")
print(f"  gap: {hired[group == 0].mean() - hired[group == 1].mean():.3f}")

print("PROXY LEAKAGE  (can 'group' be predicted from the other columns?)")
X_all = np.c_[skill, proxy]
leak = LogisticRegression().fit(X_all, group)
print(f"  accuracy predicting group from skill+proxy: {leak.score(X_all, group):.3f}")
print(f"  always-guess-group-0 baseline:              {(group == 0).mean():.3f}")

print("MODEL TRAINED WITHOUT THE GROUP COLUMN")
clf = LogisticRegression().fit(X_all, hired)
pred = clf.predict(X_all)
for g in (0, 1):
    print(f"  group {g} selection rate: {pred[group == g].mean():.3f}")
print(f"  gap: {pred[group == 0].mean() - pred[group == 1].mean():.3f}")
Output
REPRESENTATION
  group 0:  1514 records (75.7% of data)
  group 1:   486 records (24.3% of data)
TRUE ABILITY  (built to be identical, so any gap below is not merit)
  group 0: mean skill +0.004
  group 1: mean skill -0.004
LABEL BASE RATE  (fraction labelled 'hired' in the historical data)
  group 0: 0.498
  group 1: 0.183
  gap: 0.315
PROXY LEAKAGE  (can 'group' be predicted from the other columns?)
  accuracy predicting group from skill+proxy: 0.816
  always-guess-group-0 baseline:              0.757
MODEL TRAINED WITHOUT THE GROUP COLUMN
  group 0 selection rate: 0.500
  group 1 selection rate: 0.173
  gap: 0.327

What the four blocks tell you

Representation. Group 1 holds 24.3% of the records. Every metric you compute on it is roughly twice as noisy as the same metric on group 0. Small groups produce wide error bars, and wide error bars get mistaken for good news.

True ability. The mean skill difference is +0.004 against -0.004, which is nothing. This line exists to remove one escape route: whatever gap the model produces, it did not come from ability.

Label base rate. A gap of 0.315. This is the largest single source of harm in the whole pipeline, and no modelling choice will remove it. The labels record what happened, not what should have happened.

Proxy leakage. Group is predicted with 0.816 accuracy from columns that never mention it, against a 0.757 baseline from always guessing the majority. That gap is the leak. Compare against the majority-class baseline, never against 0.5, or an imbalanced dataset will look leaky when it is not.

Selection rate. The model never saw the group column, and reproduced the gap anyway: 0.327 against a historical 0.315. Blindness is not fairness. It is the same discrimination with the audit trail deleted.

Common mistakes

Dropping the sensitive attribute and calling it done. You have removed your ability to measure, not the behaviour. Store the attribute in a separate, access-controlled evaluation table, keep it out of the feature matrix, and use it only for auditing.

Rebalancing the features and forgetting the labels. Oversampling group 1 fixes representation and leaves the base-rate gap untouched. They are different defects with different fixes — see imbalanced data for the sampling half.

Auditing after training. By then the metric you care about has become a launch blocker, and launch blockers get negotiated away. Run the audit on the raw data, in the same script that loads it.

One binary group. Real attributes intersect. A model can look fine on gender and fine on region, and fail badly on one combination of the two. Report the worst intersecting cell that has enough samples to be meaningful, and state that sample count next to it.

Assuming a bigger dataset fixes it. Scaling a skewed collection process produces a bigger skewed dataset. Sampling bias does not average out with volume; it gets a tighter confidence interval around the wrong number.

Try it yourself

Set the proxy strength 1.4 to 0.0 so the proxy no longer tracks group, and rerun. The leakage accuracy should collapse toward the 0.757 baseline and the selection-rate gap should nearly vanish. Then put 1.4 back and try dropping the proxy column from X_all instead — watch what that costs you in overall accuracy. That trade is the actual decision, and it belongs to a human.

What to learn next

Researcher — Mathematics and papers.

A taxonomy that maps to fixes

Suresh and Guttag (2021), A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle, is the reference decomposition. The value of a taxonomy here is that each source admits a different intervention.

  • Historical bias — the world itself contains the disparity, and the data records it faithfully. Present even with perfect sampling and perfect labels. No data-processing fix exists; this is a policy question about whether to model the outcome at all.
  • Representation bias — the sampled population under-covers a subgroup. Fixable by targeted collection; partially mitigable by reweighting.
  • Measurement bias — the recorded feature or label is a distorted proxy for the construct. Arrest records as a proxy for crime; healthcare spending as a proxy for illness.
  • Aggregation bias — one model is fit across subpopulations with genuinely different conditional distributions $P(Y \mid X, A = a)$. Fixable with per-group models or interaction terms.
  • Evaluation bias — the benchmark itself is unrepresentative, so model selection optimises the wrong thing.
  • Deployment bias — the system is used for a purpose different from the one it was validated for.

The proxy problem, formally

Let $A$ be the protected attribute and $X$ the remaining features. Removing $A$ from the feature set — fairness through unawareness — constrains the model to $f(X)$. This provides no guarantee whenever

$$ I(X; A) > 0 $$

where $I(\cdot\,;\cdot)$ is mutual information. In real tabular data $I(X; A)$ is substantial: postcode, name, device, and browsing history are all informative about most protected attributes. Dwork et al. (2012) named and rejected unawareness for exactly this reason.

The measurable version is the leakage audit in the developer block: train an adversary $g: X \to A$ and report its skill above the majority-class baseline. Balanced accuracy or AUC is the safer statistic when groups are skewed. A useful stronger form is the residual audit — after fitting $\hat{Y} = f(X)$, test whether $A$ still predicts the residual.

Measurement bias, quantified

Obermeyer et al. (2019), Dissecting racial bias in an algorithm used to manage the health of populations (Science), is the canonical documented case. A widely deployed US risk tool predicted future healthcare cost as a stand-in for future healthcare need. Because less money had historically been spent on Black patients at equal illness severity, the model systematically under-flagged them. At a fixed risk score, Black patients had substantially more chronic conditions than white patients.

Two points generalise. The algorithm contained no race variable. And the defect lived entirely in the choice of target variable, which is a modelling decision made before any data is touched. Target selection is the highest-leverage fairness decision in most pipelines, and it is almost never reviewed.

Reweighting and its limits

Kamiran and Calders (2012) give the standard preprocessing reweighting scheme. Each record receives

$$ w(a, y) = \frac{P(A = a)\,P(Y = y)}{P(A = a, Y = y)} $$

where $P(A=a)$ is the marginal group frequency, $P(Y=y)$ the marginal label frequency, and $P(A=a, Y=y)$ the observed joint frequency. Weighting by $w$ makes $A$ and $Y$ statistically independent in the reweighted sample, which enforces demographic parity at the data level.

The limitation is structural. Reweighting adjusts the marginal association between $A$ and $Y$; it cannot repair label noise that is itself group-dependent. If labels for one group are corrupted at a different rate, no reweighting of those labels recovers the truth. Fogliato et al. (2020) analyse this for proxy outcomes and show the bias can be amplified rather than reduced.

Documentation as infrastructure

Gebru et al. (2018), Datasheets for Datasets, proposes a standard record covering motivation, composition, collection process, preprocessing, uses, distribution and maintenance. Adoption is uneven, but the format is worth following even privately, because the questions surface omissions that metrics cannot: who was excluded at collection time, and who was paid what to produce the labels.

Bender and Friedman (2018), Data Statements for NLP, is the parallel proposal for language data, with an emphasis on speaker demographics and curation rationale.

Papers

  • Suresh and Guttag, A Framework for Understanding Sources of Harm throughout the ML Life Cycle, 2021 — arxiv.org/abs/1901.10002
  • Obermeyer et al., Dissecting racial bias in an algorithm used to manage the health of populations, Science 2019
  • Dwork et al., Fairness Through Awareness, ITCS 2012 — arxiv.org/abs/1104.3913
  • Kamiran and Calders, Data preprocessing techniques for classification without discrimination, KAIS 2012
  • Gebru et al., Datasheets for Datasets, 2018 — arxiv.org/abs/1803.09010
  • Bender and Friedman, Data Statements for Natural Language Processing, TACL 2018

What to learn next

What to learn next

These follow on from what you just read.

  • AI Safety and Ethics

    Fairness metrics

    Fairness has several precise definitions that cannot all hold at once — so the job is choosing one on purpose and measuring it.

  • AI Safety and Ethics

    Explainability

    Explainability is the engineering of checkable answers to "why did the model say that" — and different methods answer different questions.

  • AI Safety and Ethics

    SHAP and LIME

    SHAP splits a single prediction into a fair share per feature; LIME fits a small readable model near one point — here both are built from scratch.