AI Safety and Ethics

Fairness metrics

Fairness has several precise definitions that cannot all hold at once — so the job is choosing one on purpose and measuring it.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. The three questions people mean by "fair"
  4. They cannot all be true at once
  5. The picture
  6. Where you have already seen the fight
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A fairness metric is a specific, countable definition of "treated equally" — and there are several, which disagree.

The analogy you have already lived

Think of one bathroom and a queue of people. What is the fair rule?

First come, first served treats everyone the same. But the person who has been waiting since morning and the person who wandered up ten seconds ago are not in the same situation. Letting the desperate one in first also feels fair.

Both rules are fair. They are fair in different senses, and you cannot follow both at the same time.

That is the whole of this lesson. Fairness is not one thing waiting to be measured. It is a family of things, and picking one means giving up another.

The three questions people mean by "fair"

Suppose a bank uses a model to approve loans. Three different people complain, and all three are right.

"You approve fewer of us." This is about the share approved in each group. Making those shares equal is called demographic parity — equal approval rates, regardless of anything else.

"You reject people from our group who would have repaid." This is about who gets missed. Making the catch-rate equal is called equal opportunity — of the people who would repay, the same fraction gets approved in each group.

"Your approvals mean something different for us." This is about what a yes is worth. Making that equal is called predictive parity — among approved people, the same fraction actually repays.

Each is reasonable. Each is measurable. And here is the hard part.

They cannot all be true at once

Suppose the two groups repay at different underlying rates. That can happen for any reason, including unfair ones inherited from the past. When it does, no model can satisfy all three at once.

This is not a limit of today's technology. It is arithmetic, proved in 2016, and it will still be true in fifty years.

So there is no fair model, only a model that is fair in a way you chose and can defend. Anyone who says their system is "unbiased" without naming a metric has not measured anything.

The picture

                 different base rates in the two groups
                                │
                ┌───────────────┼───────────────┐
                ▼               ▼               ▼
     equal approval      equal catch-rate   equal meaning
        rates             for the           of a "yes"
                          deserving
                │               │               │
                └──────► pick ONE, on purpose ◄─┘
                          then measure it
                          then publish it

Where you have already seen the fight

  • Exam cut-offs and reservation policies argue over exactly these three definitions.
  • Insurance pricing by area is demographic parity versus accuracy, in public.
  • Fraud checks that stop more cards from one region are an equal-catch-rate argument.

What is honestly hard here

Choosing the metric is not a technical decision, and engineers should not make it alone.

The right people to choose are the ones accountable for the outcome, informed by someone who understands the trade. Your job is to make the trade visible: here is what we give up, in numbers, if we pick this one.

That is a much more useful contribution than an opinion.

Remember this

  • Fairness has several precise definitions, and they conflict by arithmetic.
  • With different base rates, no model satisfies all of them at once.
  • Choose one deliberately, measure it, and write down what you gave up.

What to learn next

Developer — Code and libraries.

Four numbers, one table

Every group-fairness metric is a comparison of one rate across groups. Compute all four rates per group and you can read off any definition you need.

RateMeaningMetric that equalises it
selection rateshare of the group that gets a yesdemographic parity
TPRof those who truly qualify, share approvedequal opportunity
FPRof those who truly do not, share approved(with TPR) equalised odds
PPVof those approved, share who truly qualifypredictive parity

Setup

bash
pip install numpy

The impossibility, run on your own machine

The two groups below have different repayment base rates by construction. The scorer is group-blind and equally skilled on both groups.

fairness_audit.py
import numpy as np

rng = np.random.default_rng(3)
N = 4000

group = np.where(rng.random(N) < 0.5, "A", "B")

# Different base rates on purpose: 40% of A repay, 25% of B repay.
# This is the world the model is scored against, not an opinion about either group.
base = np.where(group == "A", 0.40, 0.25)
y = (rng.random(N) < base).astype(int)              # 1 = repaid the loan

# A single, group-blind scorer. Repayers score higher on average in BOTH groups.
score = np.clip(rng.normal(np.where(y == 1, 0.65, 0.35), 0.15), 0, 1)


def rates(mask, pred):
    yy, pp = y[mask], pred[mask]
    tp = int(((pp == 1) & (yy == 1)).sum()); fp = int(((pp == 1) & (yy == 0)).sum())
    fn = int(((pp == 0) & (yy == 1)).sum()); tn = int(((pp == 0) & (yy == 0)).sum())
    return {
        "selection": (tp + fp) / len(yy),                     # share approved
        "TPR":       tp / (tp + fn) if tp + fn else 0.0,      # of true repayers, share approved
        "FPR":       fp / (fp + tn) if fp + tn else 0.0,      # of true defaulters, share approved
        "PPV":       tp / (tp + fp) if tp + fp else 0.0,      # of approvals, share who repay
    }


def audit(title, thr_a, thr_b):
    thr = np.where(group == "A", thr_a, thr_b)
    pred = (score >= thr).astype(int)
    ra, rb = rates(group == "A", pred), rates(group == "B", pred)
    print(f"\n{title}   (thresholds A={thr_a:.2f}  B={thr_b:.2f})")
    print(f"{'metric':10s}{'group A':>10s}{'group B':>10s}{'gap':>10s}")
    for k in ("selection", "TPR", "FPR", "PPV"):
        print(f"{k:10s}{ra[k]:10.3f}{rb[k]:10.3f}{ra[k] - rb[k]:+10.3f}")


print("base rate A:", round(y[group == "A"].mean(), 3),
      " base rate B:", round(y[group == "B"].mean(), 3))

audit("ONE THRESHOLD FOR EVERYONE", 0.50, 0.50)

# Now force equal selection rates: lower B's bar until the approval shares match.
target = (score[group == "A"] >= 0.50).mean()
thr_b = float(np.quantile(score[group == "B"], 1 - target))
audit("EQUAL SELECTION RATES (demographic parity)", 0.50, thr_b)
Output
base rate A: 0.404  base rate B: 0.257

ONE THRESHOLD FOR EVERYONE   (thresholds A=0.50  B=0.50)
metric       group A   group B       gap
selection      0.438     0.343    +0.094
TPR            0.850     0.837    +0.014
FPR            0.158     0.172    -0.015
PPV            0.785     0.627    +0.158

EQUAL SELECTION RATES (demographic parity)   (thresholds A=0.50  B=0.44)
metric       group A   group B       gap
selection      0.438     0.438    +0.000
TPR            0.850     0.916    -0.066
FPR            0.158     0.272    -0.114
PPV            0.785     0.538    +0.247

This output is the entire lesson

The single threshold nearly satisfies equalised odds. TPR gap is +0.014 and FPR gap is -0.015. The scorer is equally good at ranking within both groups, so error rates line up.

The same single threshold fails demographic parity and predictive parity. Selection differs by 0.094. And PPV differs by 0.158: an approval for group A means a 78.5% chance of repayment, for group B a 62.7% chance. The same decision carries different meaning.

Forcing demographic parity fixes one gap and widens two. The selection gap goes to exactly zero. FPR gap grows from -0.015 to -0.114, and the PPV gap nearly doubles to +0.247.

Read that last line carefully. Equalising approval rates meant approving group B applicants who score lower, so more of them default, so an approval for group B now means less. Nobody made a mistake. The conflict is structural.

Line by line

np.clip(rng.normal(np.where(y == 1, 0.65, 0.35), 0.15), 0, 1) builds a score whose distribution depends only on the true label, not on the group. That is deliberate: it removes "the model is worse at group B" as an explanation, leaving only the base-rate effect.

np.quantile(score[group == "B"], 1 - target) finds the score cut-off that approves exactly the target share of group B. Per-group thresholds are the standard post-processing lever, and in many jurisdictions using them is itself legally fraught — decide with counsel, not alone.

rates() returns all four numbers together on purpose. Reporting one fairness metric in isolation is how teams accidentally trade away a metric nobody was watching.

Common mistakes

Reporting a ratio without the counts. A TPR gap of 0.30 on a group of 25 people is noise. Print n next to every group metric, and bootstrap a confidence interval when the group is under a few hundred.

Comparing rates at different thresholds without saying so. Every one of these numbers moves with the threshold. Fix the threshold, or report the whole curve.

Optimising the metric directly on the test set. Tuning per-group thresholds on your evaluation data overfits fairness the same way it overfits accuracy. Use a separate calibration split — see train, test and validation splits.

Treating a small gap as no gap. Set a threshold in advance, in writing, with a reason. "Selection-rate gap under 0.05, TPR gap under 0.05" is auditable. "Looks fine" is not.

Ignoring who bears the cost. For loans, a false positive hurts the lender and a false negative hurts the applicant. For medical screening the direction flips. The metric you equalise should be the one whose errors land on the person with the least power.

Try it yourself

Set both base rates to 0.40 and rerun. Every gap should collapse toward zero at a single threshold, because the impossibility only bites when base rates differ. Then equalise TPR instead of selection rate by tuning thr_b, and watch which other two gaps open up.

What to learn next

Researcher — Mathematics and papers.

Definitions

Let $A$ be the protected attribute, $Y \in {0,1}$ the true outcome, $R \in [0,1]$ the model score and $\hat{Y} = \mathbb{1}[R \ge t]$ the thresholded decision.

Independence (demographic parity). $\hat{Y} \perp A$, i.e. $P(\hat{Y} = 1 \mid A = a)$ is constant in $a$. Relaxations: the difference $\max_{a,a'} |P(\hat Y=1 \mid a) - P(\hat Y=1 \mid a')|$, or the ratio, which underlies the four-fifths rule of thumb used in US employment screening.

Separation (equalised odds). $\hat{Y} \perp A \mid Y$. Equivalently $P(\hat Y = 1 \mid Y = y, A = a)$ is constant in $a$ for both $y \in {0,1}$ — equal TPR and equal FPR. Equal opportunity (Hardt et al., 2016) is the relaxation to $y = 1$ only.

Sufficiency (predictive parity, calibration). $Y \perp A \mid R$. Equivalently $P(Y = 1 \mid R = r, A = a)$ is constant in $a$: a score of 0.7 means the same thing for everyone.

Barocas, Hardt and Narayanan, Fairness and Machine Learning, organise the entire literature under these three, and it is the cleanest available framing.

The impossibility results

Chouldechova (2017) proves the constraint that the developer block demonstrates. Writing $p$ for the base rate $P(Y=1 \mid A=a)$, $\text{PPV}$ for precision, and $\text{FPR}$, $\text{FNR}$ for the two error rates:

$$ \text{FPR} = \frac{p}{1-p} \cdot \frac{1 - \text{PPV}}{\text{PPV}} \cdot (1 - \text{FNR}) $$

Every quantity is per group. Fix PPV equal across groups and let $p$ differ; then FPR and FNR cannot both be equal. Equal calibration plus unequal prevalence forces unequal error rates. There is no algorithmic escape.

Kleinberg, Mullainathan and Raghavan (2016) prove the companion result for scores rather than decisions: calibration within groups, balance for the positive class, and balance for the negative class are jointly satisfiable only when base rates are equal or the predictor is perfect.

Corbett-Davies and Goel (2018), The Measure and Mismeasure of Fairness, argue from the decision-theoretic side that all three criteria can be satisfied by degrading a group's outcomes, and that the constraints therefore cannot be treated as goals in themselves. The infra-marginality critique matters: comparing rates across groups with different risk distributions can indicate a difference in the distributions rather than in the treatment.

Interventions, by pipeline stage

Pre-processing. Reweighting (Kamiran and Calders, 2012); learning fair representations (Zemel et al., 2013) minimising $I(Z; A)$ while preserving $I(Z; Y)$; disparate-impact removal (Feldman et al., 2015) by per-group quantile alignment of features. Model-agnostic, and weakest in guarantee.

In-processing. Constrained ERM — Agarwal et al. (2018), A Reductions Approach to Fair Classification, reduces fairness-constrained classification to a sequence of cost-sensitive problems solved by any base learner, with provable convergence to the constrained optimum. Adversarial debiasing (Zhang et al., 2018) trains a predictor against an adversary that predicts $A$ from the prediction. Strongest guarantees, requires retraining.

Post-processing. Hardt et al. (2016) derive the optimal equalised-odds predictor as a per-group randomised threshold, solvable by a small linear program over the group ROC curves. The achievable region is the intersection of the group ROC hulls, so the fair optimum is bounded by the worst group's curve. Cheap, needs $A$ at inference time, and often unlawful to use for that reason.

Individual and causal fairness

Group metrics are averages and can be satisfied by a policy that treats similar individuals differently. Dwork et al. (2012) propose Lipschitz individual fairness: $D(f(x), f(x')) \le d(x, x')$ for a task-specific similarity metric $d$. The metric $d$ is the whole difficulty, and constructing it is equivalent to solving the fairness problem.

Counterfactual fairness (Kusner et al., 2017) requires the prediction to be unchanged in the counterfactual world where $A$ is different, under a specified structural causal model. It makes assumptions explicit and testable, and it requires a causal graph you cannot validate from observational data alone.

Reporting protocol

  • Report all four rates per group, with $n$ and bootstrap intervals.
  • Report the intersecting cells, not only marginal groups, with counts.
  • State the threshold and the split the threshold was chosen on.
  • State which criterion the system is committed to and why, before results are available.

Fairlearn and AIF360 implement most of the above. Both are wrappers around the arithmetic here; the choice of criterion remains outside the library.

Papers

What to learn next

What to learn next

These follow on from what you just read.

  • AI Safety and Ethics

    Explainability

    Explainability is the engineering of checkable answers to "why did the model say that" — and different methods answer different questions.

  • AI Safety and Ethics

    SHAP and LIME

    SHAP splits a single prediction into a fair share per feature; LIME fits a small readable model near one point — here both are built from scratch.

  • AI Safety and Ethics

    Privacy in machine learning

    A trained model can leak the people it was trained on — and both the leak and the fix are measurable in a few lines of code.