Fairness metrics
Fairness has several precise definitions that cannot all hold at once — so the job is choosing one on purpose and measuring it.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A fairness metric is a specific, countable definition of "treated equally" — and there are several, which disagree.
The analogy you have already lived
Think of one bathroom and a queue of people. What is the fair rule?
First come, first served treats everyone the same. But the person who has been waiting since morning and the person who wandered up ten seconds ago are not in the same situation. Letting the desperate one in first also feels fair.
Both rules are fair. They are fair in different senses, and you cannot follow both at the same time.
That is the whole of this lesson. Fairness is not one thing waiting to be measured. It is a family of things, and picking one means giving up another.
The three questions people mean by "fair"
Suppose a bank uses a model to approve loans. Three different people complain, and all three are right.
"You approve fewer of us." This is about the share approved in each group. Making those shares equal is called demographic parity — equal approval rates, regardless of anything else.
"You reject people from our group who would have repaid." This is about who gets missed. Making the catch-rate equal is called equal opportunity — of the people who would repay, the same fraction gets approved in each group.
"Your approvals mean something different for us." This is about what a yes is worth. Making that equal is called predictive parity — among approved people, the same fraction actually repays.
Each is reasonable. Each is measurable. And here is the hard part.
They cannot all be true at once
Suppose the two groups repay at different underlying rates. That can happen for any reason, including unfair ones inherited from the past. When it does, no model can satisfy all three at once.
This is not a limit of today's technology. It is arithmetic, proved in 2016, and it will still be true in fifty years.
So there is no fair model, only a model that is fair in a way you chose and can defend. Anyone who says their system is "unbiased" without naming a metric has not measured anything.
The picture
different base rates in the two groups
│
┌───────────────┼───────────────┐
▼ ▼ ▼
equal approval equal catch-rate equal meaning
rates for the of a "yes"
deserving
│ │ │
└──────► pick ONE, on purpose ◄─┘
then measure it
then publish itWhere you have already seen the fight
- Exam cut-offs and reservation policies argue over exactly these three definitions.
- Insurance pricing by area is demographic parity versus accuracy, in public.
- Fraud checks that stop more cards from one region are an equal-catch-rate argument.
What is honestly hard here
Choosing the metric is not a technical decision, and engineers should not make it alone.
The right people to choose are the ones accountable for the outcome, informed by someone who understands the trade. Your job is to make the trade visible: here is what we give up, in numbers, if we pick this one.
That is a much more useful contribution than an opinion.
Remember this
- Fairness has several precise definitions, and they conflict by arithmetic.
- With different base rates, no model satisfies all of them at once.
- Choose one deliberately, measure it, and write down what you gave up.
What to learn next
- Explainability — once a gap is found, working out what drives it.
- Model cards and documentation — where these numbers get published.
- Model evaluation — the base rates and confusion matrices behind all of it.
Developer — Code and libraries.
Four numbers, one table
Every group-fairness metric is a comparison of one rate across groups. Compute all four rates per group and you can read off any definition you need.
| Rate | Meaning | Metric that equalises it |
|---|---|---|
| selection rate | share of the group that gets a yes | demographic parity |
| TPR | of those who truly qualify, share approved | equal opportunity |
| FPR | of those who truly do not, share approved | (with TPR) equalised odds |
| PPV | of those approved, share who truly qualify | predictive parity |
Setup
pip install numpyThe impossibility, run on your own machine
The two groups below have different repayment base rates by construction. The scorer is group-blind and equally skilled on both groups.
import numpy as np
rng = np.random.default_rng(3)
N = 4000
group = np.where(rng.random(N) < 0.5, "A", "B")
# Different base rates on purpose: 40% of A repay, 25% of B repay.
# This is the world the model is scored against, not an opinion about either group.
base = np.where(group == "A", 0.40, 0.25)
y = (rng.random(N) < base).astype(int) # 1 = repaid the loan
# A single, group-blind scorer. Repayers score higher on average in BOTH groups.
score = np.clip(rng.normal(np.where(y == 1, 0.65, 0.35), 0.15), 0, 1)
def rates(mask, pred):
yy, pp = y[mask], pred[mask]
tp = int(((pp == 1) & (yy == 1)).sum()); fp = int(((pp == 1) & (yy == 0)).sum())
fn = int(((pp == 0) & (yy == 1)).sum()); tn = int(((pp == 0) & (yy == 0)).sum())
return {
"selection": (tp + fp) / len(yy), # share approved
"TPR": tp / (tp + fn) if tp + fn else 0.0, # of true repayers, share approved
"FPR": fp / (fp + tn) if fp + tn else 0.0, # of true defaulters, share approved
"PPV": tp / (tp + fp) if tp + fp else 0.0, # of approvals, share who repay
}
def audit(title, thr_a, thr_b):
thr = np.where(group == "A", thr_a, thr_b)
pred = (score >= thr).astype(int)
ra, rb = rates(group == "A", pred), rates(group == "B", pred)
print(f"\n{title} (thresholds A={thr_a:.2f} B={thr_b:.2f})")
print(f"{'metric':10s}{'group A':>10s}{'group B':>10s}{'gap':>10s}")
for k in ("selection", "TPR", "FPR", "PPV"):
print(f"{k:10s}{ra[k]:10.3f}{rb[k]:10.3f}{ra[k] - rb[k]:+10.3f}")
print("base rate A:", round(y[group == "A"].mean(), 3),
" base rate B:", round(y[group == "B"].mean(), 3))
audit("ONE THRESHOLD FOR EVERYONE", 0.50, 0.50)
# Now force equal selection rates: lower B's bar until the approval shares match.
target = (score[group == "A"] >= 0.50).mean()
thr_b = float(np.quantile(score[group == "B"], 1 - target))
audit("EQUAL SELECTION RATES (demographic parity)", 0.50, thr_b)base rate A: 0.404 base rate B: 0.257 ONE THRESHOLD FOR EVERYONE (thresholds A=0.50 B=0.50) metric group A group B gap selection 0.438 0.343 +0.094 TPR 0.850 0.837 +0.014 FPR 0.158 0.172 -0.015 PPV 0.785 0.627 +0.158 EQUAL SELECTION RATES (demographic parity) (thresholds A=0.50 B=0.44) metric group A group B gap selection 0.438 0.438 +0.000 TPR 0.850 0.916 -0.066 FPR 0.158 0.272 -0.114 PPV 0.785 0.538 +0.247
This output is the entire lesson
The single threshold nearly satisfies equalised odds. TPR gap is +0.014 and FPR gap is -0.015. The scorer is equally good at ranking within both groups, so error rates line up.
The same single threshold fails demographic parity and predictive parity. Selection differs by 0.094. And PPV differs by 0.158: an approval for group A means a 78.5% chance of repayment, for group B a 62.7% chance. The same decision carries different meaning.
Forcing demographic parity fixes one gap and widens two. The selection gap goes to exactly zero. FPR gap grows from -0.015 to -0.114, and the PPV gap nearly doubles to +0.247.
Read that last line carefully. Equalising approval rates meant approving group B applicants who score lower, so more of them default, so an approval for group B now means less. Nobody made a mistake. The conflict is structural.
Line by line
np.clip(rng.normal(np.where(y == 1, 0.65, 0.35), 0.15), 0, 1) builds a score whose distribution depends only on the true label, not on the group. That is deliberate: it removes "the model is worse at group B" as an explanation, leaving only the base-rate effect.
np.quantile(score[group == "B"], 1 - target) finds the score cut-off that approves exactly the target share of group B. Per-group thresholds are the standard post-processing lever, and in many jurisdictions using them is itself legally fraught — decide with counsel, not alone.
rates() returns all four numbers together on purpose. Reporting one fairness metric in isolation is how teams accidentally trade away a metric nobody was watching.
Common mistakes
Reporting a ratio without the counts. A TPR gap of 0.30 on a group of 25 people is noise. Print n next to every group metric, and bootstrap a confidence interval when the group is under a few hundred.
Comparing rates at different thresholds without saying so. Every one of these numbers moves with the threshold. Fix the threshold, or report the whole curve.
Optimising the metric directly on the test set. Tuning per-group thresholds on your evaluation data overfits fairness the same way it overfits accuracy. Use a separate calibration split — see train, test and validation splits.
Treating a small gap as no gap. Set a threshold in advance, in writing, with a reason. "Selection-rate gap under 0.05, TPR gap under 0.05" is auditable. "Looks fine" is not.
Ignoring who bears the cost. For loans, a false positive hurts the lender and a false negative hurts the applicant. For medical screening the direction flips. The metric you equalise should be the one whose errors land on the person with the least power.
Try it yourself
Set both base rates to 0.40 and rerun. Every gap should collapse toward zero at a single threshold, because the impossibility only bites when base rates differ. Then equalise TPR instead of selection rate by tuning thr_b, and watch which other two gaps open up.
What to learn next
- Explainability — once a gap is found, working out what drives it.
- Model cards and documentation — where these numbers get published.
- Model evaluation — the base rates and confusion matrices behind all of it.
Researcher — Mathematics and papers.
Definitions
Let $A$ be the protected attribute, $Y \in {0,1}$ the true outcome, $R \in [0,1]$ the model score and $\hat{Y} = \mathbb{1}[R \ge t]$ the thresholded decision.
Independence (demographic parity). $\hat{Y} \perp A$, i.e. $P(\hat{Y} = 1 \mid A = a)$ is constant in $a$. Relaxations: the difference $\max_{a,a'} |P(\hat Y=1 \mid a) - P(\hat Y=1 \mid a')|$, or the ratio, which underlies the four-fifths rule of thumb used in US employment screening.
Separation (equalised odds). $\hat{Y} \perp A \mid Y$. Equivalently $P(\hat Y = 1 \mid Y = y, A = a)$ is constant in $a$ for both $y \in {0,1}$ — equal TPR and equal FPR. Equal opportunity (Hardt et al., 2016) is the relaxation to $y = 1$ only.
Sufficiency (predictive parity, calibration). $Y \perp A \mid R$. Equivalently $P(Y = 1 \mid R = r, A = a)$ is constant in $a$: a score of 0.7 means the same thing for everyone.
Barocas, Hardt and Narayanan, Fairness and Machine Learning, organise the entire literature under these three, and it is the cleanest available framing.
The impossibility results
Chouldechova (2017) proves the constraint that the developer block demonstrates. Writing $p$ for the base rate $P(Y=1 \mid A=a)$, $\text{PPV}$ for precision, and $\text{FPR}$, $\text{FNR}$ for the two error rates:
$$ \text{FPR} = \frac{p}{1-p} \cdot \frac{1 - \text{PPV}}{\text{PPV}} \cdot (1 - \text{FNR}) $$
Every quantity is per group. Fix PPV equal across groups and let $p$ differ; then FPR and FNR cannot both be equal. Equal calibration plus unequal prevalence forces unequal error rates. There is no algorithmic escape.
Kleinberg, Mullainathan and Raghavan (2016) prove the companion result for scores rather than decisions: calibration within groups, balance for the positive class, and balance for the negative class are jointly satisfiable only when base rates are equal or the predictor is perfect.
Corbett-Davies and Goel (2018), The Measure and Mismeasure of Fairness, argue from the decision-theoretic side that all three criteria can be satisfied by degrading a group's outcomes, and that the constraints therefore cannot be treated as goals in themselves. The infra-marginality critique matters: comparing rates across groups with different risk distributions can indicate a difference in the distributions rather than in the treatment.
Interventions, by pipeline stage
Pre-processing. Reweighting (Kamiran and Calders, 2012); learning fair representations (Zemel et al., 2013) minimising $I(Z; A)$ while preserving $I(Z; Y)$; disparate-impact removal (Feldman et al., 2015) by per-group quantile alignment of features. Model-agnostic, and weakest in guarantee.
In-processing. Constrained ERM — Agarwal et al. (2018), A Reductions Approach to Fair Classification, reduces fairness-constrained classification to a sequence of cost-sensitive problems solved by any base learner, with provable convergence to the constrained optimum. Adversarial debiasing (Zhang et al., 2018) trains a predictor against an adversary that predicts $A$ from the prediction. Strongest guarantees, requires retraining.
Post-processing. Hardt et al. (2016) derive the optimal equalised-odds predictor as a per-group randomised threshold, solvable by a small linear program over the group ROC curves. The achievable region is the intersection of the group ROC hulls, so the fair optimum is bounded by the worst group's curve. Cheap, needs $A$ at inference time, and often unlawful to use for that reason.
Individual and causal fairness
Group metrics are averages and can be satisfied by a policy that treats similar individuals differently. Dwork et al. (2012) propose Lipschitz individual fairness: $D(f(x), f(x')) \le d(x, x')$ for a task-specific similarity metric $d$. The metric $d$ is the whole difficulty, and constructing it is equivalent to solving the fairness problem.
Counterfactual fairness (Kusner et al., 2017) requires the prediction to be unchanged in the counterfactual world where $A$ is different, under a specified structural causal model. It makes assumptions explicit and testable, and it requires a causal graph you cannot validate from observational data alone.
Reporting protocol
- Report all four rates per group, with $n$ and bootstrap intervals.
- Report the intersecting cells, not only marginal groups, with counts.
- State the threshold and the split the threshold was chosen on.
- State which criterion the system is committed to and why, before results are available.
Fairlearn and AIF360 implement most of the above. Both are wrappers around the arithmetic here; the choice of criterion remains outside the library.
Papers
- Hardt, Price and Srebro, Equality of Opportunity in Supervised Learning, 2016 — arxiv.org/abs/1610.02413
- Chouldechova, Fair Prediction with Disparate Impact, 2017 — arxiv.org/abs/1703.00056
- Kleinberg, Mullainathan and Raghavan, Inherent Trade-Offs in the Fair Determination of Risk Scores, 2016 — arxiv.org/abs/1609.05807
- Agarwal et al., A Reductions Approach to Fair Classification, 2018 — arxiv.org/abs/1803.02453
- Corbett-Davies and Goel, The Measure and Mismeasure of Fairness, 2018 — arxiv.org/abs/1808.00023
- Kusner et al., Counterfactual Fairness, 2017 — arxiv.org/abs/1703.06856
What to learn next
- Explainability — once a gap is found, working out what drives it.
- Model cards and documentation — where these numbers get published.
- Model evaluation — the base rates and confusion matrices behind all of it.