Bias and consent in face recognition
One threshold produces different error rates for different groups of people, and in most of the world a face template is legally regulated biometric data — here are the measured numbers and the actual rules.
- 18 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A face system's error rate is not one number. It is a different number for each group of people, and the gap between them can be enormous.
Think of a door with a single step in front of it. The step is one height for everybody. For most people it is nothing. For a person on crutches it is a real obstacle, and for a wheelchair user it is a wall.
The step did not change. Who it stops changed. A face recognition threshold works the same way: one setting, wildly different consequences depending on who walks up.
What has actually been measured
This is not speculation. Government laboratories have measured it at scale.
The United States National Institute of Standards and Technology ran the largest study. It covered one hundred and eighty-nine algorithms from ninety-nine developers. More than eighteen million photographs went through them. The results were published in December two thousand and nineteen.
The rate of wrongly matching two different people varied between groups. The gap ran from a factor of ten to a factor of one hundred. A follow-up report in two thousand and twenty-two found the spread far wider still. It dwarfed the spread in the other kind of error.
Two findings from that work are worth holding onto.
Wrong matches vary far more than missed matches. A system might miss your face a little more often. It might wrongly match you to a stranger a hundred times more often. The second error is the one that leads to an accusation.
More accurate systems have smaller gaps. The differences were largest in weaker algorithms. Accuracy and fairness were not in opposition here; poor engineering produced both problems together.
There was one more result worth stating. Algorithms developed in different countries showed different patterns, tracking the populations in their training data. That points at the cause, and at the fix.
Why it happens
The training photos were not balanced. Large face datasets were assembled from the public internet, which is not a representative sample of humanity.
Cameras and their settings were tuned narrowly. Automatic exposure that meters for a mid-tone skin gives less usable detail on very dark or very light skin. The photograph is worse before any model sees it.
One threshold is applied to everyone. Even with equally good photographs, the score patterns differ between groups. One cut-off line lands in a different place on each.
Where the consequences land
In one-to-one checking, a wrong match means somebody unlocks something they should not. Bad, and bounded.
In one-to-many searching, a wrong match means a name and address are handed to whoever asked. By two thousand and twenty-six, more than a dozen wrongful arrests in the United States had been documented. Each followed a face recognition lead. A settlement with the Detroit police department in two thousand and twenty-four changed their policy. An arrest can no longer rest on a face recognition result alone.
The pattern in those documented cases is the same pattern the laboratory measurements predicted.
The rules, stated plainly
These are the legal facts, not opinions about them. If you build a face system, they apply to you.
In the European Union. A face template used to identify someone is special category data under the General Data Protection Regulation. Processing it is prohibited unless a specific exception applies. Explicit consent is the usual one.
Separately, the EU Artificial Intelligence Act bans four uses outright:
- Building or expanding a face recognition database by untargeted scraping of images.
- Inferring emotions in workplaces and schools.
- Using biometrics to infer race, political opinion or sexual orientation.
- Real-time remote identification in public places for law enforcement, with narrow exceptions.
Those bans have applied since February two thousand and twenty-five.
In Illinois, in the United States. The Biometric Information Privacy Act requires written, informed consent before you collect a face template. It also requires a published retention and destruction schedule, and bans selling the data. Individuals can sue directly, which is why the largest settlements come from there.
In India. The Digital Personal Data Protection Act passed in two thousand and twenty-three. Its rules were notified in November two thousand and twenty-five. It requires clear, informed and revocable consent. Use is limited to the purpose you stated when you collected the data. Organisations have until roughly the middle of two thousand and twenty-seven to comply.
Notice what is common to all three. Consent must be asked for before collection, not after. Purpose must be stated and stuck to. Deletion must be possible.
Where you have already seen this
- The consent screen before a bank app photographs your face.
- Airports offering a face-scan lane and a manual lane side by side.
- Retailers withdrawing face-based systems after regulators asked questions.
- Photo apps letting you switch face grouping off, and delete what was built.
Remember this
- Error rates differ by group, and wrong-match rates differ far more than missed-match rates.
- Measure per group, at your threshold, on your population. A single number hides the problem.
- A face template is regulated biometric data in most of the world. Consent comes first.
What to learn next
- Bias in datasets — where the imbalance enters, and what it costs downstream.
- Fairness metrics — the formal definitions, and the proofs that some cannot hold together.
- Model cards — the document that records the measured limits.
Developer — Code and libraries.
Setup
pip install numpyRun against numpy 1.26. The code shows the mechanism directly: four groups, one model, one threshold, and the arithmetic that follows.
What one threshold does to four groups
import numpy as np
rng = np.random.default_rng(42)
M = 2_000_000
# One model, one threshold, four groups. Only the score distributions differ,
# and they differ because the training data covered the groups unevenly.
groups = { # (genuine mean, sd, impostor mean, sd)
"group A": (0.620, 0.100, 0.040, 0.055),
"group B": (0.600, 0.110, 0.075, 0.062),
"group C": (0.575, 0.120, 0.095, 0.068),
"group D": (0.560, 0.125, 0.110, 0.072),
}
scores = {g: (np.clip(rng.normal(gm, gs, 200_000), -1, 1),
np.clip(rng.normal(im, isd, M), -1, 1))
for g, (gm, gs, im, isd) in groups.items()}
def matches(g, t): return int((scores[g][1] >= t).sum())
def fmr(g, t): return matches(g, t) / M
def fnmr(g, t): return float((scores[g][0] < t).mean())
pooled = np.sort(np.concatenate([s[1] for s in scores.values()]))
t = float(pooled[int(len(pooled) * (1 - 1e-4))])
print(f"one global threshold, chosen for a POOLED false match rate of 1 in 10,000: {t:.4f}\n")
print(f"{'group':>8} {'false matches':>16} {'FMR':>11} {'FNMR':>8}")
for g in scores:
print(f"{g:>8} {matches(g, t):>7,} / {M:,} {fmr(g, t):>11.6f} {fnmr(g, t):>8.4f}")
fs = {g: fmr(g, t) for g in scores}
lo, hi = min(fs.values()), max(fs.values())
print(f"\nworst group FMR is at least {hi/max(lo, 1/M):.0f}x the best (the best group measured zero)")
print(f"worst group FNMR is {max(fnmr(g,t) for g in scores)/min(fnmr(g,t) for g in scores):.0f}x the best")
print(f"\nper-group thresholds, each tuned to FMR = 1e-4 -- equal security, unequal convenience:")
print(f"{'group':>8} {'threshold':>10} {'FMR':>11} {'FNMR':>8}")
for g in scores:
tg = float(np.sort(scores[g][1])[int(M * (1 - 1e-4))])
print(f"{g:>8} {tg:>10.4f} {fmr(g, tg):>11.6f} {fnmr(g, tg):>8.4f}")
N = 50_000
print(f"\nexpected false matches per search of a {N:,}-person watchlist, global threshold:")
for g in scores:
print(f"{g:>8}: {N * fs[g]:>7.2f}")one global threshold, chosen for a POOLED false match rate of 1 in 10,000: 0.3547 group false matches FMR FNMR group A 0 / 2,000,000 0.000000 0.0040 group B 7 / 2,000,000 0.000003 0.0131 group C 136 / 2,000,000 0.000068 0.0333 group D 657 / 2,000,000 0.000329 0.0491 worst group FMR is at least 657x the best (the best group measured zero) worst group FNMR is 12x the best per-group thresholds, each tuned to FMR = 1e-4 -- equal security, unequal convenience: group threshold FMR FNMR group A 0.2449 0.000100 0.0001 group B 0.3064 0.000100 0.0039 group C 0.3477 0.000100 0.0293 group D 0.3775 0.000100 0.0713 expected false matches per search of a 50,000-person watchlist, global threshold: group A: 0.00 group B: 0.17 group C: 3.40 group D: 16.43
Reading it
The pooled figure is 1 in 10,000 and no group experiences that. Group A gets far better; group D gets more than three times worse. A pooled metric describes a population average that nobody actually lives in. This is the single most important line in the lesson.
The false-match spread is 657x; the false-non-match spread is 12x. That ratio is not an artefact of the simulation. It is the shape NIST reports: false positive differentials are far wider than false negative ones, because the impostor distribution's tail is what the threshold cuts, and tails diverge faster than means.
The 657x is a lower bound. Group A recorded zero false matches in two million comparisons, so its true rate is somewhere below $5 \times 10^{-7}$ and unmeasured. To report a real ratio you need enough comparisons for the best group to register a count. Report the count alongside every rate.
Per-group thresholds equalise one error and unbalance the other. Every group now sits at exactly $10^{-4}$ false matches — and group D's rejection rate climbs from 4.9% to 7.1%, while group A's falls to 0.01%. You have moved the burden, not removed it. Choosing which error to equalise is a policy decision, and in some jurisdictions treating people differently on the basis of a group attribute raises its own legal problems. There is no setting of a threshold that makes this go away.
The watchlist line is where it stops being abstract. At the same global threshold, a search returns on average 16.43 false candidates for a group D subject and effectively none for group A. If a human reviews the top candidates and a name leaves the building, that number is the rate at which the wrong person gets named.
What to do about it
Measure per group, always. Report FMR and FNMR per group at your operating threshold, with the raw counts. ISO/IEC 19795-10 is the standard developing for exactly this reporting.
Build an evaluation set that reflects your users. Not a public benchmark. Your cameras, your lighting, your population. This is often the highest-value week of work in the entire project.
Fix the images before the model. Under-exposed faces are the cause of a large slice of the gap. Camera placement, lighting and exposure metering are unglamorous and effective.
Raise the threshold and add a human step for 1:N. If you cannot equalise, reduce total exposure. The Detroit policy change — a face match cannot on its own support an arrest — is the operational version of this.
Write down the limits. A model card stating measured per-group rates, tested population and known failure modes is both good practice and, for high-risk systems under the EU AI Act, part of the documentation obligation.
Common mistakes
Reporting accuracy on a benchmark and calling it validated. LFW is saturated and unrepresentative. It cannot detect a demographic gap.
Balancing the training set and declaring the problem solved. Balanced training helps and does not close the gap, because image quality and threshold effects operate independently of class balance.
Collecting demographic labels without a legal basis. You need per-group data to measure fairness, and in the EU the attributes you need are themselves special category data. Plan this with counsel before collecting; do not improvise it.
Assuming consent obtained for one purpose covers another. Under GDPR, the DPDP Act and BIPA alike, purpose limitation is explicit. Face data collected for unlocking a phone does not license a watchlist.
Retaining templates indefinitely. BIPA requires a published destruction schedule. GDPR requires storage limitation. Indefinite retention is the default in code and unlawful in several jurisdictions.
Try it yourself
Add a fifth group whose genuine and impostor distributions match group A exactly, but whose images are 30% noisier — model this by widening both standard deviations by a third. Recompute everything. You will see that image quality alone, with no difference in the underlying face data, produces a demographic-looking gap. That is why camera and lighting audits belong in a fairness review.
What to learn next
- Bias in datasets — where the imbalance enters, and what it costs downstream.
- Fairness metrics — the formal definitions, and the proofs that some cannot hold together.
- Model cards — the document that records the measured limits.
Researcher — Mathematics and papers.
What the measurements say
NISTIR 8280 (Grother, Ngan, Hanaoka, December 2019), FRVT Part 3: Demographic Effects, evaluated 189 algorithms from 99 developers over 18.27 million images from four US government datasets. Headline results:
- False positive differentials across demographic groups typically spanned a factor of 10 to 100 in 1:1 verification, algorithm-dependent.
- Among US-developed algorithms, elevated false positives were observed for Asian, African American and native groups relative to Eastern European faces, with American Indian faces the highest.
- In 1:N identification against a large mugshot gallery, elevated false positives for African American females were of particular concern given the operational consequence.
- Algorithms developed in Asian countries did not show the Asian/Caucasian false-positive gap present in US-developed algorithms — evidence that the effect tracks training data composition rather than anything intrinsic to faces.
- Higher-accuracy algorithms exhibited smaller differentials.
NISTIR 8429 (Grother, July 2022), FRVT Part 8: Summarizing Demographic Differentials, formalises measurement. Its central quantitative finding: false positive differentials are widespread, present even in pristine images, and vastly exceed false negative differentials — within-group false positive rates varying by up to a factor of 7203, against roughly a factor of 3 for false negatives. The measures proposed there feed into ISO/IEC 19795-10, the standard for measuring and reporting demographic effects in biometric systems.
Buolamwini and Gebru (FAT* 2018), Gender Shades, is frequently cited in this context and measures something different: commercial gender classification, where error rates reached 34.7% for darker-skinned women against 0.8% for lighter-skinned men. It is a study of attribute classification, not identification. Citing it as evidence about recognition error rates is a common and avoidable mistake; cite NIST for recognition.
Why false positives diverge more than false negatives
Let genuine and impostor scores for group $g$ have distributions $G_g$ and $I_g$. At threshold $\tau$:
$$ \text{FNMR}_g(\tau) = G_g(\tau), \qquad \text{FMR}_g(\tau) = 1 - I_g(\tau) $$
FNMR is evaluated near the body of the genuine distribution, where modest shifts in mean produce modest changes. FMR is evaluated in the far tail of the impostor distribution, where a small change in mean or variance changes the tail mass by orders of magnitude. For Gaussian tails the sensitivity is exponential in the standardised threshold.
This is a structural explanation, and it predicts what NIST observes: differentials of order $10^3$ in FMR alongside order $10^0$ in FNMR. It also implies that fairness interventions targeting mean score separation will improve FNMR parity far more than FMR parity.
Contributing causes, separable in principle
- Training data composition. MS-Celeb-1M, VGGFace2 and successors are internet-scraped and demographically skewed. The country-of-origin effect in NISTIR 8280 is the strongest available evidence for this channel.
- Image acquisition. Automatic exposure and white balance optimised for mid-tone reflectance reduce effective dynamic range on very dark and very light skin. Cook et al. (IEEE T-BIOM 2019) measured this directly in an operational border-crossing setting: image quality metrics varied by skin reflectance, and controlling for quality reduced but did not eliminate the differential.
- Threshold effects. Even with identical per-group ROC curves, a single global threshold lands at different per-group operating points whenever the score distributions differ in location or scale.
- Evaluation population. Differentials measured on mugshot data do not transfer to visa or webcam imagery. Report the corpus.
Grother's framing is worth adopting: these are separable effects and should be reported separately, because they have different mitigations.
Mitigation, with honest limits
| Approach | Effect | Cost |
|---|---|---|
| Balanced / augmented training data | Reduces, does not eliminate | Data collection cost and its own consent problem |
| Per-group thresholds | Equalises one error rate exactly | Requires group inference at runtime; legally fraught |
| Score normalisation (cohort-based) | Reduces FMR spread without explicit labels | Needs a representative cohort |
| Image quality gating | Attacks the acquisition channel directly | Rejects more samples; can itself be unequal |
| Higher-accuracy backbone | Reduces both error types and the spread | Compute; the effect is real but not a cure |
Per-group thresholding deserves care. It requires classifying a subject into a demographic group at inference time — itself a biometric categorisation, which under Article 5(1)(g) of the EU AI Act is prohibited when used to infer protected attributes such as race. A technically effective mitigation can be an unlawful one.
Legal position, as of August 2026
EU — GDPR. Article 9(1) classifies biometric data processed for the purpose of uniquely identifying a natural person as special category. Processing is prohibited absent an Article 9(2) condition; explicit consent under 9(2)(a) is the usual commercial route. Note the purpose qualifier: a face image is not automatically Article 9 data, but a template used for identification is. Article 22 constraints on solely automated decisions with legal or similarly significant effects apply on top.
EU — AI Act (Regulation (EU) 2024/1689). Article 5 prohibitions, applicable from 2 February 2025:
- 5(1)(e) — creating or expanding facial recognition databases through untargeted scraping of facial images from the internet or CCTV footage.
- 5(1)(f) — inferring emotions in workplace and education settings, with medical and safety exceptions.
- 5(1)(g) — biometric categorisation to deduce race, political opinions, trade union membership, religious or philosophical beliefs, sex life or sexual orientation.
- 5(1)(h) — real-time remote biometric identification in publicly accessible spaces for law enforcement, subject to narrowly drawn exceptions with prior judicial or administrative authorisation.
Penalties for prohibited practices reach EUR 35 million or 7% of worldwide annual turnover.
Remote biometric identification systems that are not prohibited fall under Annex III as high-risk. Regulation (EU) 2026/1744, the AI Digital Omnibus, published in the Official Journal on 24 July 2026 and in force from 27 July 2026, deferred Annex III high-risk obligations from 2 August 2026 to 2 December 2027, and Annex I product-embedded high-risk obligations to 2 August 2028. It did not change the Article 5 prohibitions, and it added a prohibition on systems generating non-consensual intimate imagery and child sexual abuse material with a transition period to 2 December 2026. Article 50 transparency obligations remain applicable from 2 August 2026. Verify the current consolidated text before relying on any date; this area is moving.
United States. No federal biometric statute. Illinois BIPA (740 ILCS 14) requires informed written consent prior to collection, a published retention and destruction schedule, and prohibits sale; it carries a private right of action, which is why it generates the largest exposure. A 2024 amendment (SB 2979) converted per-scan damages into per-person damages, and the Seventh Circuit held in 2026 that the amendment applies retroactively to pending cases. Texas CUBI and Washington HB 1493 impose similar duties with attorney-general-only enforcement. City and state restrictions on government use vary widely and change frequently.
India. The Digital Personal Data Protection Act 2023, with the DPDP Rules notified on 14 November 2025, does not create a separate special category for biometrics. Its consent, notice, purpose-limitation, breach-notification and data-principal-rights obligations apply to face data as personal data, with a transition period running to approximately May 2027.
Operational consequence in policing. More than a dozen wrongful arrests in the United States have been publicly documented following face recognition leads. A 2024 settlement with the Detroit Police Department, arising from a 2020 wrongful arrest, produced policy changes including a prohibition on arrest or lineup based on a face recognition result alone without independent corroborating evidence, mandatory training on differential error rates, and an audit of prior cases.
Documentation
Model cards (Mitchell et al., FAT* 2019) and datasheets for datasets (Gebru et al., 2018) are the accepted vehicles for recording measured per-group performance, evaluation population, and known limits. For systems that fall under the AI Act's high-risk regime, equivalent content becomes a documentation obligation rather than a courtesy.
Papers and reports
- Buolamwini and Gebru, Gender Shades, FAT* 2018 — proceedings.mlr.press/v81/buolamwini18a.html
- Gebru et al., Datasheets for Datasets, 2018 — arxiv.org/abs/1803.09010
- Mitchell et al., Model Cards for Model Reporting, FAT* 2019 — arxiv.org/abs/1810.03993
- Grother, Ngan and Hanaoka, FRVT Part 3: Demographic Effects, NISTIR 8280, 2019 — nvlpubs.nist.gov/nistpubs/ir/2019/nist.ir.8280.pdf
- Cook et al., Demographic Effects in Facial Recognition and their Dependence on Image Acquisition, IEEE T-BIOM 2019
- Grother, FRVT Part 8: Summarizing Demographic Differentials, NISTIR 8429, 2022 — pages.nist.gov/frvt/reports/demographics/nistir_8429.pdf
What to learn next
- Bias in datasets — where the imbalance enters, and what it costs downstream.
- Fairness metrics — the formal definitions, and the proofs that some cannot hold together.
- Model cards — the document that records the measured limits.