Build a face recognition system
Find a face in a photo with OpenCV, then identify whose it is with eigenfaces and an SVM — and see why the system confidently names strangers.
- 23 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Face recognition is two separate jobs: finding a face in a picture, and then working out whose face it is.
The analogy you have already lived
You are standing on a crowded railway platform waiting for your cousin. Hundreds of people walk past. Then you spot her, from behind, thirty metres away, in bad light.
You did not measure anybody's nose. Something in your head holds a compressed impression of her face, and it matched.
Notice that you did two things without noticing. First your eyes picked out faces from the crowd — that is detection. Then one of those faces matched a person you know — that is recognition. Machines have to do both, and they are genuinely different problems.
Why this project exists
A phone that unlocks when you look at it. An airport gate that opens without a boarding pass. A photo app that groups every picture of your grandmother together.
Before this, you typed a password or showed a card. Your face is always with you and cannot be forgotten, which is convenient.
That same fact is the problem. You can change a stolen password. You cannot change your face.
This project can help people and it can harm them. That is not a footnote at the end of this lesson. It runs through the middle of it.
What you are actually going to build
You will start with four hundred small photographs — ten photos each of forty different people. Every photo is labelled with who it shows.
Your program will squeeze each face from thousands of numbers down to about a hundred. Then it learns which patterns belong to which person.
You will also draw a green box around a face in a picture, using a face-finder that has been in OpenCV since 2001.
How it works
a photograph
|
v
[ find the faces ] <- detection: WHERE is a face?
|
v
[ squeeze each face into a short list of numbers ]
|
v
[ compare against the people already in the system ]
|
v
"person 12" (86% sure) or "not in the system"That last box is the whole ballgame, and most beginner tutorials leave it out.
Where you have already seen this
- Unlocking your phone by looking at it.
- Google Photos grouping every picture of one person together.
- Airport gates in several countries that match your face to your passport.
- Attendance systems in offices and colleges.
The honest part — read this twice
Your model will never say "I do not know."
You will train it on thirty people. Then you will show it a hundred photos of ten completely different people it has never seen. It will name every single one of them as somebody from the thirty. Not once will it refuse.
That is not a bug you introduced. A classifier is built to pick the best option from a fixed list. "Nobody" is not on the list, so it cannot be chosen. Adding a "not in the system" answer is extra work you must do yourself, and you will do it in this lesson.
Now consider what that means outside a lesson. A police system that must name a suspect will name somebody. There are documented cases of people arrested after a face system matched them wrongly.
And accuracy is not the same for everybody. In 2018, a landmark study found commercial face systems that were nearly perfect on light-skinned men and wrong roughly a third of the time on dark-skinned women. A national testing body later confirmed the same pattern across a large number of systems.
The reason is unglamorous. The photos used to build these systems were not balanced, so the systems learned some faces better than others. If you build face recognition for anything that affects a person's life, testing your accuracy separately for different groups is not an optional extra.
One more: your face is data about you that you cannot revoke. Ask whether people whose faces you are storing agreed to it, and whether they could refuse without losing something they needed.
Remember this
- Face recognition is two jobs: finding a face, then identifying it.
- A face becomes a short list of numbers, and identification means comparing those numbers.
- The system always names somebody unless you build a "stranger" answer yourself — and it is not equally accurate for everyone.
What to learn next
- OpenCV — the image handling this project rests on.
- Convolutional neural networks — the architecture behind every modern face embedding.
- Image classification — the general vision task, without the identity questions.
Developer — Code and libraries.
The problem, stated precisely
Two tasks that people constantly confuse:
- Detection — given an image, return boxes around any faces. No identity involved.
- Recognition — given a cropped face, return which enrolled person it is.
A third one matters in production: verification, which asks "is this the same person as this other photo?" and answers yes or no. Phone unlock is verification, not recognition.
We build detection with a Viola-Jones cascade from OpenCV, and recognition with eigenfaces — principal component analysis followed by a support vector machine. This is the 1991 method, and it is the right thing to build first because every part of it is inspectable.
Setup
pip install scikit-learn opencv-python numpyThe face-finder XML ships inside opencv-python. The face dataset is a one-time 1.4 MB download, cached afterwards in your home directory. Everything runs on CPU in a few seconds.
The dataset
fetch_olivetti_faces() gives 400 grayscale images at 64x64 pixels: ten photos each of forty people, taken at AT&T Laboratories Cambridge between 1992 and 1994. Faces are already cropped, centred and upright, with varying expression and slight pose changes.
Being honest about what that means: this dataset is small, old, and taken under controlled lighting with a subject pool that is not demographically representative. High accuracy here does not predict accuracy on real photographs, and it especially does not predict equal accuracy across skin tones. It is a teaching dataset. Treat every number below as a demonstration of the method, not evidence the method works.
Part one — recognition
import numpy as np
from sklearn.datasets import fetch_olivetti_faces
from sklearn.model_selection import train_test_split
from sklearn.decomposition import PCA
from sklearn.svm import SVC
from sklearn.pipeline import make_pipeline
from sklearn.metrics import accuracy_score
faces = fetch_olivetti_faces() # about 1.4 MB, downloaded once then cached
X, y = faces.data, faces.target
print("images:", faces.images.shape)
print("people:", len(np.unique(y)), " photos of each:", int((y == 0).sum()))
print("one face as numbers:", X.shape[1], "pixels from %.1f to %.1f" % (X.min(), X.max()))
print()
def show(image, caption):
"""Print a 64x64 face as text. Every 2nd column and 4th row keeps the proportions."""
ramp = " .:-=+*#%@"
lo, hi = image.min(), image.max() # stretch contrast so the shape is visible
for row in image[::4]:
print("".join(ramp[min(int((p - lo) / (hi - lo) * 10), 9)] for p in row[::2]))
print(caption, "\n")
show(faces.images[0], "person 0, photo 1")
show(faces.images[5], "person 0, photo 6 -- same person, different pose")
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=0, stratify=y
)
print("training faces:", len(X_train), " test faces:", len(X_test))
# PCA squeezes 4096 pixels into 100 numbers.
# whiten=True puts every component on the same scale, which the SVM needs.
model = make_pipeline(
PCA(n_components=100, whiten=True, random_state=0),
SVC(kernel="rbf", C=10, class_weight="balanced"),
)
model.fit(X_train, y_train)
pca = model.named_steps["pca"]
print("compressed each face from %d numbers to %d" % (X.shape[1], pca.n_components_))
print("variation kept: %.1f%%" % (100 * pca.explained_variance_ratio_.sum()))
pred = model.predict(X_test)
print("test accuracy: %.4f" % accuracy_score(y_test, pred))
print("wrong:", int((pred != y_test).sum()), "of", len(y_test))
print()
show(pca.components_[0].reshape(64, 64),
"eigenface 1 -- the strongest pattern of difference between faces")images: (400, 64, 64) people: 40 photos of each: 10 one face as numbers: 4096 pixels from 0.0 to 1.0 :=+#########%%%%%%%%%%%%%####*=- -=######%##%%%@@%%%%#%%%%%%###+- :***#*#####%%%%%%%%##########*+- :****#***+*###%%%%##*++++******: :*##%++-+-+=*%%%%###++#=.:=****= *##%*#+*****+*##%%*+****+**+*##* *#########**#######*******#####* ###%#######%%###%%#############* ##%%%%%%#%%%*###%%###%########** %*#%%%%%%###%*#@%#*###%%%%%%##** .:#%%@@%%%%#*##%##**#%%%%%%%##*. ..#%%@@%%%%%##**######%%%%%%#*-. ..:#%@@%%#**+++++*+*#%%%%%%##* . ....%%@@%%%##****+++**##%%##++ . ... *#%@%@%@%%%##%##%%%###*+*+ . ....##*#%%%%%%##########**+**+ . person 0, photo 1 ***+**#########%%%%%%%%%%%%%##*+ *****####%%####%%%%%%%%%%%%%%#*+ *****#####%%%%%%%%%%#%%%#%%%###* ++++*##**##%%#%%##**+**###*##### =++*#=*-==+###%##%+++*+=--*%%### ::.+#*****++##%##*+++*#*##***#%# #*####*###%%%######%%#%#%%%%% .. +#######*#%#%##%##%%%%##%%%%% ::::#%%#%%#%%%%%####%%#%%%%%%%%# ....#%%%%%*+#%###**%%%%%%%%%%%%% ...:+%%%##*##%###*%%%%%%%%%%%%%# ..::.%%%%%%#####%%%%%%%@%%%%%%%# ::::.+%%%*+++++++===+#%%%%%%%%## ..::..+%%%%%%#%%%####%%%%%%%##** ::::...-%%%%%%#####%#%###%###*** :::......+#%%%###########*****## person 0, photo 6 -- same person, different pose training faces: 300 test faces: 100 compressed each face from 4096 numbers to 100 variation kept: 94.4% test accuracy: 0.9500 wrong: 5 of 100 =+***###%%%%%%%%%%%%@@@@@@@@@@@# +++******#####%#######%%%@@@@@@% ++++++++++*###%%##%##*****####%% ++++******###%%%%%%##***##%####% *++**+#++*%#*#%##%%####**+#*#### +*+*++***##**##**#%%#########%%# +*+****##****#***##%#####%%%%%%% +***#*****##***++**#%%%####%%%%% =++**###%%%###**##%%%@@%%%%%%%%% =+**#%%%%%#*+**#%%%%%%@@@@@@@@%% ==**#%%@%##***##%%%%%%%%@@@@@%%# =++**#%%##**####%%%%%%%@@@@@%%#* ==++*#####*+++++**+*##%%%%@%%%#* ::=++*##%###*#####%%@@@@%%%%#*++ ::--++*####%%%%%%%@@@@@%%%##**== ...:-=+***#%%%%%%@@@%@%%###**+-- eigenface 1 -- the strongest pattern of difference between faces
What that output is telling you
4096 numbers became 100, and 94.4% of the variation survived. PCA found the hundred directions along which these faces differ most, and threw away everything else. A face is now a list of a hundred coordinates.
That compression is not a side benefit — it is what makes the method work at all. With 300 training images and 4096 raw features, any classifier would fit noise. Reduce to 100 features and the problem becomes tractable.
That last picture is an eigenface. It is not a person. It is a pattern of difference, the direction in which these 400 faces vary the most. Look at it: a broad light-dark split across the image. The first component in almost every face dataset captures overall illumination, because lighting varies more between photos than bone structure does.
Every face in the dataset is reconstructed as an average face plus a weighted mix of these patterns. The hundred weights are the face's identity, as far as this model is concerned.
95% accuracy, 5 wrong out of 100. Good for a method from 1991 running in three seconds.
Part two — finding the face
Recognition assumed a cropped, centred face. Real photos do not arrive that way.
import cv2
import numpy as np
from sklearn.datasets import fetch_olivetti_faces
faces = fetch_olivetti_faces()
# This XML file ships inside opencv-python. Nothing is downloaded.
path = cv2.data.haarcascades + "haarcascade_frontalface_default.xml"
detector = cv2.CascadeClassifier(path)
print("cascade loaded:", not detector.empty())
def make_photo(index, pad_ratio):
"""Blow a 64x64 face up to 256x256 and put a grey border around it."""
small = (faces.images[index] * 255).astype("uint8")
big = cv2.resize(small, (256, 256), interpolation=cv2.INTER_CUBIC)
pad = int(256 * pad_ratio)
return cv2.copyMakeBorder(big, pad, pad, pad, pad, cv2.BORDER_CONSTANT, value=128)
for pad_ratio in (0.0, 0.3):
photo = make_photo(0, pad_ratio)
boxes = detector.detectMultiScale(photo, scaleFactor=1.1, minNeighbors=4, minSize=(60, 60))
print(f"border {int(pad_ratio*100):>3}% -> image {photo.shape}, "
f"faces found: {len(boxes)} {boxes.tolist() if len(boxes) else ''}")
# Same test across one photo of each of the 40 people.
for pad_ratio in (0.0, 0.3):
found = sum(len(detector.detectMultiScale(make_photo(i, pad_ratio), 1.1, 4, minSize=(60, 60))) > 0
for i in range(0, 400, 10))
print(f"border {int(pad_ratio*100):>3}% -> face located in {found} of 40 people")
# Draw the box and save it so you can look at the result.
photo = make_photo(0, 0.3)
colour = cv2.cvtColor(photo, cv2.COLOR_GRAY2BGR)
for (x, y, w, h) in detector.detectMultiScale(photo, 1.1, 4, minSize=(60, 60)):
cv2.rectangle(colour, (x, y), (x + w, y + h), (0, 255, 0), 3)
cv2.imwrite("detected.png", colour)
print("saved detected.png -- open it and you will see a green box on the face")cascade loaded: True border 0% -> image (256, 256), faces found: 0 border 30% -> image (408, 408), faces found: 1 [[47, 36, 315, 315]] border 0% -> face located in 2 of 40 people border 30% -> face located in 40 of 40 people saved detected.png -- open it and you will see a green box on the face
Two faces out of forty, then forty out of forty
The only change was a grey border around the image.
A Viola-Jones cascade slides a window across the picture and asks, at each position, whether the light-and-dark pattern inside looks like a face. Those patterns include the forehead-above-eyes contrast and the cheeks-beside-nose contrast. They need the region around the face to be present. A tightly cropped face has no forehead-above-eyes boundary inside the window, so the detector sees no face.
This is worth remembering, because it is the single most common reason a beginner's detection code "does not work". The face is too tight in the frame, or too close to an edge.
Two more practical notes on detectMultiScale:
scaleFactor=1.1shrinks the image by 10% per pass to catch faces of different sizes. Values closer to 1.0 are more thorough and slower.minNeighbors=4requires four overlapping detections before accepting a box. Lower it and you get more false boxes on walls and shirts; raise it and you miss real faces.
Part three — the failure that matters most
Enrol thirty people. Then show the system ten people it has never seen.
import numpy as np
from sklearn.datasets import fetch_olivetti_faces
from sklearn.model_selection import train_test_split
from sklearn.decomposition import PCA
from sklearn.svm import SVC
from sklearn.pipeline import make_pipeline
faces = fetch_olivetti_faces()
X, y = faces.data, faces.target
enrolled = y < 30 # people 0-29 are in the system
strangers = ~enrolled # people 30-39 have never been seen
X_train, X_test, y_train, y_test = train_test_split(
X[enrolled], y[enrolled], test_size=0.25, random_state=0, stratify=y[enrolled])
model = make_pipeline(
PCA(n_components=100, whiten=True, random_state=0),
SVC(kernel="rbf", C=10, class_weight="balanced", probability=True, random_state=0))
model.fit(X_train, y_train)
print("people enrolled:", len(np.unique(y_train)))
print("accuracy on enrolled people: %.3f" % (model.predict(X_test) == y_test).mean())
# Now show it 100 photos of people it has never seen.
probs = model.predict_proba(X[strangers])
guess = model.classes_[probs.argmax(axis=1)] # both numbers from the same source
conf = probs.max(axis=1)
print("\nstranger photos shown:", len(guess))
print("times it answered 'I do not know this person':", 0)
for i in range(4):
print(f" stranger -> named as person {guess[i]:>2} ({conf[i]:.0%} sure)")
known_conf = model.predict_proba(X_test).max(axis=1)
print("\nconfidence on enrolled faces: mean %.2f, lowest %.2f" % (known_conf.mean(), known_conf.min()))
print("confidence on stranger faces: mean %.2f, highest %.2f" % (conf.mean(), conf.max()))
print("\nadding a 'not in the system' answer:")
for t in (0.20, 0.30, 0.40):
print(f" threshold {t:.2f}: {(known_conf >= t).mean():.0%} of real users let in, "
f"{(conf >= t).mean():.0%} of strangers wrongly let in")people enrolled: 30 accuracy on enrolled people: 0.987 stranger photos shown: 100 times it answered 'I do not know this person': 0 stranger -> named as person 8 (8% sure) stranger -> named as person 26 (7% sure) stranger -> named as person 5 (7% sure) stranger -> named as person 5 (7% sure) confidence on enrolled faces: mean 0.42, lowest 0.08 confidence on stranger faces: mean 0.09, highest 0.19 adding a 'not in the system' answer: threshold 0.20: 87% of real users let in, 0% of strangers wrongly let in threshold 0.30: 71% of real users let in, 0% of strangers wrongly let in threshold 0.40: 43% of real users let in, 0% of strangers wrongly let in
Read this carefully
98.7% accuracy on enrolled people. One hundred strangers, one hundred names, zero refusals.
A classifier chooses from model.classes_. "Nobody" is not in that list, so predict cannot ever return it. Every tutorial that stops at the accuracy number has hidden this from you, and it is the difference between a demo and a system.
The confidence numbers do carry the information. Enrolled faces average 0.42; strangers top out at 0.19. So a threshold works:
At 0.20, no stranger gets in and 87% of genuine users are recognised. That is a real, working "not in the system" answer, built in one line.
It also costs you 13% of your genuine users, who must try again or find a human. There is no threshold that gives you both. Every face system in the world sits somewhere on that trade-off, and where it sits is a policy decision, not a technical one.
For a phone unlock, you tune toward refusing strangers and accept occasional re-tries. For a police watchlist, the same tuning produces a flood of false alarms. The number of innocent people scanned dwarfs the number of matches sought.
This is the base rate — how rare a true match actually is. It changes everything. It is why a system that tests well in a lab behaves badly in a railway station.
Common mistakes
Splitting photos randomly instead of splitting people. Our split puts some photos of person 5 in training and others in test, which measures recognition of enrolled people. That is correct here. But if you want to measure verification on unseen identities, you must hold out whole people. Getting this wrong inflates your score enormously and is the most common error in face recognition projects.
Skipping alignment. Olivetti faces are pre-cropped and centred. Feed raw phone photos to this pipeline and accuracy collapses. Detect, then rotate so the eyes are level, then crop to a fixed size. Alignment often matters more than the classifier.
Fitting PCA on all the data before splitting. PCA().fit(X) on the full set lets test images shape the components. Keep PCA inside the Pipeline, as above, so cross-validation refits it correctly per fold.
Believing predict_proba from an SVM. SVC(probability=True) fits Platt scaling by internal cross-validation. It is slow, and it can occasionally disagree with predict. The code above takes both the name and the confidence from predict_proba for that reason.
Reporting one overall accuracy. Break your accuracy down by skin tone, age and gender before you report anything. An aggregate number can hide a group the system fails badly for.
How to make this genuinely good
- Align faces before recognising. Use a landmark detector to level the eyes, then crop to a fixed template. Cheapest large improvement available.
- Switch from classification to embeddings. Modern systems map a face to a fixed vector where distance means similarity, so a new person is enrolled by storing one vector — no retraining.
face_recognition(dlib-based) or InsightFace both do this. Expect a model download in the tens to hundreds of megabytes. - Always ship a distance threshold and return "unknown" beyond it. Report your genuine-accept rate at a fixed false-accept rate, not accuracy.
- Add liveness checking before trusting anything for access. A printed photograph defeats every system in this lesson.
- Store templates, never photos, encrypt them, and give people a way to delete theirs.
- Test per demographic group and publish the breakdown. If you cannot obtain a representative test set, that is itself a finding worth reporting to whoever asked for the system.
Try it yourself
Change n_components from 100 to 10, then to 250, and watch accuracy move. Then print pca.explained_variance_ratio_[:5] and see how much of the total sits in the first few components.
Then do the experiment that matters. Re-run the stranger test with a threshold of 0.20 and count how many genuine users were turned away. Imagine those are real people at an office door on a Monday morning. That number, not the accuracy, is what your users will experience.
What to learn next
- OpenCV — the image handling this project rests on.
- Convolutional neural networks — the architecture behind every modern face embedding.
- Image classification — the general vision task, without the identity questions.
Researcher — Mathematics and papers.
Formulas here are written in plain text, since the site renders no maths typesetting library.
Eigenfaces
Let X be an n x d matrix of n vectorised face images with d = 4096 pixels, and mu the mean face. Centre the data as Xc = X - mu. PCA seeks the orthonormal basis U_k of k directions maximising retained variance, given by the top-k eigenvectors of the sample covariance:
S = (1 / (n - 1)) * Xc^T Xc (d x d)
S u_i = lambda_i u_i, lambda_1 >= lambda_2 >= ...u_iis thei-th eigenvector, reshaped to 64x64 for display as an eigenface.lambda_iis the variance captured alongu_i;explained_variance_ratio_[i] = lambda_i / sum_j lambda_j.
A face is then represented by its k coordinates z = U_k^T (x - mu), and reconstructed as x_hat = mu + U_k z.
The Turk-Pentland trick. When d >> n, forming the d x d matrix S is wasteful and its rank is at most n - 1. Instead compute eigenvectors v_i of the n x n Gram matrix Xc Xc^T; then u_i = Xc^T v_i / || Xc^T v_i ||. For our case that is a 300x300 eigenproblem rather than 4096x4096. scikit-learn achieves the same effect via randomized SVD (Halko, Martinsson and Tropp, 2011) at O(n d k).
Whitening. whiten=True rescales each coordinate to unit variance, z_i / sqrt(lambda_i). Without it the RBF kernel's isotropic distance is dominated by the first few high-variance components — which encode illumination, not identity. This single flag is often the difference between a working and a broken eigenface pipeline.
Why PCA is the wrong objective, and Fisherfaces
PCA maximises total variance, which is unsupervised. It does not know about identity. Belhumeur, Hespanha and Kriegman (1997) observed that for face data the largest variance directions correspond to lighting, so PCA preferentially retains exactly the nuisance factor.
Fisher's linear discriminant instead maximises the ratio of between-class to within-class scatter:
W* = argmax | W^T S_B W | / | W^T S_W W |
W
S_B = sum_c n_c (mu_c - mu)(mu_c - mu)^T
S_W = sum_c sum_{x in c} (x - mu_c)(x - mu_c)^Tmu_c and n_c are the mean and count of class c. S_W is singular when n < d, so the standard construction is Fisherfaces: project with PCA to n - C dimensions first, then apply LDA. LDA yields at most C - 1 discriminant directions, which is a hard ceiling.
The modern formulation: metric learning
Contemporary systems do not classify. They learn an embedding f: image -> R^128 (or 512) such that Euclidean or cosine distance encodes identity, then compare against stored templates. This is what makes enrolment cheap: adding a person stores one vector and retrains nothing.
FaceNet (Schroff, Kalenichenko and Philbin, 2015) trains directly on the triplet loss:
L = sum over triplets max( 0, || f(a) - f(p) ||^2 - || f(a) - f(n) ||^2 + alpha )a is an anchor image, p a different image of the same person, n an image of a different person, and alpha the required margin. Embeddings are L2-normalised to the unit hypersphere. The practical difficulty is triplet mining: most random triplets already satisfy the constraint and produce zero gradient, so semi-hard negative mining within each batch is essential.
ArcFace (Deng, Guo, Xue and Zafeiriou, 2019) replaced triplets with a margin applied inside the softmax, which removed the mining problem:
L = -(1/N) sum_i log [ e^{s * cos(theta_yi + m)}
/ ( e^{s * cos(theta_yi + m)} + sum_{j != yi} e^{s * cos(theta_j)} ) ]theta_jis the angle between the normalised embedding and the normalised class-jweight vector.mis an additive angular margin, typically 0.5 radians.sis a scale factor, typically 64, compensating for the bounded range of cosine.
The additive angular margin enforces a constant geodesic separation on the hypersphere, which is geometrically cleaner than the multiplicative margin of SphereFace or the cosine margin of CosFace. ArcFace and its descendants remain the standard backbone objective.
Detection followed a parallel path: Viola and Jones (2001) with Haar features and AdaBoost cascades, then MTCNN (Zhang et al., 2016) for joint detection and landmark alignment, then single-shot detectors such as RetinaFace (Deng et al., 2019).
Open-set recognition, and why the metric matters
The developer block demonstrates the closed-set assumption failing. Formally, closed-set identification assumes the probe identity is in the gallery, so argmax over gallery scores is well-defined. Open-set identification does not, and requires a reject region.
Report TAR at a fixed FAR — true accept rate at a false accept rate, typically FAR = 1e-4 or 1e-6 — never accuracy. Accuracy at a single operating point is uninterpretable because it entangles the threshold with the class balance.
For watchlist deployments the base rate dominates. With a false accept rate of 1e-5 and 100,000 people scanned per day, expect roughly one false alarm per day even if the watchlist is empty. Scale to a busy transport hub and the alarms overwhelm the true detections. This is the false positive paradox, and no improvement in FAR of the magnitude achievable today removes it — only reducing the scanned population does.
Demographic differentials
Buolamwini and Gebru (2018), Gender Shades, audited three commercial gender-classification APIs and found error rates up to 34.7% for darker-skinned women against 0.8% for lighter-skinned men. They also introduced the Pilot Parliaments Benchmark, constructed for balance across skin type and gender, because existing benchmarks were not.
Grother, Ngan and Hanaoka (2019), NIST FRVT Part 3: Demographic Effects (NISTIR 8280), evaluated 189 algorithms from 99 developers and confirmed the pattern at scale. Their key findings: false positive differentials across demographic groups spanned factors of 10 to 100 in many algorithms; the direction and magnitude varied with the training data's origin; and false positives were far more demographically variable than false negatives. Since false positives are the errors that cause wrongful accusations, that asymmetry is exactly the wrong way round.
The mechanism is not mysterious — unrepresentative training distributions, plus image pipelines (exposure, contrast) tuned on lighter skin. It is measurable, and it is fixable in direction if not fully in magnitude. The obligation is therefore to measure and publish per-group rates rather than an aggregate.
Security and governance
Presentation attacks. A printed photo, a screen replay or a silicone mask defeats a matcher that has no liveness check. ISO/IEC 30107-3 defines the evaluation framework. Report attack presentation classification error rate alongside matching accuracy, or the accuracy figure is not meaningful for access control.
Template protection. Embeddings are not anonymous. Model inversion can reconstruct a recognisable face from an embedding (Mai et al., 2018, on face reconstruction from deep templates). Treat templates as biometric data with the same protections as the images, and prefer cancellable-biometric schemes that allow a compromised template to be revoked and reissued.
Legal position. Face data is special-category personal data under GDPR Article 9, requiring an explicit lawful basis. Illinois BIPA requires written consent before collection and has produced substantial settlements. The EU AI Act classifies remote biometric identification as high-risk with significant restrictions on real-time use in public spaces. India's Digital Personal Data Protection Act, 2023 governs processing of personal data including biometrics, with consent and purpose-limitation obligations. Determine the applicable regime before collecting a single face.
Papers
- Sirovich and Kirby, Low-Dimensional Procedure for the Characterization of Human Faces, JOSA A, 1987.
- Turk and Pentland, Eigenfaces for Recognition, Journal of Cognitive Neuroscience, 1991.
- Belhumeur, Hespanha and Kriegman, Eigenfaces vs. Fisherfaces, IEEE TPAMI, 1997.
- Viola and Jones, Rapid Object Detection using a Boosted Cascade of Simple Features, CVPR 2001.
- Huang, Ramesh, Berg and Learned-Miller, Labeled Faces in the Wild, UMass TR 07-49, 2007.
- Taigman, Yang, Ranzato and Wolf, DeepFace, CVPR 2014.
- Schroff, Kalenichenko and Philbin, FaceNet, CVPR 2015 — arxiv.org/abs/1503.03832
- Zhang, Zhang, Li and Qiao, Joint Face Detection and Alignment using MTCNN, IEEE SPL, 2016 — arxiv.org/abs/1604.02878
- Buolamwini and Gebru, Gender Shades, FAccT 2018.
- Deng, Guo, Xue and Zafeiriou, ArcFace, CVPR 2019 — arxiv.org/abs/1801.07698
- Grother, Ngan and Hanaoka, Face Recognition Vendor Test Part 3: Demographic Effects, NISTIR 8280, 2019.
What to learn next
- OpenCV — the image handling this project rests on.
- Convolutional neural networks — the architecture behind every modern face embedding.
- Image classification — the general vision task, without the identity questions.