How it works
How face recognition works
Face recognition finds a face, straightens it, compresses it into a list of numbers, and decides identity by measuring how close that list is to a stored one.
- 3 min read
- Updated
On this page 7
Your phone unlocks the moment it sees you, and Google Photos groups twenty years of pictures by person. Neither system stores what you "look like" in any human sense. They store a short list of numbers, and compare distances. Here is the pipeline that gets there.
The pipeline at a glance
photo or camera frame
|
v
[1. detect] -> find the face box in the image
|
v
[2. align] -> locate eyes, nose, mouth; straighten the crop
|
v
[3. embed] -> one face -> one list of ~512 numbers
|
v
[4. compare] -> distance to stored numbers
|
v
[5. decide] -> close enough? same personStage 1 — find the face
Before recognising anyone, the system must find where faces are. A face detector — a small, fast neural network — scans the image and returns a box around each face. Detection answers "is there a face here?", nothing about whose face it is.
This split matters. Detection is cheap and runs constantly, on every camera frame. The expensive identity work only runs on the cropped boxes the detector hands over.
Stage 2 — straighten the crop
Faces arrive tilted, turned and at every size. The system locates landmarks — standard reference points like eye corners, nose tip and mouth edges — then rotates and scales the crop so those points sit in standard positions. Eyes level, face centred.
This is like a passport-photo booth insisting you face forward under even light. With every face normalised the same way, the next stage compares like with like.
Stage 3 — the face becomes numbers
The heart of the system. The aligned crop goes through a deep network, historically a CNN — a network built for images. It outputs a face embedding: a list of roughly 128 to 512 numbers summarising the face.
The network was trained on millions of photos with one goal. Two photos of the same person must land close together in that number space. Photos of different people must land far apart. Lighting, glasses, beard growth and ageing shift the numbers a little. The gap between two different people is bigger. That trained separation is the entire trick.
Note what is stored: the numbers, not the photo. You cannot reconstruct a face from a good embedding, which matters for privacy.
Stage 4 — measure the distance
Recognition is now arithmetic. To check "is this the phone's owner?", compute the distance between the fresh embedding and the enrolled one. This one-to-one check is verification. Photo apps do the one-to-many version — compare against every known person, take the nearest — called identification, usually with a vector index doing the fast lookup.
Stage 5 — the threshold decides
A distance is not a decision. A threshold converts it: closer than the line means same person. Set the line tight and the owner gets rejected on a dim morning. Set it loose and a sibling unlocks the phone. Every deployment tunes this trade-off, and phone makers tune it very tight, preferring a second try over a false unlock.
Two honest caveats. Accuracy is not uniform: systems trained on unbalanced photo sets make more errors on the groups they saw least, which has caused real wrongful arrests. And a printed photo can fool a plain camera, so phones add depth sensors and liveness checks before trusting a match.
Which lessons teach each stage
- The field this all lives in: What is computer vision?
- Stages 1 and 3, networks that read images: CNNs and Image classification
- Stage 3, meaning as coordinates: Embeddings
- Stage 4, the fast nearest-neighbour lookup: Vector databases