How it works

How face recognition works

Face recognition finds a face, straightens it, compresses it into a list of numbers, and decides identity by measuring how close that list is to a stored one.

On this page 7
  1. The pipeline at a glance
  2. Stage 1 — find the face
  3. Stage 2 — straighten the crop
  4. Stage 3 — the face becomes numbers
  5. Stage 4 — measure the distance
  6. Stage 5 — the threshold decides
  7. Which lessons teach each stage

Your phone unlocks the moment it sees you, and Google Photos groups twenty years of pictures by person. Neither system stores what you "look like" in any human sense. They store a short list of numbers, and compare distances. Here is the pipeline that gets there.

The pipeline at a glance

 photo or camera frame
        |
        v
 [1. detect]     ->  find the face box in the image
        |
        v
 [2. align]      ->  locate eyes, nose, mouth; straighten the crop
        |
        v
 [3. embed]      ->  one face  ->  one list of ~512 numbers
        |
        v
 [4. compare]    ->  distance to stored numbers
        |
        v
 [5. decide]     ->  close enough? same person

Stage 1 — find the face

Before recognising anyone, the system must find where faces are. A face detector — a small, fast neural network — scans the image and returns a box around each face. Detection answers "is there a face here?", nothing about whose face it is.

This split matters. Detection is cheap and runs constantly, on every camera frame. The expensive identity work only runs on the cropped boxes the detector hands over.

Stage 2 — straighten the crop

Faces arrive tilted, turned and at every size. The system locates landmarks — standard reference points like eye corners, nose tip and mouth edges — then rotates and scales the crop so those points sit in standard positions. Eyes level, face centred.

This is like a passport-photo booth insisting you face forward under even light. With every face normalised the same way, the next stage compares like with like.

Stage 3 — the face becomes numbers

The heart of the system. The aligned crop goes through a deep network, historically a CNN — a network built for images. It outputs a face embedding: a list of roughly 128 to 512 numbers summarising the face.

The network was trained on millions of photos with one goal. Two photos of the same person must land close together in that number space. Photos of different people must land far apart. Lighting, glasses, beard growth and ageing shift the numbers a little. The gap between two different people is bigger. That trained separation is the entire trick.

Note what is stored: the numbers, not the photo. You cannot reconstruct a face from a good embedding, which matters for privacy.

Stage 4 — measure the distance

Recognition is now arithmetic. To check "is this the phone's owner?", compute the distance between the fresh embedding and the enrolled one. This one-to-one check is verification. Photo apps do the one-to-many version — compare against every known person, take the nearest — called identification, usually with a vector index doing the fast lookup.

Stage 5 — the threshold decides

A distance is not a decision. A threshold converts it: closer than the line means same person. Set the line tight and the owner gets rejected on a dim morning. Set it loose and a sibling unlocks the phone. Every deployment tunes this trade-off, and phone makers tune it very tight, preferring a second try over a false unlock.

Two honest caveats. Accuracy is not uniform: systems trained on unbalanced photo sets make more errors on the groups they saw least, which has caused real wrongful arrests. And a printed photo can fool a plain camera, so phones add depth sensors and liveness checks before trusting a match.

Which lessons teach each stage