Faces, People and Pose

Face detection models

A face detector answers one question — where are the faces — by scoring thousands of pre-placed boxes and then deleting the duplicates it created on purpose.

On this page 9
  1. The short answer
  2. Why the split matters
  3. How it works
  4. Why guess-boxes instead of searching
  5. The five points
  6. Where you have already seen this
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A face detector finds where the faces are in a picture. It does not know who they are.

Think of a crowded railway platform and a friend scanning it for anyone waving. Her eyes sweep the whole platform, marking each waving figure. She has not recognised anybody yet. She has only marked positions.

That is the whole job of a face detector. Mark the positions, hand them on, stop.

Why the split matters

Every face system in the world is built as a chain of small steps. Detection is the first link.

Detection finds the boxes. Alignment straightens each face. Recognition decides who it is.

Keeping them apart is a design choice, not an accident. Each step can be swapped, tested and measured on its own. It also means you can run detection alone and never touch identity at all.

Blurring faces in street photography needs detection only. Counting how many people entered a shop needs detection only. Both stop before anything personal starts.

How it works

A modern detector does not search the picture bit by bit. Instead, it lays down a fixed grid of guess-boxes before it looks at anything. These are called anchors or priors — boxes of standard size, placed at every position, waiting to be judged.

   photo
     │
     ▼
   lay down thousands of guess-boxes, everywhere, at several sizes
     │
     ▼
   for each guess-box the network says:
        "is there a face here?"   (a confidence)
        "nudge me: left a bit, wider a bit"   (a correction)
        "the eyes, nose and mouth are here"   (five points)
     │
     ▼
   throw away the low-confidence ones
     │
     ▼
   many boxes now cover the SAME face
     │
     ▼
   keep the best, delete its near-copies
     │
     ▼
   final list of face boxes

That last deleting step has a name: non-maximum suppression, meaning keep the strongest box and remove anything overlapping it heavily.

Why guess-boxes instead of searching

Older detectors slid a window over the picture and asked the same question at every position, one at a time. That was slow, and it had to be repeated at many sizes.

Guess-boxes turn the search into a single pass. The network sees the whole picture once and grades every box together. A phone can do this many times a second.

The cost is those duplicates. Several nearby boxes fire on the same face, so the cleanup step is not optional.

The five points

Face detectors usually return more than a box. They also return five landmark points: two eyes, the nose tip and the two mouth corners.

This is a small detail with a large consequence. The next step in the chain needs those five points to straighten the face. Getting them for free during detection saves an entire extra model.

Where you have already seen this

  • The yellow squares your phone camera draws before you press the shutter.
  • Blurred faces on Google Street View.
  • Auto-cropping a group photo so nobody's head is cut off.
  • The count of "people in frame" on a video-conference dashboard.

What is honestly hard here

Detection sounds easy and stops being easy fast. Small faces in the background. Faces turned sideways. Faces half behind a shoulder. Faces in near-darkness or motion blur. Every one is a known weak spot.

There is a standard hard test set for this, and the best models still miss faces on it. A detector that scores well on posed portraits can perform far worse on a real CCTV frame.

Remember this

  • A detector finds where, never who. That boundary is deliberate.
  • It grades thousands of pre-placed guess-boxes in one pass, then deletes the duplicates.
  • Most detectors hand back five landmark points as well, which the next stage needs.

What to learn next

  • Object detection — the general machinery this is a special case of.
  • YOLO — anchors, NMS and single-shot heads in their most-used form.
  • OpenCV — the library that ships the detector you will reach for first.

Developer — Code and libraries.

Setup

bash
pip install numpy

The runnable example below is pure NumPy. It builds the three pieces every single-shot face detector has: priors, box decoding, and non-maximum suppression. Understanding these three makes every detector config file readable.

The three pieces, written out

detector_core.py
import numpy as np

# ---- 1. Priors ("anchors"): fixed boxes the detector measures faces against.
def make_priors(image_size, stride, sizes):
    """One prior per (cell, size). Returns cx, cy, w, h in pixels."""
    n = image_size // stride
    centres = (np.arange(n) + 0.5) * stride          # cell centre, not cell corner
    cx, cy = np.meshgrid(centres, centres, indexing="xy")
    priors = []
    for s in sizes:
        priors.append(np.stack([cx.ravel(), cy.ravel(),
                                np.full(cx.size, s, float),
                                np.full(cx.size, s, float)], axis=1))
    return np.concatenate(priors, axis=0)

priors = make_priors(image_size=64, stride=8, sizes=[16, 32])
print("feature map 8x8, 2 sizes  ->", priors.shape[0], "priors")
print("first prior (cx, cy, w, h):", priors[0])

# ---- 2. Decode: the network never predicts pixels, it predicts nudges to a prior.
def decode(priors, deltas, var=(0.1, 0.2)):
    cx = priors[:, 0] + deltas[:, 0] * var[0] * priors[:, 2]
    cy = priors[:, 1] + deltas[:, 1] * var[0] * priors[:, 3]
    w  = priors[:, 2] * np.exp(deltas[:, 2] * var[1])
    h  = priors[:, 3] * np.exp(deltas[:, 3] * var[1])
    return np.stack([cx - w / 2, cy - h / 2, cx + w / 2, cy + h / 2], axis=1)

# Three priors fire on one real face. Their raw nudges differ slightly.
idx    = np.array([27, 28, 91])
deltas = np.array([[ 0.9,  0.4, 1.9, 1.9],
                   [-0.6,  0.5, 1.8, 2.0],
                   [ 0.2, -0.3, 0.4, 0.5]])
scores = np.array([0.97, 0.93, 0.88])
boxes  = decode(priors[idx], deltas)
print("\ndecoded boxes (x1, y1, x2, y2):")
for b, s in zip(boxes, scores):
    print(f"  score {s:.2f}  [{b[0]:6.2f} {b[1]:6.2f} {b[2]:6.2f} {b[3]:6.2f}]")

# ---- 3. NMS: three boxes, one face. Keep the best, drop its near-duplicates.
def iou(a, boxes):
    x1 = np.maximum(a[0], boxes[:, 0]); y1 = np.maximum(a[1], boxes[:, 1])
    x2 = np.minimum(a[2], boxes[:, 2]); y2 = np.minimum(a[3], boxes[:, 3])
    inter = np.clip(x2 - x1, 0, None) * np.clip(y2 - y1, 0, None)
    area  = lambda b: (b[..., 2] - b[..., 0]) * (b[..., 3] - b[..., 1])
    return inter / (area(a) + area(boxes) - inter)

def nms(boxes, scores, thresh=0.4):
    order, keep = np.argsort(-scores), []
    while order.size:
        i = order[0]; keep.append(i)
        overlap = iou(boxes[i], boxes[order[1:]])
        order = order[1:][overlap < thresh]      # drop anything hugging the winner
    return keep

print("\nIoU of box 0 against boxes 1 and 2:", np.round(iou(boxes[0], boxes[1:]), 3))
kept = nms(boxes, scores, thresh=0.4)
print("kept after NMS:", kept, " -> ", len(kept), "face(s) reported")
print("NMS at a loose 0.7 threshold:", nms(boxes, scores, thresh=0.7),
      "-> the same face reported more than once")
Output
feature map 8x8, 2 sizes  -> 128 priors
first prior (cx, cy, w, h): [ 4.  4. 16. 16.]

decoded boxes (x1, y1, x2, y2):
  score 0.97  [ 17.74  16.94  41.14  40.34]
  score 0.93  [ 23.57  16.87  46.51  40.73]
  score 0.88  [ 11.31   9.36  45.97  44.72]

IoU of box 0 against boxes 1 and 2: [0.601 0.447]
kept after NMS: [0]  ->  1 face(s) reported
NMS at a loose 0.7 threshold: [0, 1, 2] -> the same face reported more than once

Reading that output

128 priors from an 8x8 grid. Sixty-four cells, two sizes each. A real detector runs this over three or four feature maps at strides 8, 16 and 32, which is how a few thousand priors appear from one small image.

The network predicts nudges, never coordinates. Look at decode. The centre moves by a fraction of the prior's own width; the size is multiplied by an exponential. This keeps every predicted number near zero and near unit variance, which is far easier to train than raw pixel values. The var=(0.1, 0.2) pair is the variance encoding, and it must match between training and inference or every box lands in the wrong place.

IoU of 0.601 and 0.447. Both are above the 0.4 threshold, so both duplicates die. Box 2 came from a different prior size at the same location — cross-scale duplicates are the common case, not the exception.

The loose threshold reports one face three times. This is the single most common visible bug in a detection pipeline. If your face count is inflated, check the NMS threshold before you blame the model.

Running a real detector

OpenCV ships a face detector class in the main package. It needs a small ONNX file, around 230 KB, that you download separately.

yunet.py
import cv2

# Model: face_detection_yunet_2023mar.onnx from github.com/opencv/opencv_zoo
detector = cv2.FaceDetectorYN.create(
    "face_detection_yunet_2023mar.onnx", "", (320, 320),
    score_threshold=0.9, nms_threshold=0.3, top_k=5000)

img = cv2.imread("group_photo.jpg")
detector.setInputSize((img.shape[1], img.shape[0]))   # MUST be set before detect()
_, faces = detector.detect(img)

if faces is not None:
    for f in faces:
        x, y, w, h = f[:4].astype(int)
        landmarks = f[4:14].reshape(5, 2)             # eyes, nose, mouth corners
        print(x, y, w, h, round(float(f[14]), 3))

No output block for this one, and that is deliberate. The result depends entirely on which photo you feed it. Printing invented boxes would teach you to expect numbers that will not appear.

Verified against opencv-python 4.10 and 4.11; cv2.FaceDetectorYN has been in the main package since 4.5.4.

Common mistakes

Forgetting setInputSize. YuNet is created with a fixed input size and needs telling when the real image differs. Skip this and you get either an exception or boxes scaled to the wrong frame.

Mismatched variance constants. If you port weights between repositories, the (0.1, 0.2) encoding may differ. Boxes will appear roughly right and consistently a few pixels off.

Raising the confidence threshold to reduce false positives. It also deletes every small and side-facing face. Tune NMS first; it fixes duplicates without costing recall.

Testing only on portraits. Detection accuracy on posed faces tells you almost nothing about a CCTV frame. Build a small evaluation set from images that look like your real input, including the hard ones.

Feeding BGR to a model trained on RGB. OpenCV loads BGR. Most PyTorch models expect RGB. The detector still works, a bit worse, and nothing errors — which is why this one survives to production.

Try it yourself

Add a fourth prediction that lands on a genuinely different face, well away from the first three. Confirm NMS keeps two boxes rather than one. Then shrink the NMS threshold to 0.1 and watch a legitimate second face disappear when two people stand shoulder to shoulder.

What to learn next

  • Object detection — the general machinery this is a special case of.
  • YOLO — anchors, NMS and single-shot heads in their most-used form.
  • OpenCV — the library that ships the detector you will reach for first.

Researcher — Mathematics and papers.

The single-shot formulation

Face detection is class-agnostic object detection with one class and a strong scale prior. Given an image $I$, a network produces for each prior $a_i = (a_x, a_y, a_w, a_h)$ a tuple $(p_i, \delta_i, \ell_i)$: a face probability, four box deltas, and ten landmark offsets.

Decoding follows the SSD parameterisation (Liu et al., 2016):

$$ \hat{x} = a_x + \sigma_c\, \delta_x\, a_w, \qquad \hat{w} = a_w \exp(\sigma_s\, \delta_w) $$

$\sigma_c$ and $\sigma_s$ are the centre and size variances, conventionally $0.1$ and $0.2$. They rescale the regression targets to roughly unit variance; they are a normalisation constant, not a learned quantity.

Training uses a multi-task loss. RetinaFace (Deng et al., CVPR 2020, arxiv.org/abs/1905.00641) writes it as:

$$ L = L_{\text{cls}} + \lambda_1 L_{\text{box}} + \lambda_2 L_{\text{pts}} + \lambda_3 L_{\text{pixel}} $$

$L_{\text{cls}}$ is softmax or focal loss over face/background, $L_{\text{box}}$ a smooth-$L_1$ on the four deltas, $L_{\text{pts}}$ a smooth-$L_1$ on the five landmarks, and $L_{\text{pixel}}$ a self-supervised 3D mesh decoder branch. The paper's ablation shows the landmark branch improves the box AP as well — a supervision signal, not only an output.

Positive assignment and the small-face problem

A prior is positive when its IoU with a ground-truth box exceeds $0.5$ (typically $0.35$ for faces, which are small). WIDER FACE (Yang et al., CVPR 2016) splits its validation set into Easy, Medium and Hard, where Hard is dominated by faces under 20 pixels.

The core difficulty is statistical. Tiny faces match very few priors, so they contribute almost nothing to the loss. Two responses dominate:

  • Sample redistribution. SCRFD (Guo et al., 2021, arxiv.org/abs/2105.04714) reallocates positive samples toward shallow, high-resolution stages, and searches how to redistribute compute across backbone, neck and head under a fixed FLOP budget.
  • Denser priors at fine strides. Stride-8 or even stride-4 levels, at the cost of a much larger prior count.

Cost and accuracy, roughly

DetectorYearWIDER Hard APNote
MTCNN (Zhang et al.)2016~0.61Three-stage cascade, CPU-friendly, now dated
RetinaFace (R50)2020~0.91Heavy multi-scale configuration reaches ~0.918
SCRFD-0.5GF2021~0.6850.508 GFLOPs; the low end of the family
SCRFD-10GF2021~0.831Around 4.9 ms at VGA on a mid-range GPU
YuNet2023~0.811Roughly 75k parameters; ~1.6 ms at 320x320 on a desktop CPU

Reported latencies are hardware- and batch-dependent; treat the ordering as reliable and the absolute numbers as indicative. YuNet is documented in Wu, Peng et al., YuNet: A Tiny Millisecond-level Face Detector, Machine Intelligence Research, 2023.

BlazeFace (Bazarevsky et al., 2019, arxiv.org/abs/1907.05047) takes a different route again: a near-frontal, single-face-dominant model designed for mobile GPU, with a tie-resolution strategy replacing NMS to avoid jitter between video frames. It is the detector inside the MediaPipe face pipeline.

NMS and its alternatives

Greedy NMS is $O(n^2)$ in the worst case and non-differentiable. Two variants matter in practice:

  • Soft-NMS (Bodla et al., ICCV 2017) decays overlapping scores rather than deleting them, which helps in crowds where two real faces genuinely overlap.
  • Anchor-free heads (FCOS-style, and YOLO-family face variants such as YOLO5Face, Qi et al., 2021) predict distances to box edges from each location, cutting prior tuning entirely. They still need NMS.

Evaluation caveats

WIDER FACE AP is computed at IoU $0.5$ and rewards recall of very small faces heavily. Two consequences follow.

First, a detector tuned for WIDER Hard may return many low-confidence boxes that are useless downstream. Second, AP says nothing about landmark quality, which is what actually limits recognition accuracy in a full pipeline. Measure the metric your system depends on, which is usually alignment error, not box AP.

Where detection stops

A detector is scope-limiting infrastructure. Detection alone supports counting, autofocus, cropping and redaction, and produces no biometric identifier. Adding an embedding model changes the legal character of the system entirely, bringing biometric-data obligations that detection alone does not trigger — see responsible deployment.

Papers

What to learn next

  • Object detection — the general machinery this is a special case of.
  • YOLO — anchors, NMS and single-shot heads in their most-used form.
  • OpenCV — the library that ships the detector you will reach for first.