Faces, People and Pose

Face embeddings and ArcFace

A face embedding turns a face crop into a list of numbers on a sphere, and ArcFace is the training trick that forces every person's numbers into a tight, well-separated cluster.

On this page 8
  1. The short answer
  2. Why it exists
  3. How it works
  4. The trick that made it work
  5. Where you have already seen this
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A face embedding is a list of numbers standing in for a face. Two photos of one person give two nearby lists.

Think of seating guests at a wedding. You do not label anyone. You put people who get along near each other, and people who do not far apart. Then "are these two from the same family?" becomes "are they sitting close?"

An embedding does this with faces. Every face gets a seat. Same person, seats almost on top of each other. Different people, seats across the hall.

Why it exists

The first instinct is to train a classifier with one output per person. That works and then breaks the moment reality arrives.

New people keep appearing. A phone unlock model cannot be retrained every time somebody buys a phone.

Most people have one or two photos. A classifier needs many examples per class.

There are billions of people. An output layer with a billion slots is not a thing you build.

The seating-plan idea sidesteps all three. Learn a good seating rule from a few million faces. Then apply that rule to a person the model has never seen. You do not need a slot for them; you need a seat, and the rule provides one.

This is why a phone can learn your face from a handful of scans. It is not learning you. It already knows how to seat faces, and it is writing down where yours sits.

How it works

   aligned face crop
        │
        ▼
   neural network
        │
        ▼
   a list of numbers  ("the embedding")
        │
        ▼
   scale it so every list has the same length
        │
        ▼
   compare two lists by the angle between them
        │
        ▼
   small angle → same person.  Large angle → different people.

That "same length" step matters more than it looks. Once every list has the same length, only direction is left. Brightness, contrast, camera quality — the things that change how strong the numbers are — stop mattering. Direction carries identity.

The trick that made it work

For years the seating plans were mediocre. Same-person seats were near, but not near enough, and the clusters bled into each other.

The fix, called ArcFace, is a piece of deliberate unfairness during training. When the model is learning, its answer for the correct person is handicapped. The model is told the true seat is further away than it really is. It must get the answer right anyway.

Because the training is harder than reality, the model overshoots. It pushes each person's cluster tighter and further from its neighbours than the plain task required. At test time the handicap is removed, and there is now a comfortable gap between clusters.

It is the same idea as training with a weighted bat, or practising a speech with a noisy fan running. Make practice harder than the real thing, and the real thing has margin in it.

Where you have already seen this

  • Google Photos grouping every picture of one cousin, without ever being told her name.
  • Face unlock recognising you after a single enrolment.
  • Airport e-gates comparing your face to the chip in your passport.
  • "Is this the same person?" checks during online account verification.

What is honestly hard here

An embedding is a biometric identifier. It is not anonymous, and calling it a harmless list of numbers is wrong.

Research has repeatedly shown that a recognisable face can be reconstructed from an embedding. Storing embeddings is storing biometric data, with the legal duties that come with it in most jurisdictions. Treat an embedding database like a password database: encrypted, access-logged, and deletable on request.

Remember this

  • An embedding is a seat in a space where distance means "same person or not".
  • Scaling every embedding to the same length leaves direction, and direction carries identity.
  • ArcFace trains with a handicap, so the finished clusters have a real gap between them.

What to learn next

  • Embeddings — the same idea, applied to words instead of faces.
  • Loss functions — where the margin lives and what it does to gradients.
  • Vector databases — searching millions of embeddings without comparing every one.

Developer — Code and libraries.

Setup

bash
pip install torch

Run against torch 2.5.1 on CPU. The example is a handful of tensors; no GPU and no download.

The ArcFace head, in full

arcface.py
import torch
import torch.nn.functional as F

torch.manual_seed(0)

D, C = 8, 4            # 8-dim embedding, 4 identities in the training set
s, m = 64.0, 0.50      # scale and angular margin from the ArcFace paper

emb = torch.randn(2, D)                 # two face crops out of the backbone
W   = torch.randn(C, D)                 # one prototype column per identity
y   = torch.tensor([0, 2])              # who they actually are

e = F.normalize(emb, dim=1)             # every embedding onto the unit sphere
w = F.normalize(W,   dim=1)             # every prototype too
cos = e @ w.T                           # so the dot product IS the cosine
print("cosine of each face against each identity prototype:")
print(cos.round(decimals=4))
print("angles in degrees:")
print(torch.rad2deg(torch.acos(cos.clamp(-1, 1))).round(decimals=2))

theta_y = torch.acos(cos[torch.arange(2), y].clamp(-1 + 1e-7, 1 - 1e-7))
target  = torch.cos(theta_y + m)        # push the true class further away, then ask it to win
logits  = s * cos.clone()
logits[torch.arange(2), y] = s * target

plain = F.cross_entropy(s * cos, y)
arc   = F.cross_entropy(logits, y)
print("\nangle to the CORRECT identity :", torch.rad2deg(theta_y).round(decimals=2).tolist())
print("after adding the 0.5 rad margin:",
      torch.rad2deg(theta_y + m).round(decimals=2).tolist())
print("cos before margin:", cos[torch.arange(2), y].round(decimals=4).tolist())
print("cos after  margin:", target.round(decimals=4).tolist())
print(f"\nloss without margin: {plain.item():.4f}")
print(f"loss with    margin: {arc.item():.4f}   <- harder, so the model must separate further")

# Verification uses cosine between two embeddings. Nothing else.
a, b = F.normalize(torch.randn(1, D), dim=1), F.normalize(torch.randn(1, D), dim=1)
print("\ncosine between two unrelated embeddings:", round(float(a @ b.T), 4))
print("cosine of an embedding with itself      :", round(float(a @ a.T), 4))
print("squared L2 distance = 2 - 2*cos         :",
      round(float(torch.cdist(a, b) ** 2), 4), "vs", round(2 - 2 * float(a @ b.T), 4))

# Why real face embeddings are 512-dimensional and not 8.
for dim in (8, 128, 512):
    v = F.normalize(torch.randn(20000, dim), dim=1)
    c = (v[:10000] * v[10000:]).sum(1)
    print(f"dim {dim:4d}: random-pair cosine  mean {c.mean():+.4f}  sd {c.std():.4f}  "
          f"|max| {c.abs().max():.4f}")
Output
cosine of each face against each identity prototype:
tensor([[-0.1366,  0.4571, -0.0041,  0.5253],
        [-0.0718,  0.7553,  0.1741,  0.0753]])
angles in degrees:
tensor([[97.8500, 62.8000, 90.2300, 58.3100],
        [94.1200, 40.9500, 79.9700, 85.6800]])

angle to the CORRECT identity : [97.8499984741211, 79.97000122070312]
after adding the 0.5 rad margin: [126.5, 108.62000274658203]
cos before margin: [-0.13660000264644623, 0.17409999668598175]
cos after  margin: [-0.5947999954223633, -0.31929999589920044]

loss without margin: 39.7858
loss with    margin: 70.2377   <- harder, so the model must separate further

cosine between two unrelated embeddings: 0.7331
cosine of an embedding with itself      : 1.0
squared L2 distance = 2 - 2*cos         : 0.5338 vs 0.5338
dim    8: random-pair cosine  mean -0.0052  sd 0.3495  |max| 0.9548
dim  128: random-pair cosine  mean +0.0000  sd 0.0888  |max| 0.3659
dim  512: random-pair cosine  mean +0.0010  sd 0.0437  |max| 0.1610

Reading that output carefully

The correct identity sits at 97.85 degrees. That is worse than a right angle, which is what a random guess gives. This is an untrained network with random weights, and the output is telling the truth about it. Training is the process of driving that number down toward zero.

The margin turns a 97.85-degree angle into 126.5 degrees. The loss rises from 39.79 to 70.24. That gap is the learning pressure. Every gradient step now has to overcome the handicap as well as the real task.

The scale s = 64 is not decoration. Cosines live in a range of length two. Cross-entropy over logits bounded that tightly saturates and produces almost no gradient. Multiplying by 64 stretches them into a range where softmax has something to work with. Remove s and the model trains to nothing.

A distance and a cosine are the same fact. 0.5338 from both formulas. For unit vectors, squared Euclidean distance equals two minus twice the cosine, so ranking by one ranks identically by the other. Any claim that a system "uses distance, not similarity" is describing a display choice.

The dimension table is the punchline. In 8 dimensions two random vectors reach a cosine of 0.95 by luck — the space is too small to keep strangers apart. In 512 dimensions the largest of ten thousand random pairs reaches only 0.16. That headroom is what lets a threshold near 0.3 hold at a false match rate of one in a million. It is why production face embeddings are 512-dimensional.

Common mistakes

Normalising at training and forgetting at inference. Cosine on unnormalised vectors is not cosine. Errors are subtle and scale with image brightness, which makes them look like a data problem.

Comparing embeddings from two different models. Face embedding spaces are arbitrary up to rotation. Model A's vector and model B's vector have no relationship. Re-enrol everyone when you change models; there is no conversion.

Applying the margin at inference. The margin belongs to the loss and only the loss. Applying it when scoring lowers every genuine score and destroys your threshold.

Setting the margin too high, too early. With m = 0.5 from step zero, many models diverge. Standard practice is a warm-up: train briefly with m = 0, then ramp it. If your loss goes to NaN in the first epoch, this is the first thing to check.

Missing the clamp before acos. Floating-point error pushes a cosine to 1.0000001, acos returns NaN, and the whole batch's gradient becomes NaN. The clamp(-1 + 1e-7, 1 - 1e-7) above exists for exactly this.

Try it yourself

Set m = 0.0 and rerun. The two losses become identical, which confirms the margin is the only difference. Then sweep s over 1, 8, 32 and 64, and watch how much the loss changes when the margin is applied. You will see why the scale has to be large before the margin does anything at all.

What to learn next

  • Embeddings — the same idea, applied to words instead of faces.
  • Loss functions — where the margin lives and what it does to gradients.
  • Vector databases — searching millions of embeddings without comparing every one.

Researcher — Mathematics and papers.

From softmax to angular margin

Standard softmax cross-entropy on the final layer is

$$ L = -\log \frac{e^{W_{y}^{\top} x + b_{y}}}{\sum_{j=1}^{C} e^{W_{j}^{\top} x + b_{j}}} $$

$x \in \mathbb{R}^d$ is the embedding, $W_j$ the $j$-th class weight column, $b_j$ its bias, $C$ the number of identities, $y$ the true label. The decision boundary depends on both magnitude and angle, so nothing constrains the geometry of the learned features.

Setting $b = 0$ and normalising both sides, $\lVert W_j \rVert = 1$ and $\lVert x \rVert = s$, reduces the logit to $s\cos\theta_j$ where $\theta_j$ is the angle between $x$ and $W_j$. All information now lives on a hypersphere of radius $s$.

ArcFace (Deng, Guo, Xue, Zafeiriou, CVPR 2019, arxiv.org/abs/1801.07698) adds an additive angular margin to the target logit:

$$ L_{\text{arc}} = -\log \frac{e^{s\cos(\theta_y + m)}}{e^{s\cos(\theta_y + m)} + \sum_{j \neq y} e^{s\cos\theta_j}} $$

$m$ is the margin in radians, paper default $0.5$; $s$ the scale, paper default $64$. Because $\cos$ is monotonically decreasing on $[0, \pi]$, adding $m$ inside strictly lowers the target logit, so the network must reduce $\theta_y$ below what the unmargined objective would demand.

The margin family, compared

LossLogit for the true classMargin acts on
Softmax$\lVert W_y \rVert \lVert x \rVert \cos\theta_y$nothing
SphereFace (Liu et al., 2017)$s\cos(m\theta_y)$angle, multiplicative
CosFace (Wang et al., 2018)$s(\cos\theta_y - m)$cosine, additive
ArcFace (Deng et al., 2019)$s\cos(\theta_y + m)$angle, additive

ArcFace's argument for the additive angular form is geometric: the margin is a constant arc length on the hypersphere, so the induced separation is uniform across the space. CosFace's margin, being additive in cosine, produces angular separation that varies with $\theta$. SphereFace's multiplicative margin is difficult to optimise and needed an annealing schedule.

Empirically the three are close on saturated benchmarks and separate on hard ones. Do not expect large differences on LFW, which sits above 99.8% for all of them.

Why the scale is necessary

With $\lVert x \rVert = 1$, logits lie in $[-1, 1]$ and the softmax over $C$ classes cannot approach a one-hot distribution. The lower bound on the loss is roughly $\log C - 2$, and gradients stay small regardless of correctness.

Wang et al. (2017), NormFace, analyse this and derive a lower bound on $s$ as a function of $C$ and the target confidence. For $C \sim 10^5$ identities, $s$ in the range 30 to 64 is standard.

Complexity

The head is a single matrix product, $O(BCd)$ for batch $B$, embedding dimension $d$ and $C$ classes. With $C = 10^7$ identities and $d = 512$, the weight matrix alone is 20 GB in fp32, which dominates everything.

Two standard responses:

  • Partial FC (An et al., CVPR 2022, arxiv.org/abs/2203.15565) samples a subset of negative class centres per step. Sampling 10% of classes loses negligible accuracy and cuts the head's memory by an order of magnitude, which is what made ten-million-identity training practical.
  • Sub-centre ArcFace (Deng et al., ECCV 2020) gives each identity $K$ sub-centres, tolerating label noise by letting outlier images attach to a minority sub-centre instead of dragging the main one.

Evaluation, and what the numbers mean

LFW is saturated; treat any LFW number above 99.5% as uninformative. The benchmarks that still discriminate are IJB-B and IJB-C, reported as TAR at FAR $10^{-6}$ — true accept rate when only one impostor pair in a million is accepted. IJB-C provides roughly 15.6 million impostor comparisons, which is what makes a FAR of $10^{-6}$ measurable at all.

At those operating points the differences between recent methods are typically one to two percentage points, and are sensitive to training set, backbone and alignment. Compare methods only when all three are held fixed.

Embeddings are biometric data

Template inversion is a solved problem in the research literature: Mai et al. (TPAMI 2019), On the Reconstruction of Face Images from Deep Face Templates, recover recognisable faces from embeddings alone, including from models they did not have access to. Later work with generative priors has improved this considerably.

The practical consequences are concrete. An embedding database is a biometric database under GDPR Article 9 and under India's DPDP Act, and is a "biometric identifier" under Illinois BIPA. Protection strategies exist — cancellable biometrics, homomorphic matching, secure enclaves — and each costs accuracy or latency. See privacy in ML for the wider picture on what leaks out of a stored representation.

Papers

What to learn next

  • Embeddings — the same idea, applied to words instead of faces.
  • Loss functions — where the margin lives and what it does to gradients.
  • Vector databases — searching millions of embeddings without comparing every one.