Verification vs identification
Verification compares two faces and identification searches a whole database, and the same model at the same threshold is safe for one and dangerous for the other.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Verification asks "is this the same person as that one?" Identification asks "who, out of everybody we have, is this?"
Think of two different doors. At the first, a guard holds your ID card, looks at your face, and compares the two. One comparison, one answer.
At the second door there is no card. The guard flips through a folder of ten thousand photographs, hunting for your face. Same guard, same eyesight, an entirely different job.
The first door is verification. The second is identification. The gap between them is the most misunderstood thing in this whole field.
Why the difference is dangerous
The guard makes a mistake now and then. Say he wrongly matches strangers one time in ten thousand.
At the first door that is fine. One comparison, one chance to slip, one in ten thousand.
At the second door he makes ten thousand comparisons for a single visitor. Now a wrong match is close to certain on every visit. Nothing about the guard changed. The number of chances changed.
This is why a face system can unlock your phone flawlessly. The same system can be unfit for searching a city database. The same accuracy, applied to a far larger number of comparisons, produces a completely different outcome.
How it works
VERIFICATION (one to one)
your face ─┐
├─► one comparison ─► score ─► above the line? yes / no
the card ─┘
IDENTIFICATION (one to many)
your face ──► compared against EVERY stored face
│
▼
thousands of scores
│
▼
keep the highest one
│
▼
above the line? ─► "this is who you are"The extra danger sits in that "keep the highest one" step. Out of thousands of strangers, one of them will resemble you more than the rest, purely by chance. The system then reports that stranger with total confidence.
The two ways to be wrong
Every face system has two failure modes, and they pull against each other.
Letting in the wrong person. The score was high enough, and it should not have been.
Rejecting the right person. Your own face scored below the line, so the door stayed shut.
Move the line up and strangers stop getting in, while more genuine users get locked out. Move it down and genuine users sail through, while strangers do too. There is no setting that removes both.
So the question is never "is this system accurate?" It is "which mistake do we prefer, and how often can we afford it?" A phone lock and a border gate answer that question very differently.
Where you have already seen this
- Verification: unlocking your phone, a bank app matching your selfie to your ID, an airport e-gate reading your passport chip.
- Identification: photo apps grouping people by face, a shop's watchlist system, police searching a mugshot database.
Notice the second list is where nearly all the public controversy sits. That is not a coincidence.
What is honestly hard here
A vendor quoting a single accuracy figure has told you almost nothing. Accurate at which of the two tasks? Against how large a database? At which setting of the line? On which population?
Without those four things, the number is marketing. Ask for the two error rates, at a stated line, on a database the size of yours. If the answer does not arrive in that form, the vendor has not measured it.
Remember this
- Verification is one comparison. Identification is one comparison per stored person.
- Errors that are rare in one comparison become common across thousands.
- Every system trades letting the wrong person in against shutting the right person out.
What to learn next
- Model evaluation — precision, recall, ROC and where thresholds come from.
- Fairness metrics — what happens when one threshold meets several populations.
- Speaker identification — the identical 1:1 and 1:N split, applied to voices.
Developer — Code and libraries.
Setup
pip install numpyRun against numpy 1.26. The scores below are simulated so the code runs anywhere, and the shapes of the distributions are chosen to resemble a decent modern model.
Measuring both tasks with one model
import numpy as np
rng = np.random.default_rng(7)
# Simulated cosine scores from a decent face model.
genuine = np.clip(rng.normal(0.62, 0.11, 200_000), -1, 1) # pairs that ARE the same person
impostor = np.clip(rng.normal(0.06, 0.06, 2_000_000), -1, 1) # pairs that are NOT
g_sorted, i_sorted = np.sort(genuine), np.sort(impostor)
def fmr(t): # false match rate: a stranger accepted
return 1.0 - np.searchsorted(i_sorted, t, "left") / i_sorted.size
def fnmr(t): # false non-match rate: the right person rejected
return np.searchsorted(g_sorted, t, "left") / g_sorted.size
print("VERIFICATION (1:1, 'is this the person on the card?')")
print(f"{'threshold':>10} {'FMR':>12} {'FNMR':>10}")
for t in (0.20, 0.26, 0.32, 0.38, 0.44):
print(f"{t:>10.2f} {fmr(t):>12.6f} {fnmr(t):>10.4f}")
grid = np.linspace(0.0, 0.9, 9001)
F, N = np.array([fmr(t) for t in grid]), np.array([fnmr(t) for t in grid])
k = int(np.argmin(np.abs(F - N)))
print(f"\nEER {0.5*(F[k]+N[k]):.5f} at threshold {grid[k]:.4f}")
for target in (1e-3, 1e-4, 1e-5):
j = int(np.argmax(F <= target))
n_above = int(round(F[j] * i_sorted.size))
print(f"FMR <= {target:g}: threshold {grid[j]:.4f} FNMR {N[j]:.4f} "
f"(estimated from {n_above} impostor pairs above it)")
# ---- Same model, same threshold, now used for 1:N search.
j = int(np.argmax(F <= 1e-4))
t_op, fmr_op = float(grid[j]), float(F[j])
print(f"\nIDENTIFICATION (1:N, 'who is this?') at threshold {t_op:.4f}")
print(f"{'gallery N':>12} {'predicted FPIR':>16} {'measured FPIR':>16}")
for size in (100, 10_000, 1_000_000):
hits = 0
for _ in range(200): # 200 searches where the person is NOT enrolled
best = -1.0
for start in range(0, size, 200_000): # chunked, so this fits in a laptop's RAM
n = min(200_000, size - start)
best = max(best, float(np.clip(rng.normal(0.06, 0.06, n), -1, 1).max()))
hits += best >= t_op
print(f"{size:>12,} {1-(1-fmr_op)**size:>16.4f} {hits/200:>16.4f}")VERIFICATION (1:1, 'is this the person on the card?')
threshold FMR FNMR
0.20 0.009765 0.0001
0.26 0.000426 0.0005
0.32 0.000007 0.0033
0.38 0.000000 0.0147
0.44 0.000000 0.0509
EER 0.00047 at threshold 0.2584
FMR <= 0.001: threshold 0.2451 FNMR 0.0003 (estimated from 1993 impostor pairs above it)
FMR <= 0.0001: threshold 0.2804 FNMR 0.0009 (estimated from 200 impostor pairs above it)
FMR <= 1e-05: threshold 0.3125 FNMR 0.0027 (estimated from 20 impostor pairs above it)
IDENTIFICATION (1:N, 'who is this?') at threshold 0.2804
gallery N predicted FPIR measured FPIR
100 0.0100 0.0100
10,000 0.6321 0.6800
1,000,000 1.0000 1.0000This run takes around a minute, mostly on the million-strong gallery. The seed makes it reproducible, but note that 200 trials is a small sample: the measured FPIR for the 100-person gallery carries roughly a 70 percent relative error.
Reading that output, line by line
The verification table looks superb. At threshold 0.32, seven impostor pairs in a million are accepted and 0.33% of genuine users get a retry. For a phone lock that is more than good enough.
The EER of 0.00047 is a benchmark number, not an operating point. Equal error rate is where the two mistakes are equally likely. Almost no real system runs there. A door wants FMR far below FNMR; a photo-grouping feature wants the reverse.
Notice the shrinking sample counts. FMR of $10^{-5}$ was estimated from 20 impostor pairs. Twenty. That estimate has a standard error of roughly 22%, so quoting it to two significant figures is dishonest. If you want to claim a false match rate of one in a million, you need tens of millions of impostor comparisons to measure it. This is why IJB-C ships 15.6 million impostor pairs.
Then the identification block, and this is the whole lesson. The same threshold that gave one wrong match in ten thousand gives a false alarm on 68% of searches against a ten-thousand-person gallery. Against a million, on every single search.
Predicted 0.6321, measured 0.6800, and the gap is real. The prediction uses the measured FMR of exactly $1.0 \times 10^{-4}$; the true tail probability of the underlying distribution at that threshold is closer to $1.2 \times 10^{-4}$. A 20% error in an FMR estimate becomes a visible error in the 1:N prediction. Estimation noise does not stay small when you exponentiate it.
The vocabulary, since it is inconsistent everywhere
| Term | Also written | Means |
|---|---|---|
| FMR | FAR, FPR | Impostor comparison accepted |
| FNMR | FRR, FNR | Genuine comparison rejected |
| FPIR | — | A 1:N search returns a candidate when the person is not enrolled |
| FNIR | — | A 1:N search misses the person who is enrolled |
| Rank-1 | Top-1 | The correct person is the highest-scoring candidate |
| CMC | — | Rank-$k$ accuracy plotted against $k$ |
NIST uses FMR/FNMR for 1:1 and FPIR/FNIR for 1:N. Vendor material tends to use FAR/FRR for everything. When someone quotes FAR, ask which of the four they mean.
Common mistakes
Reporting accuracy. With a balanced test set, "accuracy" is the average of two rates you actually need separately. Report both rates at a stated threshold, always.
Reusing the 1:1 threshold for 1:N. The output above is what that costs. A 1:N system needs its threshold set from measured FPIR at your gallery size, not derived from a verification benchmark.
Testing 1:N on a small gallery. A watchlist demo on 500 people tells you nothing about 500,000. FPIR grows with gallery size; you must test at production scale or at least extrapolate deliberately.
Ignoring rank when a human reviews the output. If an analyst inspects the top five candidates, rank-5 accuracy and top-5 FPIR are your metrics, not rank-1. Measure the thing your workflow actually consumes.
Pooling every demographic group into one number. A single global FMR hides per-group rates that can differ by two orders of magnitude. That is measured in the next-but-one lesson, and it changes what these numbers mean.
Try it yourself
Set the impostor distribution's standard deviation to 0.09 instead of 0.06, leaving everything else alone. This models a weaker or lower-quality-image model. Recompute the threshold for FMR of $10^{-4}$, then rerun the identification block. Watch how a modest widening of the impostor tail wipes out the usable gallery size.
What to learn next
- Model evaluation — precision, recall, ROC and where thresholds come from.
- Fairness metrics — what happens when one threshold meets several populations.
- Speaker identification — the identical 1:1 and 1:N split, applied to voices.
Researcher — Mathematics and papers.
Formal definitions
Let $s(a, b) \in \mathbb{R}$ be the comparison score between two biometric samples, and $\tau$ the decision threshold.
Verification (1:1). For a claimed identity, accept when $s \geq \tau$.
$$ \text{FMR}(\tau) = \Pr\big[s \geq \tau \mid \text{different subjects}\big], \qquad \text{FNMR}(\tau) = \Pr\big[s < \tau \mid \text{same subject}\big] $$
The DET curve plots FNMR against FMR on a normal-deviate scale; the ROC plots $1 - \text{FNMR}$ against FMR. Equal error rate is the $\tau$ where $\text{FMR}(\tau) = \text{FNMR}(\tau)$.
Open-set identification (1:N). Given a gallery $G$ of size $N$, compute $s_i$ for each enrolled subject and take $s_{\max} = \max_i s_i$.
$$ \text{FPIR}(\tau, N) = \Pr\big[s_{\max} \geq \tau \mid \text{probe not in } G\big], \qquad \text{FNIR}(\tau, N) = \Pr\big[s_{\text{mate}} < \tau \ \text{or}\ \text{rank} > R\big] $$
If the $N$ non-mate scores were independent, then
$$ \text{FPIR}(\tau, N) = 1 - \big(1 - \text{FMR}(\tau)\big)^{N} $$
Independence is the assumption that fails in practice. Real galleries contain relatives, repeated enrolments of the same person, and demographic clusters, all of which correlate the scores. NIST's own analysis, Relating 1:1 and 1:N False Positive Rates (part of the FRVT demographics material), finds the naive formula usefully approximate for algorithm comparison and unreliable as an absolute prediction. Measure FPIR directly on a gallery drawn from your deployment population.
Closed set versus open set
Closed-set identification assumes the probe is enrolled, so the metric is rank-$k$ accuracy and the CMC curve. It is the easier problem and the less useful one.
Almost every deployed 1:N system is open-set: most people who walk past a camera are not on the watchlist. Reporting rank-1 accuracy for an open-set deployment is a category error, because rank-1 says nothing about what happens for the overwhelming majority of probes that have no mate at all.
The threshold is a policy decision
At a fixed $\tau$, the ratio of the two error types is a value judgement about which harm is worse. Expected cost makes this explicit:
$$ \mathbb{E}[C] = \pi\, C_{\text{FN}}\, \text{FNMR}(\tau) + (1 - \pi)\, C_{\text{FP}}\, \text{FMR}(\tau) $$
$\pi$ is the prior probability that a presentation is genuine, $C_{\text{FN}}$ and $C_{\text{FP}}$ the costs of each error. In a watchlist scenario $\pi$ is minuscule, so the false-positive term dominates regardless of how good the model is. This is the base-rate problem, and no improvement in FMR removes it — it only shifts the gallery size at which the system becomes unusable.
Score normalisation
Because $s_{\max}$ over $N$ scores has a distribution that depends on $N$, comparing raw maxima across galleries of different sizes is invalid. Two standard corrections:
- Z-normalisation and T-normalisation, borrowed from speaker recognition, rescale a probe's score by the mean and variance of its scores against a cohort.
- Extreme value modelling (Scheirer et al., TPAMI 2011, Meta-Recognition) fits a Weibull to the non-match tail and converts $s_{\max}$ into a calibrated probability of being a true match.
The second is the principled route for open-set operation and is under-used in practice.
Confidence intervals on very low error rates
Estimating $\text{FMR} = 10^{-6}$ requires enough impostor comparisons for the count above threshold to be non-trivial. With $n$ independent comparisons and $k$ false matches, a Poisson interval is adequate: the relative standard error is roughly $1/\sqrt{k}$.
Ten false matches gives about 32% relative error. One hundred gives 10%. This is why serious evaluations are built from tens of millions of pairs, and why a benchmark reporting $10^{-6}$ FMR from a hundred thousand comparisons is reporting noise.
Note also that impostor pairs from a fixed set of subjects are not independent — each subject appears in many pairs — so the effective sample size is smaller than the pair count. Bootstrapping over subjects, not over pairs, is the correct resampling unit.
Standards and reference reading
- ISO/IEC 19795-1 — biometric performance testing and reporting; the source of the FMR/FNMR vocabulary.
- ISO/IEC 2382-37 — biometrics vocabulary.
- Grother et al., NIST FRVT ongoing reports — pages.nist.gov/frvt — the only large-scale, independent, continuously updated evaluation of commercial algorithms at both 1:1 and 1:N.
- Maze et al., IARPA Janus Benchmark-C, ICB 2018 — the IJB-C protocol, 3,531 subjects, 19,557 genuine and 15,638,932 impostor comparisons.
What to learn next
- Model evaluation — precision, recall, ROC and where thresholds come from.
- Fairness metrics — what happens when one threshold meets several populations.
- Speaker identification — the identical 1:1 and 1:N split, applied to voices.