Face anti-spoofing and liveness
Liveness detection asks whether a real person is in front of the camera at all, and it is measured with a separate pair of error rates where only the worst attack type counts.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Liveness detection checks that a real, living person is in front of the camera, rather than a photo of one.
Think of a security guard who has your photograph and compares it to the face at the door. He is excellent at comparing. Now somebody walks up holding a printed picture of you at arm's length in front of their own head.
The guard's comparison still succeeds. The photo matches the photo. He was never asked whether the thing in front of him was a person.
That missing question is the whole of liveness detection.
Why it exists
Face recognition and liveness detection are separate models solving separate problems, and it is worth being precise about why.
A recognition model is trained on pairs of images and asked whether they show the same person. It has never been shown a printed photo and told to reject it. Feeding it a good photo of the right person gives a confident, correct answer to the question it was asked.
So a face-matching system with no liveness check can be defeated by a printout. This has been demonstrated over and over, on shipping products.
The attacks, from cheapest to worst
Each of these is called a presentation attack — something held up to the camera to fool it.
- A printed photo. Costs a few rupees. Blocked by nearly every system now.
- A photo with the eyes cut out, worn as a mask so real eyes blink through.
- A video replayed on a phone screen. Adds motion and blinking.
- A high-quality curved print wrapped to give the face depth.
- A silicone or resin mask, moulded from a real face. Costs a great deal and defeats most defences.
The list is ordered by cost, and defence quality tends to follow the same order. A system that stops printouts and fails against masks is the normal case, not a scandal. You must know which one you bought.
How it works
camera
│
▼
find the face (detection)
│
├──────────────► is this a live person? ──► if no, STOP
│ (liveness)
▼
is it the right person? (recognition)
│
▼
let them inLiveness sits before recognition and can veto it. What the liveness model looks for depends on the design.
Passive checks study a single image or a short clip without asking the user to do anything. Screens have reflections and pixel grids. Prints have paper texture. Real skin scatters light in a way flat surfaces do not.
Active checks ask you to act: turn your head, blink, read a number aloud, or follow a moving dot. Harder to fake, and slower and more annoying for the user.
Hardware checks use a second sensor. An infrared camera, or a depth sensor that projects dots and measures how they land. This is what makes phone face unlock hard to beat with a photograph. It is the most reliable option by some distance.
The newer problem
There is an attack that does not go in front of the camera at all.
Instead of holding something up, the attacker replaces the camera. Software on the device pretends to be a camera and feeds in a video the attacker made. This is called an injection attack — the fake image is inserted into the data stream directly.
Every liveness model in the world is looking at the picture. If the picture is a well-made fake, the model sees a perfect live person. As far as it can tell, it is one. The blinks are there. The head turns are there.
Defending against this is not a computer-vision problem. It needs device attestation, secure camera paths, and checks that the app has not been tampered with. Anyone who sells you a liveness model as protection against injection has misunderstood the attack.
Where you have already seen this
- The prompt to turn your head during a bank's video verification.
- Your phone refusing to unlock from a printed photo of you.
- An exam proctoring tool checking you are still at the desk.
- A ride-hailing app asking a driver for a selfie before a shift.
Remember this
- Recognition asks who. Liveness asks whether anyone is really there.
- Defences are ranked by the attack they stop, and masks are far harder than printouts.
- An injected fake video bypasses liveness entirely, and needs device security to stop.
What to learn next
- Adversarial attacks — the wider family of inputs designed to defeat a model.
- Model evaluation — thresholds, error trade-offs and why one number is never enough.
- Responsible deployment — shipping a system whose failures have consequences.
Developer — Code and libraries.
Setup
pip install numpyRun against numpy 1.26. The code below computes the metrics from the international standard for this task, ISO/IEC 30107-3. Getting these right matters more than any model choice, because they are how you compare vendors and how you decide the system is safe to ship.
Scoring liveness the way the standard requires
import numpy as np
rng = np.random.default_rng(3)
# A liveness model returns a "this is a real, present person" score in [0, 1].
bona_fide = np.clip(rng.normal(0.86, 0.10, 5_000), 0, 1) # genuine users at the camera
# Attacks, by species (ISO/IEC 30107-3 calls each one a Presentation Attack Instrument).
attacks = {
"printed photo": np.clip(rng.normal(0.12, 0.09, 2_000), 0, 1),
"phone replay": np.clip(rng.normal(0.24, 0.12, 2_000), 0, 1),
"paper eye-cut": np.clip(rng.normal(0.31, 0.13, 2_000), 0, 1),
"silicone mask": np.clip(rng.normal(0.63, 0.15, 2_000), 0, 1), # the expensive one
}
def bpcer(t): # real people wrongly called an attack
return float((bona_fide < t).mean())
def apcer_per_species(t): # attacks wrongly called real, one rate per species
return {k: float((v >= t).mean()) for k, v in attacks.items()}
print(f"{'thresh':>7} {'BPCER':>7} | " + " ".join(f"{k:>14}" for k in attacks) + f" {'APCER':>7} {'ACER':>7}")
for t in (0.30, 0.40, 0.50, 0.60, 0.70):
per = apcer_per_species(t)
ap = max(per.values()) # ISO 30107-3: APCER is the WORST species, never the average
bp = bpcer(t)
print(f"{t:>7.2f} {bp:>7.4f} | " + " ".join(f"{per[k]:>14.4f}" for k in attacks)
+ f" {ap:>7.4f} {(ap+bp)/2:>7.4f}")
# The number a vendor quotes, and the number you should ask for instead.
mean_ap = {t: float(np.mean(list(apcer_per_species(t).values()))) for t in (0.5,)}
print(f"\nAt threshold 0.50 the honest APCER is {max(apcer_per_species(0.5).values()):.4f}")
print(f"The average across species is {mean_ap[0.5]:.4f} <- flattering, and not the standard")
# Operating point people actually buy: fix the user pain, then read the security.
grid = np.linspace(0, 1, 2001)
bps = np.array([bpcer(t) for t in grid])
for budget in (0.01, 0.05):
j = int(np.argmax(bps >= budget)) - 1
per = apcer_per_species(float(grid[j]))
print(f"\nBPCER capped at {budget:.0%}: threshold {grid[j]:.4f}, BPCER {bps[j]:.4f}")
for k, v in per.items():
print(f" APCER {k:>14}: {v:.4f}") thresh BPCER | printed photo phone replay paper eye-cut silicone mask APCER ACER
0.30 0.0000 | 0.0210 0.3075 0.5255 0.9905 0.9905 0.4953
0.40 0.0000 | 0.0010 0.0890 0.2440 0.9475 0.9475 0.4738
0.50 0.0000 | 0.0000 0.0145 0.0760 0.8140 0.8140 0.4070
0.60 0.0058 | 0.0000 0.0000 0.0135 0.5795 0.5795 0.2927
0.70 0.0574 | 0.0000 0.0000 0.0005 0.3250 0.3250 0.1912
At threshold 0.50 the honest APCER is 0.8140
The average across species is 0.2261 <- flattering, and not the standard
BPCER capped at 1%: threshold 0.6220, BPCER 0.0096
APCER printed photo: 0.0000
APCER phone replay: 0.0000
APCER paper eye-cut: 0.0110
APCER silicone mask: 0.5275
BPCER capped at 5%: threshold 0.6935, BPCER 0.0496
APCER printed photo: 0.0000
APCER phone replay: 0.0000
APCER paper eye-cut: 0.0015
APCER silicone mask: 0.3395Reading that output, which is the point of the lesson
APCER is the maximum over attack types, not the mean. That is what the standard specifies, and the reason is visible in the numbers. At threshold 0.50 the honest APCER is 0.8140. Average the four species and you get 0.2261. The same model, the same data, and a figure nearly four times better because three easy attacks dilute one hard one.
An attacker does not sample uniformly from attack types. They pick the one that works. The maximum is the only number that describes what they will do.
The whole table is the silicone mask. Printouts and replays are handled at any reasonable threshold. The mask column is what determines the operating point, and it never gets good. This is a realistic shape for a passive, camera-only model.
BPCER at 1% still leaves the mask at 0.5275. Tighten the user pain to one rejection in a hundred, and better than half of mask attacks still succeed. Pushing to 5% rejection — which most product owners will refuse — buys the mask rate down to 0.3395. Camera-only liveness does not solve masks, and no threshold choice makes it.
ACER is a summary, and a poor one. Averaging APCER and BPCER treats a locked-out customer and a defeated security control as equally important. They are not. Quote the pair at a fixed BPCER; use ACER for a leaderboard and nothing else.
The vocabulary, in the standard's own terms
| Term | Meaning |
|---|---|
| Bona fide presentation | A genuine user presenting themselves normally |
| PAI (Presentation Attack Instrument) | The physical object used to attack: print, screen, mask |
| PAI species | A category of instrument, tested and reported separately |
| APCER | Attacks classified as bona fide, reported as the maximum over species |
| BPCER | Bona fide presentations classified as attacks |
| ACER | Mean of APCER and BPCER; a convenience number |
| IAPMR | For full-system tests: attacks that both pass liveness and match the target identity |
Common mistakes
Reporting one APCER without saying which species were tested. An APCER of 0.001 against printouts alone is not a security claim. List the species and their individual rates.
Testing on a public dataset and shipping. Public spoof datasets are recorded with a handful of cameras under a handful of lighting conditions. A model tuned on them generalises poorly to your device population. This is the same generalisation failure covered in the next lesson.
Building liveness into the recognition model. A single model that outputs both identity and liveness is difficult to threshold sensibly and impossible to audit separately. Keep them as two models with two thresholds.
Trusting the client. If liveness runs in the browser or the app and sends back a boolean, the attacker sends back true. Score on the server, from raw frames, or accept that you have no protection at all.
Treating injection attacks as a model problem. A virtual camera driver feeding synthetic frames defeats every passive and active check, because the frames are internally consistent. The mitigations are device attestation, hardware-backed key attestation, secure camera APIs and tamper detection. Industry reporting through 2025 and 2026 has consistently found injection to be the fastest-growing attack vector against remote identity verification.
Try it yourself
Add a fifth species — a "3D-printed resin mask" scoring around 0.75 with a spread of 0.12 — and rerun. Watch what a single new attack type does to the APCER column at every threshold, while three of the four original species remain at zero. That is the argument for reporting the maximum, made concrete.
What to learn next
- Adversarial attacks — the wider family of inputs designed to defeat a model.
- Model evaluation — thresholds, error trade-offs and why one number is never enough.
- Responsible deployment — shipping a system whose failures have consequences.
Researcher — Mathematics and papers.
Standardised metrics
ISO/IEC 30107-3 defines PAD performance for subsystem evaluation. For a threshold $\tau$, with PAI species indexed by $s$:
$$ \text{APCER}_s(\tau) = \frac{1}{N_s}\sum_{i=1}^{N_s} \mathbb{1}\big[ f(x_i^{(s)}) \geq \tau \big], \qquad \text{APCER}(\tau) = \max_s \text{APCER}_s(\tau) $$
$$ \text{BPCER}(\tau) = \frac{1}{N_{bf}}\sum_{i=1}^{N_{bf}} \mathbb{1}\big[ f(x_i^{bf}) < \tau \big] $$
$N_s$ is the number of attack presentations of species $s$, $N_{bf}$ the number of bona fide presentations, and $f$ the liveness score. The maximum over species is normative, not a convention.
For full-system evaluation the standard uses IAPMR — impostor attack presentation match rate — the proportion of attacks that both survive PAD and produce a biometric match against the targeted identity. IAPMR is the number that describes end-to-end risk; APCER describes the PAD subsystem alone.
BPCER@APCER=x% and APCER@BPCER=x% are the reporting forms that carry information. A single ACER figure does not.
Method families
Texture and quality cues. Early work used LBP, colour-space statistics and image-quality measures (Boulkenafet et al., 2015; Galbally et al., 2014). Cheap, interpretable, and brittle across capture devices.
rPPG. Remote photoplethysmography extracts the faint periodic colour change caused by blood flow. A print or mask has no pulse. Liu et al., CVPR 2018, Learning Deep Models for Face Anti-Spoofing, supervise a network with both a depth map and an rPPG signal, using auxiliary physical targets rather than a binary label. rPPG degrades badly with compression, low frame rate and motion.
Depth supervision. Predicting a pseudo-depth map — flat for attacks, facial geometry for bona fide — is a stronger training signal than a binary label, and it survives the loss of a real depth sensor at deployment.
Multi-channel input. George et al., IEEE TIFS 2019, Biometric Face Presentation Attack Detection with Multi-Channel CNN (arxiv.org/abs/1909.08848), combine colour, depth, near-infrared and thermal. This is the strongest published direction against masks, and it requires hardware the deployment must actually have.
Domain generalisation. Because the failure mode is cross-dataset, the field has moved toward meta-learning and domain-adversarial training over multiple capture domains. Progress is real and modest; leave-one-dataset-out remains far below within-dataset performance.
Benchmarks and their limits
- OULU-NPU (Boulkenafet et al., 2017) — four protocols isolating unseen lighting, unseen PAI, unseen camera, and all three.
- SiW-M (Liu et al., 2019) — 13 spoof types, designed for leave-one-type-out evaluation.
- CelebA-Spoof (Zhang et al., ECCV 2020) — over 600k images with rich annotation.
- CASIA-SURF — multi-modal RGB, depth and infrared.
Protocol 4 of OULU-NPU is the one to read. Within-protocol numbers on the easier protocols are not predictive of deployment.
Independent certification against ISO/IEC 30107-3 is performed by accredited labs, most visibly iBeta, at Level 1 (low-cost instruments) and Level 2 (higher-effort instruments including custom masks). A Level 2 confirmation letter is a meaningful signal about mask resistance; a self-reported APCER is not.
Injection attacks
The threat model that PAD does not address is injection: substituting the capture stream rather than the presented artefact. A virtual camera, a hooked capture API, an emulator, or a manipulated network payload delivers frames that are internally self-consistent, so texture, depth and rPPG cues are all reproducible by the generator.
Reporting through 2025 and 2026 has consistently placed injection among the fastest-growing vectors against remote identity verification, and the World Economic Forum's 2026 material on digital identity verification documents virtual-camera injection defeating a broad range of active liveness implementations.
Mitigations sit outside computer vision:
- Hardware-backed key attestation and platform integrity attestation.
- Secure camera paths and trusted execution environments.
- Server-side capture with signed frames and timestamps.
- Detection of emulators, rooted devices and hooked APIs.
- Cryptographic binding of the capture session to the account.
The correct framing is that PAD defends the sensor's front side and attestation defends its back side. Deploying one without the other leaves an open path.
Papers and standards
- ISO/IEC 30107-1:2016 and 30107-3:2023 — PAD framework and testing methodology.
- Galbally et al., Image Quality Assessment for Fake Biometric Detection, IEEE TIP 2014
- Boulkenafet et al., Face Anti-Spoofing Based on Color Texture Analysis, ICIP 2015
- Liu et al., Learning Deep Models for Face Anti-Spoofing, CVPR 2018 — arxiv.org/abs/1803.11097
- George et al., Biometric Face Presentation Attack Detection with Multi-Channel CNN, IEEE TIFS 2019 — arxiv.org/abs/1909.08848
- Zhang et al., CelebA-Spoof, ECCV 2020 — arxiv.org/abs/2007.12342
What to learn next
- Adversarial attacks — the wider family of inputs designed to defeat a model.
- Model evaluation — thresholds, error trade-offs and why one number is never enough.
- Responsible deployment — shipping a system whose failures have consequences.