Deepfake detection
Deepfake detectors find traces the generator left behind, they score beautifully on the fakes they were trained on, and they collapse on fakes made by a method they have never seen.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A deepfake detector looks for the marks a fake-making tool left behind, rather than judging whether the face looks real.
Think of spotting a counterfeit note. You do not decide by looking at the picture of Gandhi. You tilt it for the watermark, feel the raised print, hold it to the light for the thread.
A skilled forger fixes whichever mark you are checking. So the check has to keep moving, and a check that worked last year can fail this year.
Deepfake detection is that same race, and the forger has been winning ground.
Why it is hard, stated plainly
Every fake-making method leaves its own marks. Train a detector on fakes from one method and it learns that method's marks extremely well.
Then somebody uses a different method. The marks are different. The detector, which never learned to judge realness in general, has nothing to go on. Accuracy falls off a cliff — sometimes below a coin toss.
This is not a bug that better engineering removes. It is the shape of the problem. New generation methods appear faster than detectors can be retrained, and a detector's training data is always about the past.
Where the marks come from
Building marks. A generator builds a face from a small grid and enlarges it step by step. That enlarging leaves a regular pattern, like the weave in cheap cloth. It is invisible to you and measurable in seconds.
Joining marks. A face swap pastes a generated face into a real video frame. Around the join, the noise, sharpness and colour do not quite agree. Skin from one camera, background from another.
Behaviour marks. Early fakes blinked at the wrong rate. Teeth stayed identical between frames. A head turned without the neck following. Each of these got fixed once it was published.
real video frame
│
▼
a generated face is pasted in
│
▼
detector examines: the weave pattern in the fake region
the seam where the two sources meet
whether the motion is physically sensible
│
▼
"generated" / "genuine" (with a confidence)The thing that ruins it
Here is the awkward part. Those marks are delicate.
Upload a video anywhere and it gets re-encoded and resized. Screenshot it and it is re-encoded again. Each step scrubs away exactly the faint traces the detector needs.
So a detector can score brilliantly on a clean generated file. Send that clip through a messaging app. Now it is close to useless. The fakes people actually see are the compressed ones.
The other direction
Because catching fakes afterwards is losing, the industry is putting effort into the opposite approach: proving what is real.
The idea is to attach a signed record to a picture as it is captured. The record lists the camera and every edit since. Anyone can check the signature. A file with no record is not proven fake, but a file with a valid record is proven genuine.
This is called content provenance, and the main open standard for it is C2PA. Some cameras and phones now sign images at capture. It changes the question from "can we spot the fake?" to "can this file prove where it came from?"
Where you have already seen this
- The "AI-generated" label on posts in social apps.
- News organisations verifying a video before broadcasting it.
- Banks refusing a video call as proof of identity.
- Content credentials shown in some photo editing tools.
What is honestly hard here
Be careful with detector output on a single file. A confident score from a detector that has not been tested on that generation method is not evidence.
Real verification work still leans on the boring things. Where did this file come from? Who posted it first? Does the metadata survive scrutiny? Is there an independent recording of the same event? The model is one input among several, and never the last word.
Remember this
- Detectors find generator traces, not realness, so an unfamiliar generator defeats them.
- Compression and resizing destroy the traces, and real-world files are always compressed.
- Proving what is genuine, using signed provenance records, is the more durable direction.
What to learn next
- GAN — the generator architecture whose upsampling leaves the traces above.
- Diffusion models — the newer family, with fainter traces and a different failure profile.
- Adversarial attacks — why every detector is evadable by a determined attacker.
Developer — Code and libraries.
Setup
pip install numpy opencv-python scikit-learnRun against numpy 1.26, opencv-python 4.10 and scikit-learn 1.7. Two demonstrations follow: a real generator artefact you can measure, and the generalisation failure that limits every detector.
Part 1: an upsampling artefact you can measure
Generators build images by repeatedly enlarging a small grid. Nearest-neighbour enlargement by a factor of four multiplies the image spectrum by a kernel that is exactly zero at frequencies that are multiples of one quarter. That is a hard, checkable fingerprint.
import numpy as np
import cv2
rng = np.random.default_rng(11)
N = 256
# A camera-like image: real photographs have smoothly falling frequency content.
f = np.fft.fftfreq(N); fx, fy = np.meshgrid(f, f, indexing="xy")
r = np.sqrt(fx**2 + fy**2); r[0, 0] = 1e-6
sp = rng.normal(size=(N, N)) + 1j * rng.normal(size=(N, N))
camera = np.real(np.fft.ifft2(sp / r))
camera = (camera - camera.min()) / np.ptp(camera)
# A generator-like image: pixels built by repeating a small grid, twice over.
def upsample(x, k=2):
return np.repeat(np.repeat(x, k, axis=0), k, axis=1)
built = upsample(upsample(camera[::4, ::4])) # net factor of 4
def column_energy(img):
"""Mean spectral magnitude at each horizontal frequency, averaged over vertical ones."""
S = np.abs(np.fft.fftshift(np.fft.fft2(img - img.mean())))
return (S / S.mean()).mean(axis=0)
ec, eb = column_energy(camera), column_energy(built)
print(f"{'freq (cycles/px)':>17} {'camera':>9} {'built':>9}")
for u in (0.1875, 0.2344, 0.2500, 0.2656, 0.3125, 0.4844, 0.4961):
j = N // 2 + int(round(u * N))
print(f"{u:>17.4f} {ec[j]:>9.4f} {eb[j]:>9.4f}")
def lattice_ratio(img, k=4):
"""Energy exactly on the multiples-of-1/k grid, against energy two bins off it."""
e = column_energy(img)
on = e[N // 2::N // k][1:] # 1/4, 2/4 ... skipping DC
off = e[N // 2 + 2::N // k][1:]
return float(on.mean() / off.mean())
print(f"\nlattice ratio, camera-like : {lattice_ratio(camera):.4f}")
print(f"lattice ratio, built : {lattice_ratio(built):.4f} <- a hard spectral zero")
ok, buf = cv2.imencode(".jpg", (built * 255).astype(np.uint8),
[int(cv2.IMWRITE_JPEG_QUALITY), 55])
jpeg = cv2.imdecode(buf, cv2.IMREAD_GRAYSCALE).astype(float) / 255.0
print(f"built, then JPEG quality 55: {lattice_ratio(jpeg):.4f}")
small = cv2.resize(built, (231, 231), interpolation=cv2.INTER_AREA)
back = cv2.resize(small, (N, N), interpolation=cv2.INTER_LINEAR)
print(f"built, then resized 256->231->256: {lattice_ratio(back):.4f}") freq (cycles/px) camera built
0.1875 1.0236 0.8493
0.2344 0.8766 0.2795
0.2500 0.8195 0.0000
0.2656 0.7995 0.2534
0.3125 0.7126 0.5675
0.4844 0.5188 0.1879
0.4961 0.5019 0.0618
lattice ratio, camera-like : 1.0821
lattice ratio, built : 0.0000 <- a hard spectral zero
built, then JPEG quality 55: 0.8735
built, then resized 256->231->256: 0.6951Reading it
The camera image has energy at every frequency; the built image has exactly zero at 0.25. That zero is the null of the upsampling kernel, and it is not statistical. It is arithmetic. On raw generator output this fingerprint is close to a perfect detector.
JPEG at quality 55 takes the ratio from 0.0000 to 0.8735. The fingerprint is essentially gone after one ordinary save. A single resize takes it to 0.6951. Anything that touches the pixel grid — a messaging app, a screenshot, a re-upload — degrades it.
This is the honest summary of frequency-based detection. It is excellent on files that came straight from a generator, and it is weak on the files people actually encounter. Published work on this artefact family includes Wang et al., CVPR 2020, CNN-generated images are surprisingly easy to spot... for now, and Frank et al., ICML 2020, Leveraging Frequency Analysis for Deep Fake Image Recognition, both of which report exactly this compression sensitivity.
Part 2: the generalisation failure
Two features per image: how strong the upsampling lattice trace is, and how visible the blending seam is. Different generator families leave different combinations.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
rng = np.random.default_rng(5)
# Two features per image: [upsampling-lattice score, blending-seam score].
def sample(n, mean):
return rng.normal(mean, 0.35, (n, 2))
real = sample(3000, [0.0, 0.0])
family_A = sample(3000, [2.2, 0.2]) # a generator that leaves a strong lattice fingerprint
family_B = sample(3000, [-0.4, 2.2]) # a different one that leaves a seam instead
X = np.vstack([real, family_A]); y = np.r_[np.zeros(3000), np.ones(3000)]
clf = LogisticRegression().fit(X, y)
def auc(fakes):
Xt = np.vstack([real, fakes]); yt = np.r_[np.zeros(len(real)), np.ones(len(fakes))]
return roc_auc_score(yt, clf.predict_proba(Xt)[:, 1])
print("trained on family A only")
print(f" AUC on family A (seen) : {auc(family_A):.4f}")
print(f" AUC on family B (unseen) : {auc(family_B):.4f}")
print(f" learned weights : {np.round(clf.coef_[0], 3).tolist()}")
clf2 = LogisticRegression().fit(np.vstack([real, family_A, family_B]),
np.r_[np.zeros(3000), np.ones(6000)])
def auc2(fakes):
Xt = np.vstack([real, fakes]); yt = np.r_[np.zeros(len(real)), np.ones(len(fakes))]
return roc_auc_score(yt, clf2.predict_proba(Xt)[:, 1])
print("\ntrained on both families")
print(f" AUC on family A: {auc2(family_A):.4f}")
print(f" AUC on family B: {auc2(family_B):.4f}")
family_C = sample(3000, [-0.5, -0.4])
print(f" AUC on family C (still unseen): {auc2(family_C):.4f}")trained on family A only AUC on family A (seen) : 1.0000 AUC on family B (unseen) : 0.3941 learned weights : [7.921, 0.955] trained on both families AUC on family A: 0.9989 AUC on family B: 0.9989 AUC on family C (still unseen): 0.0994
Reading it
AUC 1.0000 on the seen family. A detector paper reporting this is not lying. It is reporting a within-distribution result.
AUC 0.3941 on the unseen family. Below 0.5, which means the detector is anti-correlated with the truth. It has learned "high lattice score means fake", and family B has a slightly lower lattice score than real images. The detector confidently marks family B fakes as more genuine than real photographs.
Training on both families fixes both, and family C gives 0.0994. Adding data closes the gap for the families you added and does nothing for the next one. Family C leaves fewer traces than a camera does — a real phenomenon as generators improve — and the detector's decision rule inverts against it.
That is the central result of this field, reproduced in miniature. Cross-dataset AUC is the metric that matters, and papers reporting within-dataset numbers are answering a different question. Recent work in this direction includes Yermakov et al., WACV 2026, Deepfake Detection that Generalizes Across Benchmarks, evaluated across fourteen benchmarks spanning 2019 to 2025.
Common mistakes
Evaluating within-dataset. Train on FaceForensics++, test on FaceForensics++, report 0.99. This tells you nothing about deployment. Use leave-one-manipulation-out and report the worst held-out family.
Testing on uncompressed video. FaceForensics++ ships raw, c23 and c40 compression levels for a reason. Report c40. If your number is only good at raw, say so.
Ignoring the base rate. In a stream where one video in ten thousand is fake, a detector at 99% AUC still produces a flood of false accusations. This is the same base-rate arithmetic as 1:N face search.
Treating a score as evidence about a person. A detector output is a weak signal about a file. Calling a specific video fake, in public, on a model score alone, is a serious thing to get wrong.
Assuming a fixed threshold survives. Score distributions shift with every new generator and every new codec. A threshold set six months ago is measuring something else now. Monitor the score distribution, not the accuracy.
Try it yourself
Change the JPEG quality in Part 1 from 55 to 90, then to 30. Plot the lattice ratio against quality. You will find the fingerprint's survival is not gradual — there is a quality below which it is gone. That number is the operating limit of every frequency-based detector you will ever build.
What to learn next
- GAN — the generator architecture whose upsampling leaves the traces above.
- Diffusion models — the newer family, with fainter traces and a different failure profile.
- Adversarial attacks — why every detector is evadable by a determined attacker.
Researcher — Mathematics and papers.
Problem framing
Deepfake detection is binary classification under covariate shift where the shift is adversarial and unbounded. Formally, training draws from manipulation families $\mathcal{M}{\text{train}}$ and deployment draws from $\mathcal{M}{\text{test}}$, with no guarantee that $\mathcal{M}{\text{test}} \subseteq \mathcal{M}{\text{train}}$ and an adversary actively selecting $\mathcal{M}_{\text{test}}$.
This rules out standard i.i.d. generalisation guarantees. The relevant evaluation is therefore leave-one-manipulation-out cross-dataset AUC, not within-dataset accuracy.
Artefact families
Upsampling and spectral traces. Transposed convolution and nearest/bilinear upsampling impose periodic structure on the spectrum. Zhang et al. (2019), Detecting and Simulating Artifacts in GAN Fake Images, characterise it; Frank et al. (ICML 2020) build a DCT-domain classifier on it; Durall et al. (CVPR 2020) show the azimuthal spectral falloff differs systematically. Corvi et al. (ICASSP 2023) extend the analysis to diffusion models and find weaker but still present traces.
Blending boundaries. Face X-ray (Li et al., CVPR 2020, arxiv.org/abs/1912.13458) predicts the blending mask rather than a fake/real label, on the argument that almost every face-swap pipeline composites a generated region into a real frame. It is trained on self-blended images and generalises better than label-supervised baselines, because the supervision target is method-agnostic.
Identity inconsistency. ICT (Dong et al., CVPR 2022) compares inner-face and outer-face identity, exploiting that a swap changes one and not the other. Requires a reference identity, which restricts the setting.
Temporal and physiological. Blink rate, rPPG consistency across facial regions, and head-pose/landmark dynamics. All were effective and all were closed once published, which is the structural problem with any published cue.
Benchmarks
| Dataset | Year | Note |
|---|---|---|
| FaceForensics++ | 2019 | 4 (later 5) manipulation methods, three compression levels; the standard training set |
| Celeb-DF (v2) | 2020 | Higher visual quality; detectors trained on FF++ drop sharply here |
| DFDC | 2020 | Facebook challenge, ~100k clips; winning entry reached ~0.65 precision on a held-out set |
| DeeperForensics-1.0 | 2020 | Real-world perturbations applied deliberately |
| DeepfakeBench | 2023 | Unified codebase and protocol; the right starting point for comparable numbers |
The DFDC result is the sobering one. The winning solution in an open competition with a large purse reached a level that is far from usable as automated evidence, on a held-out set drawn from the same collection effort.
Generalisation strategies
- Method-agnostic supervision. Face X-ray's blending mask, and self-supervised blending augmentation (SBI, Shiohara and Yamasaki, CVPR 2022), synthesise training fakes from real images alone, removing the dependence on any particular generator.
- Frozen foundation encoders with minimal adaptation. Yermakov et al. (WACV 2026) fine-tune only LayerNorm parameters of a pre-trained vision encoder — roughly 0.03% of weights — with L2-normalised features and a metric-learning objective, reporting state-of-the-art average cross-dataset AUROC across fourteen benchmarks from 2019 to 2025. The finding that training on older, more diverse manipulation sets generalises better than training on recent ones is worth internalising.
- Reconstruction and one-class methods. Model the real-image manifold and flag deviation, avoiding the need for fake examples entirely. Weaker in absolute terms, more robust to novel generators.
Adversarial robustness
Detectors are differentiable classifiers and inherit every adversarial vulnerability. Carlini and Farid (CVPR Workshops 2020), Evading Deepfake-Image Detectors with White- and Black-Box Attacks, reduce a detector's AUC from 0.95 to below 0.1 with perturbations invisible to a human, in both white-box and black-box settings. Any deployed detector should be assumed evadable by a motivated adversary.
Provenance as the structural answer
Detection is a losing asymmetry: the defender must generalise to all future generators, the attacker needs one that works. Provenance inverts the burden.
C2PA (Coalition for Content Provenance and Authenticity) specifies cryptographically signed manifests binding capture device, edit history and assertions to the asset. Capture-time signing is shipping in some camera hardware. The limitations are real and should be stated: an unsigned file is not evidence of manipulation, signatures are stripped by most platform processing pipelines, and a signed image of a screen showing a fake is a valid signature of a false scene.
Regulation is moving in the same direction. Article 50 of the EU AI Act, applicable from 2 August 2026, requires providers of generative systems to mark synthetic output in a machine-readable form and requires deployers to disclose deep fakes. That obligation was left in place by the AI Digital Omnibus (Regulation (EU) 2026/1744), which deferred the high-risk obligations but not the transparency ones. Marking is a labelling duty on compliant actors; it does not constrain a non-compliant one.
Papers
- Rössler et al., FaceForensics++, ICCV 2019 — arxiv.org/abs/1901.08971
- Li et al., Celeb-DF, CVPR 2020 — arxiv.org/abs/1909.12962
- Li et al., Face X-ray for More General Face Forgery Detection, CVPR 2020 — arxiv.org/abs/1912.13458
- Wang et al., CNN-generated images are surprisingly easy to spot... for now, CVPR 2020 — arxiv.org/abs/1912.11035
- Frank et al., Leveraging Frequency Analysis for Deep Fake Image Recognition, ICML 2020 — arxiv.org/abs/2003.08685
- Carlini and Farid, Evading Deepfake-Image Detectors, CVPR Workshops 2020 — arxiv.org/abs/2004.00622
- Dolhansky et al., The DeepFake Detection Challenge Dataset, 2020 — arxiv.org/abs/2006.07397
- Shiohara and Yamasaki, Detecting Deepfakes with Self-Blended Images, CVPR 2022 — arxiv.org/abs/2204.08376
- Yermakov et al., Deepfake Detection that Generalizes Across Benchmarks, WACV 2026 — arxiv.org/abs/2508.06248
What to learn next
- GAN — the generator architecture whose upsampling leaves the traces above.
- Diffusion models — the newer family, with fainter traces and a different failure profile.
- Adversarial attacks — why every detector is evadable by a determined attacker.