Image Generation and Restoration

Watermarking and provenance

Watermarking hides a mark inside the pixels and provenance attaches a signed history to the file, and both are now legal requirements in the EU rather than optional extras.

On this page 10
  1. The short answer
  2. The analogy
  3. Why anyone needs this now
  4. The invisible mark
  5. The signed record
  6. This is now the law in Europe
  7. The honest position
  8. Where you have seen this
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A watermark hides a mark inside the picture itself. Provenance attaches a signed record of where the picture came from.

The analogy

Think about two ways of proving a hundred-rupee note is genuine.

One is the watermark held into the paper. You cannot see it flat on, and you cannot photocopy it. It travels with the note wherever the note goes.

The other is the serial number and the printer's records. That is a separate document, held elsewhere, that says this note was made here on this date. Tear a corner off the note and the number goes with it.

Both approaches exist for pictures, and they solve different halves of the same problem.

Why anyone needs this now

Until recently, a photograph was weak evidence but it was some evidence. Someone had been there with a camera.

That is no longer true. A convincing picture of an event that never happened takes seconds to make.

Two things follow. Fake pictures spread as if real. And real pictures get dismissed as fake, which is arguably worse, because it gives everyone an excuse.

The invisible mark

The first approach changes the pixels a tiny amount, in a pattern only a matching detector can find.

Done well, your eye sees nothing. A detector that knows the pattern reads a hidden message: which model made this, and when.

The strength of this approach is that it survives the file being copied, re-saved and posted.

Its weakness is that it does not survive very much else. You will see this measured next, with real numbers, and the results are humbling.

The signed record

The second approach adds a small sealed record to the file. Made by this camera or this tool, on this date, edited in these ways.

It is signed, the way a bank message is signed, so tampering is detectable. This is the C2PA standard, branded as Content Credentials. Camera makers and editing tools are building it in.

Its strength is that it can carry a rich, verifiable history.

Its weakness is that it is a separate parcel travelling with the file. Screenshot the image, or upload it to a site that strips extra data, and the record is gone. The picture survives; its history does not.

   invisible watermark              signed record
   ------------------              -------------
   lives IN the pixels             lives BESIDE the pixels
   survives a screenshot           does not survive a screenshot
   carries a few bits              carries a full history
   easy to damage                  easy to remove entirely
   cannot be forged easily         cryptographically signed

Neither is enough alone, which is why serious systems use both.

This is now the law in Europe

This stopped being a good idea and became a duty. Article 50 of the EU AI Act covers systems that generate images, audio, video or text. Their providers must mark the outputs in a machine-readable form. The marks must make the content detectable as AI-generated.

Those duties apply from the second of August, 2026. Systems already on the market before that date have longer. They have until the second of December, 2026 for the machine-readable marking.

The European Commission published finalised guidance on the twentieth of July, 2026. A Code of Practice on transparency for AI-generated content came with it.

If you ship a generator to European users, this is a compliance requirement, not a feature request.

The honest position

None of this is solved.

Invisible watermarks can be removed by someone who is trying. Signed records can be stripped by anyone at all. Neither proves a picture is true; they only say where the file came from.

They still help. A world where most AI images carry a mark beats one where none do. That holds even though a determined person can defeat both.

Where you have seen this

  • "Made with AI" labels on social media posts.
  • The Content Credentials icon in some photo editors and news sites.
  • Camera makers advertising signed capture in professional cameras.

Remember this

  • Watermarks hide a mark in the pixels and survive copying but not much editing.
  • Provenance attaches a signed history that is rich but easily stripped.
  • In the EU, marking generated content is now a legal duty, not an option.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy==1.26.4 scipy==1.14.1 pillow==11.0.0

A real invisible watermark, and how easily it breaks

This embeds 1024 bits into a 256-by-256 image. It compares pairs of mid-frequency DCT coefficients in each 8-by-8 block. That is roughly how classical robust watermarking works. Then it attacks it.

watermark.py
import numpy as np, io, math
from scipy.fft import dctn, idctn
from PIL import Image

rng = np.random.default_rng(0)
N, BLOCK = 256, 8
y, x = np.mgrid[0:N, 0:N]
img = np.clip(120 + 50 * np.sin(x / 17) + 30 * np.sin(y / 11) + rng.normal(0, 6, (N, N)), 0, 255)

BITS = rng.integers(0, 2, (N // BLOCK) ** 2)      # 1024 bits hidden in the picture
A, B = (3, 4), (4, 3)                             # two mid-frequency coefficients per block
STRENGTH = 12.0

def blocks(a):
    return a.reshape(N // BLOCK, BLOCK, N // BLOCK, BLOCK).swapaxes(1, 2).reshape(-1, BLOCK, BLOCK)

def unblock(b):
    return b.reshape(N // BLOCK, N // BLOCK, BLOCK, BLOCK).swapaxes(1, 2).reshape(N, N)

def embed(a, bits):
    c = dctn(blocks(a.astype(float)), axes=(1, 2), norm="ortho")
    for i, bit in enumerate(bits):
        u, v = c[i][A], c[i][B]
        hi, lo = max(u, v), min(u, v)
        # bit 1 -> coefficient A must exceed B by STRENGTH; bit 0 -> the other way round
        c[i][A], c[i][B] = (hi + STRENGTH/2, lo - STRENGTH/2) if bit else (lo - STRENGTH/2, hi + STRENGTH/2)
    return np.clip(unblock(idctn(c, axes=(1, 2), norm="ortho")), 0, 255)

def extract(a):
    c = dctn(blocks(np.asarray(a, float)), axes=(1, 2), norm="ortho")
    return (c[:, A[0], A[1]] > c[:, B[0], B[1]]).astype(int)

wm = embed(img, BITS)
psnr = 10 * math.log10(255.0 ** 2 / ((wm - img) ** 2).mean())
print(f"visual damage from the watermark: PSNR {psnr:.1f} dB, "
      f"largest pixel change {int(np.abs(wm - img).max())}")
print(f"bits recovered from the untouched file: {(extract(wm) == BITS).mean():.1%}\n")

def jpeg(a, q):
    buf = io.BytesIO(); Image.fromarray(a.astype(np.uint8)).save(buf, "JPEG", quality=q)
    return np.array(Image.open(buf).convert("L"), float)

def resize_round_trip(a, f):
    im = Image.fromarray(a.astype(np.uint8))
    return np.array(im.resize((int(N*f), int(N*f)), Image.BICUBIC).resize((N, N), Image.BICUBIC), float)

def crop_and_restore(a, px):
    return np.array(Image.fromarray(a.astype(np.uint8)).crop((px, px, N, N)).resize((N, N), Image.BICUBIC), float)

print("attack                          bits correct")
print(f"{'nothing':30s} {(extract(wm) == BITS).mean():14.1%}")
for q in (95, 75, 50, 30):
    print(f"{'JPEG quality ' + str(q):30s} {(extract(jpeg(wm, q)) == BITS).mean():14.1%}")
print(f"{'brightness +20':30s} {(extract(np.clip(wm + 20, 0, 255)) == BITS).mean():14.1%}")
print(f"{'resize to 50% and back':30s} {(extract(resize_round_trip(wm, 0.5)) == BITS).mean():14.1%}")
print(f"{'crop 8 px off two sides':30s} {(extract(crop_and_restore(wm, 8)) == BITS).mean():14.1%}")
print(f"{'crop 3 px (breaks alignment)':30s} {(extract(crop_and_restore(wm, 3)) == BITS).mean():14.1%}")
print("\n50% is the score you get from guessing.")
Output
visual damage from the watermark: PSNR 42.5 dB, largest pixel change 11
bits recovered from the untouched file: 100.0%

attack                          bits correct
nothing                                100.0%
JPEG quality 95                        100.0%
JPEG quality 75                         66.4%
JPEG quality 50                         48.4%
JPEG quality 30                         50.1%
brightness +20                         100.0%
resize to 50% and back                  61.3%
crop 8 px off two sides                 49.8%
crop 3 px (breaks alignment)            53.0%

Reading the output

The mark really is invisible. 42.5 dB and a maximum change of 11 grey levels out of 255. Nobody will spot this by eye.

It survives what it was designed to survive. Perfect recovery on the untouched file and on JPEG at quality 95. Perfect recovery after a brightness shift of +20. The scheme compares two coefficients against each other, not against a fixed threshold. Relative encodings survive global changes; absolute ones do not.

It dies at everything else. JPEG at quality 75 drops to 66.4%. At quality 50 it reaches 48.4%, which is worse than a coin toss. Resize costs it. Any crop destroys it completely. The 8-by-8 block grid shifts, and every block is read at the wrong offset.

The crop rows are the important ones. Cropping three pixels off a picture is not an attack. It is what happens when someone frames a screenshot slightly differently. Real watermarking systems spend most of their engineering effort on geometric synchronisation. That means recovering the grid alignment before reading, and it is genuinely hard.

Take the honest conclusion from this. A watermark you write in an afternoon is defeated by ordinary handling. Never mind by someone actively trying. Use a published scheme, and do not tell your users the mark is permanent.

What production systems actually do

Google SynthID covers text, image, audio and video. The text variant is open source, shipped in Hugging Face Transformers as SynthIDTextWatermarkLogitsProcessor with a Bayesian detector. The detector returns three states, watermarked, not watermarked, or uncertain, with two configurable thresholds. Google's own documentation states the text watermark survives cropping, small edits and mild paraphrase. Detector confidence drops substantially under heavy rewriting or translation. The image, audio and video variants are not open source; verification goes through Google's SynthID Detector portal.

Report those limits when you build on it. A detector with an "uncertain" state is telling you something. Collapse it into a yes or no and you have made the detector worse.

C2PA, the provenance side

The Coalition for Content Provenance and Authenticity publishes the Content Credentials specification. Version 2.3 was published in January 2026, and version 2.4 in April 2026. A conformance programme lists products validated against the specification and its security requirements.

The data model, in the specification's own terms:

ConceptWhat it is
AssertionOne declaration about an asset: metadata, an action taken, a thumbnail, a content binding
ClaimA signed structure referencing a set of assertions
Claim signatureThe digital signature over the claim, made with the signer's private key
ManifestThe container: assertions, one claim, and the claim signature
IngredientA referenced asset that was incorporated into a composed asset
Hard bindingCryptographic hashes identifying the asset or part of it, so tampering is detectable
Soft bindingA content identifier that is not statistically unique: a fingerprint, or an embedded invisible watermark

The relationship is the part to internalise. Assertions populate the manifest. The claim references them and is signed. Hard bindings prove integrity. And soft bindings exist precisely to survive the loss of the manifest. That is where watermarking re-enters. It is the fallback that matches a stripped file back to its provenance record.

This is the design answer to the screenshot problem. Neither mechanism is sufficient; the standard specifies both because each covers the other's failure.

The regulatory requirement

Article 50 of the EU AI Act covers AI systems generating synthetic audio, image, video or text. Providers must ensure outputs are marked in a machine-readable format. The outputs must also be detectable as artificially generated. The obligations apply from 2 August 2026. Generative systems already on the market before that date have until 2 December 2026 for the machine-readable marking requirement. That extension comes from the AI Omnibus provisional agreement of May 2026.

The European Commission published finalised guidelines on 20 July 2026. A Code of Practice on Transparency of AI-Generated Content came with them.

Two points engineers get wrong. The obligation is not restricted to high-risk systems. It applies to any system used in the situations Article 50 covers. And "machine-readable" means a detector can find it. An on-image caption alone does not discharge the duty.

This page is not legal advice. If you ship to the EU, read the Commission guidelines and talk to someone qualified.

Common mistakes

Treating a visible caption as compliance. It is stripped by a crop and is not machine-readable.

Relying on the file's metadata alone. Most upload pipelines strip EXIF and XMP by default, often for privacy reasons.

Claiming a watermark is tamper-proof. The output above shows what happens to a naive one, and published attacks exist against production schemes too.

Treating absence of a watermark as proof of authenticity. It means the file has no mark. Cameras without C2PA support, older files, and any stripped file all look identical to a genuine unmarked photo.

Verifying only the signature and not the certificate chain. A validly signed manifest from an untrusted signer proves nothing about origin.

Try it yourself

Change A and B to lower-frequency coefficients such as (1, 2) and (2, 1) and rerun. JPEG survival improves markedly, because low frequencies are quantised less. Then check the PSNR: the mark becomes visible. That single sweep is the robustness-versus-imperceptibility trade-off every watermarking paper is arguing about.

What to learn next

Researcher — Mathematics and papers.

The three requirements, and why they conflict

Any watermarking scheme is judged on capacity (bits carried), imperceptibility (distortion added) and robustness (survival under transformation). These trade against one another, and the trade is information-theoretic rather than an implementation detail.

The standard formalisation is Moulin and O'Sullivan (2003), Information-theoretic analysis of information hiding. It frames watermarking as communication over a channel controlled by an adversary. Watermark capacity under distortion constraints, $D_1$ for the embedder and $D_2$ for the attacker, is a game-theoretic quantity. Capacity falls to zero once the attacker's distortion budget matches the embedder's.

The practical consequence is visible in the developer output. An attacker willing to degrade the image as much as the watermark does can always remove the mark.

Families of scheme

Spatial and transform-domain classical methods. LSB substitution (fragile, zero robustness), spread spectrum (Cox et al., 1997), and quantisation index modulation (Chen and Wornell, 2001). QIM is the theoretically interesting one. It embeds by quantising a host feature to one of several interleaved lattices. Capacity approaches the informed-encoder bound under additive attacks.

Learned end-to-end. HiDDeN (Zhu et al., 2018) trains an encoder-decoder pair. A differentiable noise layer between them simulates crops, blur and JPEG. StegaStamp (Tancik et al., 2020) extends this to survive physical printing and re-photography. The pattern is now standard: put the attack in the training loop, and the embedder learns a robust code.

In-model watermarking. Rather than post-processing the output, alter the generative process. Stable Signature (Fernandez et al., 2023) fine-tunes the latent decoder so every image it produces carries a specific key. Tree-Ring (Wen et al., 2023) plants a pattern in the initial noise of the diffusion process. DDIM inversion detects it. The pattern lives in the Fourier domain of the latent, so it survives crops and rotations. Gaussian Shading (Yang et al., 2024) claims performance-lossless embedding in the initial noise.

In-model methods are the direction of travel. A user who declines the post-processing step cannot skip them.

Attacks

Attacks fall into three groups.

Removal: valumetric (compression, noise, filtering), geometric (crop, rotate, scale, warp), and regeneration. Regeneration is the strongest general attack. Encode and decode through a different autoencoder, or noise and denoise through a diffusion model. Most pixel-domain marks do not survive that. Zhao et al. (2023), Invisible Image Watermarks Are Provably Removable Using Generative AI, formalise this and give bounds.

Forgery: applying a target watermark to an unmarked image so a genuine photograph is flagged as generated. This is under-discussed and arguably more dangerous than removal, because it attacks the credibility of real evidence.

Detector inference: Jovanović et al. (2024), Watermark Stealing in LLMs, probe a detector's decision boundary. Enough of the key can be reconstructed to both remove and forge. The image analogue applies wherever a public detector API exists.

Bypass in deployment. Publicly documented tools strip or defeat production watermarks. Such tools were reported against SynthID during 2026. Treat any public claim of unbreakable watermarking as marketing.

Provenance as the complementary mechanism

C2PA takes the opposite approach: rather than hiding data in the pixels, attach a signed manifest. The version 2.4 specification describes it this way. The manifest binds assertions to an asset through a claim, signed by the actor that created it. Hard bindings, which are cryptographic hashes, prove integrity. Soft bindings, which are fingerprints or watermarks, enable recovery when the manifest is lost.

Three unresolved problems in deployment.

Stripping is trivial and undetectable. Any re-encode drops the manifest, and the resulting file is indistinguishable from one that never had one. The soft binding is the designed mitigation and inherits every watermark weakness above.

Trust reduces to certificate policy. A manifest is only as meaningful as the signer's identity and the verifier's trust list. The C2PA conformance programme holds generator products, validator products and certification authorities to the specification. The resulting conformance list is the trust anchor.

Capture-time signing has a hardware analogue-hole problem. A signed camera photographing a screen produces a genuinely signed image of a fake, and the signature is valid. Provenance says where the file came from, never whether the scene was real.

The evaluation standard

Report, at minimum:

  • Bit accuracy or detection AUC under an explicit, named attack suite, with attack strengths stated.
  • Imperceptibility as PSNR and a perceptual measure (LPIPS or SSIM), because PSNR alone hides structured artefacts.
  • The false-positive rate on a large corpus of genuine unmarked images. A detector with 0.1 percent false positives, applied at platform scale, mislabels enormous numbers of real photographs.

That last point is the one most papers under-report and the one that decides whether a system is deployable.

Sadasivan et al. (2023) and the broader impossibility literature argue reliable detection of generated content is not achievable in the general case. The defensible reading is narrower and still useful. Watermarking and provenance raise the cost of undetected misuse. They provide an audit trail for cooperative actors. They are not a truth oracle, and systems that treat them as one will cause harm.

References

  • Cox et al., Secure Spread Spectrum Watermarking for Multimedia, IEEE TIP, 1997
  • Chen and Wornell, Quantization Index Modulation, IEEE Trans. Inf. Theory, 2001
  • Moulin and O'Sullivan, Information-Theoretic Analysis of Information Hiding, IEEE Trans. Inf. Theory, 2003
  • Zhu et al., HiDDeN: Hiding Data With Deep Networks, 2018 — arxiv.org/abs/1807.09937
  • Tancik et al., StegaStamp, 2020 — arxiv.org/abs/1904.05343
  • Fernandez et al., The Stable Signature, 2023 — arxiv.org/abs/2303.15435
  • Wen et al., Tree-Ring Watermarks, 2023 — arxiv.org/abs/2305.20030
  • Zhao et al., Invisible Image Watermarks Are Provably Removable Using Generative AI, 2023 — arxiv.org/abs/2306.01953
  • C2PA, Content Credentials: C2PA Technical Specification 2.4 — spec.c2pa.org
  • European Commission, Guidelines on transparency obligations for providers and deployers of certain AI systems, 20 July 2026 — digital-strategy.ec.europa.eu

What to learn next