Image Generation and Restoration

Inpainting and outpainting

Inpainting fills a hole inside a picture and outpainting extends it past its edges, and both work by keeping the known pixels pinned at every step.

On this page 9
  1. The short answer
  2. The analogy
  3. What you supply
  4. How the model keeps the rest of the picture
  5. Why the old repair tools were not enough
  6. Outpainting is the same thing
  7. Where you have seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Inpainting fills a chosen area of a picture with something new, keeping the rest untouched.

The analogy

Think about repairing a torn cotton bedsheet. A tailor does not weave a whole new sheet. They cut a patch, match the thread colour, and stitch it so the join disappears.

The good repair is not the one with the nicest patch. It is the one you cannot find afterwards.

Outpainting is the same craft, applied at the edges. You have a photo that got cropped too tight, and you want more sky above and more road below. The tailor is adding cloth to the hem.

What you supply

Two things. The picture, and a mask, which is a black-and-white stencil marking which pixels may be changed.

White means "you may repaint here". Black means "leave this alone".

   the photo            the mask              the result
   +--------+          +--------+          +--------+
   | sky    |          |........|          | sky    |
   |  [car] |    +     |..####..|    ->    |        |
   | road   |          |........|          | road   |
   +--------+          +--------+          +--------+
                     white = repaint      the car is gone,
                                          road continues through

How the model keeps the rest of the picture

Here is the mechanism, and it is neater than most people expect.

A picture generator works by starting from noise and cleaning it up over many rounds. For inpainting, something happens at every single round. The model's work outside the mask is thrown away. In its place goes the real photo, degraded to match that round's noise level.

So the known part of the image is re-asserted hundreds of times. The model never gets to change it. But the model does see it, every round. Its guesses inside the hole are shaped by what surrounds them.

That is why the fill matches the lighting and the texture around it. It was looking at them the whole way.

Why the old repair tools were not enough

Photo editors have had a healing brush for decades. It works by spreading nearby colour into the hole.

For a scratch or a dust speck, that is perfect. For a large hole in a brick wall, it produces a smooth, empty smear. Spreading colour cannot invent bricks.

You will see that measured with real numbers in the next section, and the gap is large.

Outpainting is the same thing

Pad the picture with blank space. Mark all the blank space as repaintable. Run inpainting.

That is genuinely all outpainting is. Every "expand this image" feature works this way.

The one honest caveat: extend too far in one go and the model loses the thread. Better to extend a little, accept it, then extend again.

Where you have seen this

  • "Magic eraser" tools that remove a person from a holiday photo.
  • Object removal in phone photo editors.
  • Widening a portrait photo into a banner.
  • Restoring damaged old family photographs.

Remember this

  • You supply a mask saying which pixels may change.
  • The known pixels are re-asserted at every step, so they cannot drift.
  • Outpainting is inpainting on padding added around the picture.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install opencv-python==4.10.0 numpy==1.26.4 torch==2.5.1

Where classical inpainting works, and where it stops

Before reaching for a diffusion model, know what the cheap method gets you. OpenCV ships two classical algorithms, both of which propagate colour inward from the hole boundary.

classical_inpaint.py
import cv2, numpy as np

rng = np.random.default_rng(0)
# A synthetic photo: a smooth gradient sky plus a striped brick wall.
y, x = np.mgrid[0:128, 0:128]
img = (60 + y * 0.6).astype(np.uint8)
img[70:, :] = (140 + 40 * ((x[70:, :] // 8) % 2)).astype(np.uint8)
img = np.clip(img.astype(np.int16) + rng.normal(0, 3, img.shape), 0, 255).astype(np.uint8)

def hole(size, top, left):
    m = np.zeros(img.shape, np.uint8)
    m[top:top+size, left:left+size] = 255
    return m

print("hole   region        Telea MSE   Navier-Stokes MSE   copy-the-mean MSE")
for size, top, left, where in [(6, 20, 20, "sky, tiny"), (6, 90, 20, "wall, tiny"),
                               (40, 15, 40, "sky, large"), (40, 80, 40, "wall, large")]:
    m = hole(size, top, left)
    damaged = img.copy(); damaged[m > 0] = 0
    inside = m > 0
    t = cv2.inpaint(damaged, m, 3, cv2.INPAINT_TELEA)
    n = cv2.inpaint(damaged, m, 3, cv2.INPAINT_NS)
    flat = np.full_like(img, int(img[~inside].mean()))
    mse = lambda a: float(((a[inside].astype(float) - img[inside].astype(float)) ** 2).mean())
    print(f"{size:4d}   {where:12s} {mse(t):11.1f} {mse(n):19.1f} {mse(flat):19.1f}")

# Why generative inpainting exists: measure how much TEXTURE the fill contains.
m = hole(40, 80, 40)                       # the large hole in the brick wall
damaged = img.copy(); damaged[m > 0] = 0
filled = cv2.inpaint(damaged, m, 3, cv2.INPAINT_TELEA)
inside = m > 0
detail = lambda a: float(np.abs(cv2.Laplacian(a.astype(np.float32), cv2.CV_32F))[inside].std())
print("")
print("texture energy inside the hole, real bricks:    %.2f" % detail(img))
print("texture energy inside the hole, classical fill: %.2f" % detail(filled))
print("the fill is not wrong at the edges. It is empty in the middle.")
Output
hole   region        Telea MSE   Navier-Stokes MSE   copy-the-mean MSE
   6   sky, tiny           10.2                 9.3              1897.3
   6   wall, tiny          86.9                61.8              1841.5
  40   sky, large          24.4                22.4              1588.9
  40   wall, large        572.9               641.8              3260.0

texture energy inside the hole, real bricks:    16.23
texture energy inside the hole, classical fill: 3.03
the fill is not wrong at the edges. It is empty in the middle.

Reading the output

Small holes are a solved problem. Six pixels of sky, error 10.2. Do not train a model for dust specks and scratches. cv2.inpaint runs in microseconds and is good enough.

Large holes in smooth regions are also fine. A 40-pixel hole in the sky costs 24.4, barely worse than the tiny one. Smooth content is what colour propagation is designed for.

Large holes in textured regions fail. 572.9 on the brick wall, more than twenty times the sky error. The fallback of filling with the image mean scores 3260. Classical inpainting is doing something, but not enough.

The texture measurement says why. Real bricks in that region have a detail energy of 16.23. The fill has 3.03, about one fifth. The algorithm did not put the wrong bricks there. It put no bricks there.

Colour propagation solves a smoothness equation. Smoothness is exactly what texture is not. That is the gap generative inpainting exists to close. It also tells you which of your own images will need a model.

The masking mechanism, on a distribution where correctness is checkable

The claim is about replacing the known region at every denoising step, rather than pasting at the end. Doing so makes the fill respect the joint structure of the data. Here is that claim as a number.

The data has three clusters and no cluster at (+2, -2). So if you are told x = +2, the only valid y is +2. A method that ignores the dependency will produce both.

diffusion_inpaint.py
import torch, torch.nn as nn, time
torch.manual_seed(0)

# Three clusters. There is no cluster at (+2, -2), so "x is +2" implies "y must be +2".
CENTRES = torch.tensor([[-2., -2.], [-2., 2.], [2., 2.]])
def data(n): return CENTRES[torch.randint(0, 3, (n,))] + 0.35 * torch.randn(n, 2)

T = 400
betas = torch.linspace(1e-4, 0.02, T); abar = torch.cumprod(1 - betas, 0)
net = nn.Sequential(nn.Linear(3, 128), nn.SiLU(), nn.Linear(128, 128), nn.SiLU(), nn.Linear(128, 2))
eps = lambda x, t: net(torch.cat([x, t[:, None].float() / T], -1))
opt = torch.optim.Adam(net.parameters(), lr=2e-3)
t0 = time.time()
for _ in range(3000):
    x0 = data(256); t = torch.randint(0, T, (256,)); n = torch.randn_like(x0)
    a = abar[t][:, None]
    loss = ((eps(a.sqrt() * x0 + (1 - a).sqrt() * n, t) - n) ** 2).mean()
    opt.zero_grad(); loss.backward(); opt.step()
print(f"trained in {time.time()-t0:.0f}s")

KNOWN = torch.tensor([True, False])          # x is given, y must be filled in
X_GIVEN = 2.0

@torch.no_grad()
def run(replace_every_step, n=4000):
    torch.manual_seed(3)
    x = torch.randn(n, 2)
    known_clean = torch.full((n, 2), X_GIVEN)
    for i in reversed(range(T)):
        if replace_every_step:
            # overwrite the known pixels with the TRUE value, noised to this step's level
            a = abar[i]
            x[:, KNOWN] = (a.sqrt() * known_clean + (1 - a).sqrt() * torch.randn(n, 2))[:, KNOWN]
        t = torch.full((n,), i)
        e = eps(x, t)
        mean = (x - betas[i] / (1 - abar[i]).sqrt() * e) / (1 - betas[i]).sqrt()
        x = mean + (betas[i].sqrt() * torch.randn_like(x) if i else 0)
    x[:, KNOWN] = X_GIVEN                     # the final paste, in both variants
    return x

for label, rep in [("paste only at the end", False), ("replace at every step", True)]:
    out = run(rep)
    up = (out[:, 1] > 0).float().mean()
    print(f"{label:24s}  filled y is +2 in {up:5.1%} of samples, "
          f"mean |y| {out[:,1].abs().mean():.2f}")
print("\ncorrect answer: y must be +2 in 100% of samples, because (+2, -2) is not in the data")
Output
trained in 7s
paste only at the end     filled y is +2 in 70.5% of samples, mean |y| 2.02
replace at every step     filled y is +2 in 99.6% of samples, mean |y| 1.90

correct answer: y must be +2 in 100% of samples, because (+2, -2) is not in the data

Pasting only at the end gives 70.5%. That is close to the two-in-three you get from ignoring the condition and sampling all three clusters. Replacing at every step gives 99.6%.

In image terms, this is the difference between a fill that ignores its surroundings and one that continues them. The first produces a visible patch. The second produces a repair.

The two ways real tools do this

Replacement, no retraining. Exactly the loop above, applied to latents. Works with any checkpoint, which is why every "inpaint with any model" feature exists. Weaker at large holes. The model only ever sees the known region as noised context, never as a clean signal.

A dedicated inpainting model. The UNet's input convolution is widened from 4 latent channels to 9. That is 4 for the noisy latent, 4 for the encoded masked image, and 1 for the downsampled mask. The model is then fine-tuned on masked data. This is what runwayml/stable-diffusion-inpainting and the SDXL inpainting checkpoints are. Better at large holes, and it will not load into a standard pipeline because the channel count differs.

Practical details that decide whether it looks right

Dilate the mask. Object shadows, reflections and colour fringing extend past the object. A mask drawn tightly around a car leaves a car-shaped shadow behind. Grow it by 5 to 15 pixels with cv2.dilate.

Blur the mask edge. A hard mask boundary in latent space becomes a visible seam after decoding. One latent pixel covers an 8-by-8 block of image pixels. A few pixels of Gaussian blur on the mask removes it.

Watch the latent resolution. At an 8x downsample, a 20-pixel mask is 2.5 latent pixels. Masks smaller than roughly 32 image pixels are barely representable. For those, use classical inpainting.

Set strength deliberately in img2img-style inpainting. strength=1.0 ignores the original content inside the mask completely. Lower values keep some of it, which is what you want when altering an object rather than removing it.

For outpainting, use cv2.copyMakeBorder with BORDER_REFLECT, not black. The reflected content gives the model a plausible starting point and reduces the dark halo at the join.

Common mistakes

Inverted mask polarity. White is the region to change. Half of all inpainting bugs are this, and the symptom is unmistakable: everything except your selection gets repainted.

Mask not the same size as the image. Silent resizing shifts the region.

Expecting a large fill to match a texture perfectly. Bricks, tiles, printed fabric and text are the hard cases. The model produces something plausible, but not aligned to the surrounding grid. Inspect the join.

Loading an inpainting checkpoint into a text-to-image pipeline. The 9-channel input convolution will not accept 4 channels.

Outpainting a whole extra image in one pass. Extend by 25 to 50 percent at a time.

Try it yourself

Change KNOWN to [False, True] so the roles swap: y is given as +2, and x must be filled. Now both -2 and +2 are valid answers for x. Check that the replacement method produces roughly a two-way split rather than collapsing to one. That confirms it is sampling the true conditional, not memorising.

What to learn next

Researcher — Mathematics and papers.

Replacement-based inpainting

Let $\mathbf{m}$ be a binary mask with 1 on known pixels. Song et al. (2021) and, in developed form, Lugmayr et al. (2022) (RePaint) define the sampling step as:

$$ \mathbf{x}{t-1}^{\text{known}} \sim \mathcal{N}!\left(\sqrt{\bar{\alpha}{t-1}}\,\mathbf{x}0,\; (1 - \bar{\alpha}{t-1})\mathbf{I}\right) $$ $$ \mathbf{x}{t-1}^{\text{unknown}} \sim p\theta(\mathbf{x}_{t-1} \mid \mathbf{x}t) $$ $$ \mathbf{x}{t-1} = \mathbf{m} \odot \mathbf{x}{t-1}^{\text{known}} + (1 - \mathbf{m}) \odot \mathbf{x}{t-1}^{\text{unknown}} $$

The known region is resampled from the exact forward marginal at each step. No retraining is required, and any unconditional diffusion model becomes an inpainter.

RePaint's central observation is a failure of this scheme. The model sees the known region only through $\mathbf{x}_t$. The reverse step conditions on it for one step, then the composition overwrites it. The generated content is locally plausible but globally disharmonious with the context.

Their fix is resampling. After computing $\mathbf{x}_{t-1}$, diffuse it back to $\mathbf{x}_t$ and denoise again. This runs $U$ times per step, and they use $U = 10$. Each round lets the unknown region re-condition on the known one. The paper shows harmonisation improving monotonically with $U$. The cost is $U$ times the compute.

Dedicated inpainting models

The alternative conditions the network directly. The input to the UNet becomes $[\mathbf{z}_t; \mathcal{E}(\mathbf{x} \odot \mathbf{m}); \mathrm{down}(\mathbf{m})]$. That concatenation along the channel axis gives $4 + 4 + 1 = 9$ channels for a Stable Diffusion latent. The extra input-convolution weights are zero-initialised and the model is fine-tuned with random masks.

Mask distribution during training is the design decision that matters. Training on small random rectangles produces a model that cannot fill large regions. Published recipes mix wide free-form strokes, large rectangles, full-object masks derived from segmentation, and whole-image masks. LaMa (Suvorov et al., 2021) makes the same point for the non-diffusion case. Their large-mask training regime is as important as their architecture.

Blended Latent Diffusion (Avrahami et al., 2023) sits between the two. It performs the blend in latent space at each step. An optimisation stage fixes the imprecision caused by the 8x downsampling of the mask.

Non-diffusion baselines worth knowing

  • Telea (2004) and Bertalmio et al. (2001), the two OpenCV algorithms, propagate isophotes inward. Both solve smoothness problems and neither synthesises texture.
  • PatchMatch (Barnes et al., 2009) finds approximate nearest-neighbour patch correspondences and copies real texture from elsewhere in the image. It remains excellent for texture continuation and is what Photoshop's Content-Aware Fill was built on. It cannot invent structure that does not already appear in the image.
  • LaMa (Suvorov et al., 2021) uses fast Fourier convolutions to give early layers an image-wide receptive field. That lets it complete periodic structures such as brickwork and windows, across very large masks. For pure removal tasks it is often better and vastly cheaper than a diffusion model.

The practical ordering is this. Classical for specks. PatchMatch or LaMa for removal on structured backgrounds. Diffusion when new semantic content must be invented.

Evaluation

Inpainting has no single correct answer. Reference metrics are only valid on synthetic masks over known ground truth. Even then they penalise plausible alternatives. Standard practice:

  • LPIPS and FID over the whole image. Add U-IDS / P-IDS (Zhao et al., 2021). Both measure how often a linear classifier confuses real and inpainted images.
  • Masked-region-only metrics, since whole-image PSNR is dominated by the unchanged pixels and looks impressive for no reason.
  • Boundary consistency, measuring gradient discontinuity across the mask edge.
  • Human two-alternative forced choice, which remains the reference standard.

Report which mask distribution was used. Results on thin strokes and on 40-percent-area masks are not comparable. Papers reporting only one are reporting the easy case.

Papers

  • Telea, An Image Inpainting Technique Based on the Fast Marching Method, 2004
  • Barnes et al., PatchMatch, SIGGRAPH 2009
  • Suvorov et al., Resolution-robust Large Mask Inpainting with Fourier Convolutions (LaMa), 2021 — arxiv.org/abs/2109.07161
  • Lugmayr et al., RePaint: Inpainting using Denoising Diffusion Probabilistic Models, 2022 — arxiv.org/abs/2201.09865
  • Rombach et al., High-Resolution Image Synthesis with Latent Diffusion Models, 2022 — arxiv.org/abs/2112.10752
  • Avrahami et al., Blended Latent Diffusion, 2023 — arxiv.org/abs/2206.02779
  • Saharia et al., Palette: Image-to-Image Diffusion Models, 2022 — arxiv.org/abs/2111.05826

What to learn next