Image Generation and Restoration

Super-resolution

Super-resolution makes a small picture bigger by inventing plausible detail, and the standard quality score actively rewards the blurry answer.

On this page 9
  1. The short answer
  2. The analogy
  3. Why plain enlarging looks bad
  4. The trap in measuring this
  5. The two kinds of tool
  6. The honest limit
  7. Where you have seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Super-resolution enlarges a small picture by inventing the detail that was never recorded.

The analogy

Think about a friend describing a wedding you missed. They tell you the hall was big, the food was good, the bride wore red.

Your mind fills in a hall. It is a plausible hall. It is not the actual hall, because your friend never described the pillars or the ceiling.

Super-resolution does the same thing. The small picture is the description. The model fills in a plausible version of what was never captured.

That is worth holding on to. The extra detail is invented, not recovered. It was thrown away when the picture was made small, and it is not coming back.

Why plain enlarging looks bad

The simple way to enlarge is to guess each new pixel from its neighbours. Average them, smoothly.

The result is soft. Every edge that used to be crisp is now a gentle ramp. Nothing looks wrong exactly, and nothing looks sharp either.

That is because averaging is the safest answer. The model does not know whether a pixel is dark or bright. Splitting the difference is the least wrong single choice.

   original:  dark | bright        a real edge

   averaged:  dark - grey - bright     safe, and blurry

   invented:  dark | bright             sharp, and maybe in the wrong place

The trap in measuring this

Here comes the surprising part, and it explains a lot of confusing research.

The usual way to score an enlarged picture is to compare it pixel by pixel against the true one. Closer means better.

But that score rewards blurring. A soft, safe picture is on average nearer to the truth. A sharp picture whose edge is one pixel off scores worse. So a model trained to score well learns to be soft.

You will see this measured next. A deliberately blurred picture scores higher than a perfectly sharp one, shifted by a single pixel. To your eye, the sharp one is far better.

That single fact shaped fifteen years of research in this area.

The two kinds of tool

Faithful upscalers try to stay close to the truth. Safe, soft, and good when the picture is evidence.

Realistic upscalers produce convincing texture. Sharp, attractive, and willing to invent things that were never there.

You must pick based on what the picture is for. A wedding photo wants the second. A licence plate at a crime scene wants neither. Both are inventing, and a court should not see invented detail.

The honest limit

Ask a model to enlarge a face that is twelve pixels across. It will give you a face. That face is not the person.

Every detail is a guess drawn from the faces the model was trained on. Guesses drawn from a training set carry that set's biases. This has caused real public controversy.

If the picture is going to be used as evidence about a person, do not upscale it.

Where you have seen this

  • Phone camera "digital zoom" beyond the lens's real range.
  • Streaming boxes that upscale older shows to fit a bigger screen.
  • Old family photographs restored to print size.
  • Game settings that render at lower resolution and upscale for speed.

Remember this

  • The detail is invented, never recovered.
  • Pixel-comparison scores reward blur, which is why many upscalers look soft.
  • Never upscale a picture that has to serve as evidence about a person.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy==1.26.4 pillow==11.0.0 scikit-image==0.25.2 scipy==1.14.1 torch==2.5.1

The metric trap, measured

sr_metrics.py
import numpy as np, math
from PIL import Image
from skimage.metrics import structural_similarity as ssim

rng = np.random.default_rng(0)
# Ground truth: smooth shading plus fine texture, the two things SR has to handle differently.
y, x = np.mgrid[0:256, 0:256]
hr = (110 + 60 * np.sin(x / 30) + 30 * np.sin(y / 7.0) + rng.normal(0, 4, (256, 256)))
hr = np.clip(hr, 0, 255).astype(np.uint8)

lr = Image.fromarray(hr).resize((64, 64), Image.BICUBIC)      # the 4x downscale
def psnr(a, b): return 10 * math.log10(255.0 ** 2 / ((a.astype(float) - b.astype(float)) ** 2).mean())

print("upscaler        PSNR (dB)   SSIM")
for name, f in [("nearest", Image.NEAREST), ("bilinear", Image.BILINEAR),
                ("bicubic", Image.BICUBIC), ("lanczos", Image.LANCZOS)]:
    up = np.array(lr.resize((256, 256), f))
    print(f"{name:12s} {psnr(hr, up):11.2f} {ssim(hr, up, data_range=255):6.3f}")

print("\n--- the trade-off that makes SR benchmarks misleading ---")
bicubic = np.array(lr.resize((256, 256), Image.BICUBIC)).astype(float)
# candidate A: blur the bicubic result even more. Definitely worse to look at.
from scipy.ndimage import gaussian_filter
blurry = gaussian_filter(bicubic, 2.0)
# candidate B: the true image, shifted by one pixel. All the detail, slightly misplaced.
sharp_shifted = np.roll(hr.astype(float), 1, axis=0)
for name, cand in [("blurrier than bicubic", blurry), ("perfectly sharp, 1px off", sharp_shifted)]:
    print(f"{name:26s} PSNR {psnr(hr, cand):6.2f} dB   SSIM {ssim(hr, cand.astype(np.uint8), data_range=255):.3f}   "
          f"high-frequency energy {np.abs(np.diff(cand, axis=0)).mean():5.2f}")
print(f"{'the real image itself':26s} PSNR {'inf':>6}      SSIM 1.000   "
      f"high-frequency energy {np.abs(np.diff(hr.astype(float), axis=0)).mean():5.2f}")
Output
upscaler        PSNR (dB)   SSIM
nearest            33.50  0.827
bilinear           36.03  0.890
bicubic            36.20  0.892
lanczos            36.20  0.892

--- the trade-off that makes SR benchmarks misleading ---
blurrier than bicubic      PSNR  35.90 dB   SSIM 0.888   high-frequency energy  2.57
perfectly sharp, 1px off   PSNR  31.66 dB   SSIM 0.794   high-frequency energy  5.23
the real image itself      PSNR    inf      SSIM 1.000   high-frequency energy  5.13

Reading the output

The interpolation ranking is what you would expect. Nearest is worst, bicubic and Lanczos are best and effectively tied. Nothing surprising, and this is your baseline: any learned upscaler must beat 36.20 dB here to be worth running.

Then the third block breaks the metric. A picture that has been deliberately blurred beyond bicubic scores 35.90 dB. A picture with every detail perfectly intact, shifted by one single row, scores 31.66 dB.

PSNR prefers the blur by more than 4 dB. So does SSIM, 0.888 against 0.794.

The high-frequency column shows what each candidate actually contains. The real image has 5.13. The shifted-sharp version has 5.23, essentially identical, because it is the real image. The blurred version has 2.57, half the detail.

So the metric prefers the candidate with half the texture. It rejects the one with all of it, slightly misplaced. Any model trained to minimise pixel error will learn this preference and produce soft images. That is not a flaw in the models. It is the loss function working correctly on a badly chosen objective.

What sharp models do instead

Add a term that does not compare pixels.

  • Perceptual loss (Johnson et al., 2016): compare activations of a pretrained network rather than raw pixels. Two images with the same texture in slightly different positions look similar in feature space.
  • Adversarial loss (SRGAN, Ledig et al., 2017): a discriminator scores whether the output looks like a real photograph.
  • Both, weighted, plus a small pixel term for stability. That combination is ESRGAN and is still the recipe behind most practical upscalers.

The building block: sub-pixel convolution

Upsampling inside the network predicts extra channels at low resolution, then rearranges them. That beats interpolating, and it beats transposed convolution.

espcn.py
import torch, torch.nn as nn, math, time
torch.manual_seed(0)

# The sub-pixel shuffle: predict r*r channels at LOW resolution, then rearrange
# them into one channel at HIGH resolution. No interpolation, no transposed conv.
x = torch.arange(16).float().view(1, 4, 2, 2)
print("input ", tuple(x.shape))
print("output", tuple(nn.PixelShuffle(2)(x).shape), " <- 4 channels at 2x2 became 1 channel at 4x4")
print(nn.PixelShuffle(2)(x)[0, 0].int().tolist())

R = 4
def batch(n, size=64):
    g = torch.arange(size).float()
    yy, xx = torch.meshgrid(g, g, indexing="ij")
    ph = torch.rand(n, 1, 1) * 6.28
    hr = (0.5 + 0.25 * torch.sin(xx / 5 + ph) + 0.25 * torch.sin(yy / 7 + ph)).unsqueeze(1)
    # A REALISTIC degradation, not a clean resize: downscale, then add sensor noise.
    lr = torch.nn.functional.interpolate(hr, scale_factor=1 / R, mode="area")
    lr = lr + 0.05 * torch.randn_like(lr)
    return lr, hr

espcn = nn.Sequential(nn.Conv2d(1, 32, 5, padding=2), nn.Tanh(),
                      nn.Conv2d(32, 32, 3, padding=1), nn.Tanh(),
                      nn.Conv2d(32, R * R, 3, padding=1), nn.PixelShuffle(R))
opt = torch.optim.Adam(espcn.parameters(), lr=2e-3)
t0 = time.time()
for _ in range(800):
    lr, hr = batch(16)
    loss = ((espcn(lr) - hr) ** 2).mean()
    opt.zero_grad(); loss.backward(); opt.step()

torch.manual_seed(99)
lr, hr = batch(64)
psnr = lambda a, b: 10 * math.log10(1.0 / ((a - b) ** 2).mean().item())
with torch.no_grad():
    bic = torch.nn.functional.interpolate(lr, scale_factor=R, mode="bicubic", align_corners=False)
    print(f"\nbicubic PSNR      {psnr(bic, hr):.2f} dB")
    print(f"tiny ESPCN PSNR   {psnr(espcn(lr), hr):.2f} dB")
    print(f"parameters        {sum(p.numel() for p in espcn.parameters()):,}")
    print(f"training seconds  {time.time()-t0:.0f} (varies by machine)")
Output
input  (1, 4, 2, 2)
output (1, 1, 4, 4)  <- 4 channels at 2x2 became 1 channel at 4x4
[[0, 4, 1, 5], [8, 12, 9, 13], [2, 6, 3, 7], [10, 14, 11, 15]]

bicubic PSNR      26.98 dB
tiny ESPCN PSNR   30.85 dB
parameters        14,704
training seconds  9 (varies by machine)

Exact PSNR values depend on your PyTorch build. The ordering reproduces.

The shuffle pattern is worth reading carefully. Channel 0 supplies positions (0,0), (0,2), (2,0), (2,2), channel 1 supplies the odd columns, and so on. Every output pixel comes from exactly one input channel. That is why this has no checkerboard artefacts. Unlike transposed convolution, no output pixel sums an uneven number of contributions.

The learned model beat bicubic by 3.9 dB, and the reason matters. In the earlier script the degradation was a clean bicubic downscale. A tiny learned model loses to bicubic there, because there is nothing left to learn. Here the degradation also adds sensor noise, so the model can denoise as well as upscale. That is the difference between a toy benchmark and a real photograph. It is exactly why Real-ESRGAN's contribution was a degradation pipeline, not an architecture.

Common mistakes

Training on bicubic-downscaled data and deploying on real photos. Real low-resolution images carry compression artefacts, sensor noise, motion blur and prior resizing. A model trained only on clean downscales sharpens the artefacts confidently. This is the single most common reason an upscaler works on the benchmark and fails on your images.

Reporting PSNR only. See the first table. Report PSNR, SSIM, LPIPS and a no-reference score together, or say plainly that you optimised for one of them.

Upscaling before, rather than after, other processing. Upscale last. Denoise, colour-correct and crop on the small image. Everything is cheaper there, and the artefacts are not yet magnified.

Applying face-restoration models to non-faces. GFPGAN and similar models have a strong face prior. Run them on a hand or a plate of food and you may get a face.

Ignoring tile seams. Large images are upscaled in tiles for memory. Without overlap and blending, tile boundaries show as faint grid lines.

Try it yourself

Remove the lr = lr + 0.05 * torch.randn_like(lr) line and retrain. Bicubic will now beat the learned model. That one line is the whole argument for realistic degradation modelling. Running it both ways fixes the idea permanently.

What to learn next

Researcher — Mathematics and papers.

The problem is one-to-many

Take a low-resolution $\mathbf{y} = (\mathbf{x} \otimes \mathbf{k}) \downarrow_s + \mathbf{n}$. Here $\mathbf{k}$ is a blur kernel, $\downarrow_s$ is decimation by factor $s$, and $\mathbf{n}$ is noise. Many high-resolution $\mathbf{x}$ are consistent with the same $\mathbf{y}$. The task is not inversion; it is sampling or selecting from a set.

Minimising $\mathbb{E}|\mathbf{x} - \hat{\mathbf{x}}|^2$ returns the conditional mean $\mathbb{E}[\mathbf{x} \mid \mathbf{y}]$. That is the average of all consistent images, and so it is blurred by construction. This is not an approximation error. It is the exact minimiser of the stated objective.

The perception-distortion trade-off

Blau and Michaeli (2018) prove this is a fundamental bound, not an engineering limitation. Define distortion as $\mathbb{E}[\Delta(\mathbf{x}, \hat{\mathbf{x}})]$, for any full-reference measure $\Delta$. Define perceptual quality as $d(p_{\hat{\mathbf{x}}}, p_{\mathbf{x}})$, a divergence between the output distribution and the distribution of natural images. Their result:

The perception-distortion function $P(D) = \min_{p_{\hat{\mathbf{x}}|\mathbf{y}}} d(p_{\mathbf{x}}, p_{\hat{\mathbf{x}}}) \;\text{ s.t. }\; \mathbb{E}[\Delta] \le D$ is monotonically non-increasing and convex.

Improving distortion past a point necessarily worsens perceptual quality, for every distortion measure. The PIRM 2018 challenge was designed around this, scoring methods on a perception-distortion plane rather than a single axis.

The practical reading is short. A paper claiming both the best PSNR and the best perceptual quality is doing one of two things. It is either evaluating where the bound is loose, or measuring something incorrectly.

Architecture lineage

YearMethodContribution
2014SRCNNThree convolutions on a bicubic-upsampled input
2016ESPCNSub-pixel convolution; process at low resolution, upsample last
2016VDSRDeep network learning the residual over bicubic
2017SRGAN / SRResNetPerceptual plus adversarial loss; the perceptual turn
2017EDSRRemoves batch norm from residual blocks; large gains
2018RCANChannel attention, very deep residual-in-residual
2018ESRGANRRDB blocks, relativistic discriminator, pre-activation perceptual features
2021SwinIRSwin Transformer blocks for restoration
2021Real-ESRGANHigh-order random degradation pipeline for real images
2021BSRGANRandomly shuffled degradations, same motivation
2023 onwardStableSR, DiffBIR, SUPIRDiffusion priors for blind restoration

EDSR's batch-norm removal is worth stating because it generalises to all restoration. Batch norm normalises away the per-image contrast and colour statistics. In classification that is a feature. In restoration it is destruction of the signal. Restoration networks use no normalisation, or layer norm.

Real-ESRGAN and BSRGAN are the practical turning point. Both argue the architecture was never the bottleneck. Models trained on bicubic downscales fail on real photographs. The degradation assumed at training time does not match reality. Their solution is a randomised composition of blur, resize, noise and JPEG compression, applied in several rounds. The NTIRE 2026 mobile real-world super-resolution challenge report confirms this remains standard. Entries follow Real-ESRGAN-style degradation simulation and BasicSR engineering practice as the stable baseline.

Diffusion-based restoration

Diffusion priors sit at the far perceptual end of the curve. StableSR, DiffBIR and SUPIR condition a pretrained text-to-image diffusion model on the degraded input. They produce textures no regression model can match. The cost is low fidelity, high compute, and hallucinated content.

The 2026 literature reports hybrid designs converging on this. They combine the structural reliability of GAN-based methods with the texture generation of diffusion models. Residual refinement holds the two ends together. Evaluation in these challenges leans on no-reference metrics: NIQE, NRQM, PI, CLIP-IQA, MUSIQ. The full-reference metrics penalise the behaviour these methods are designed for.

Evaluation, and what each metric actually rewards

  • PSNR: pure distortion. Rewards the conditional mean. Correlates poorly with human judgement above about 30 dB.
  • SSIM / MS-SSIM: local structure comparison. Better than PSNR, still a distortion measure and still on the wrong side of the bound.
  • LPIPS (Zhang et al., 2018): distance in the features of a pretrained network, calibrated against human judgements. Better correlated, and it is still a reference metric, so it retains a distortion character.
  • NIQE, PI, MUSIQ, CLIP-IQA: no-reference. They score whether the output looks like a natural image, with no ground truth involved. A model can score well here while inventing content freely, which is both the point and the danger.

The honest report is a table with at least one metric from each of the three groups. Add the operating point on the perception-distortion plane.

The evidential problem

Upscaling is generation, and generated faces are drawn from the training distribution. The 2020 PULSE incident is the widely cited example. A downsampled photograph of Barack Obama upscaled to a white face. The underlying cause is structural rather than a bug. The model samples a plausible face from its prior, and the prior reflects its training data.

This is a hard constraint on deployment. Super-resolution output must not be presented as recovered evidence about an individual, in forensic, medical or identification contexts. Menon et al. (2020) added a statement to PULSE's own paper acknowledging exactly this failure mode.

Papers

What to learn next