Image Generation and Restoration
Super-resolution
Super-resolution makes a small picture bigger by inventing plausible detail, and the standard quality score actively rewards the blurry answer.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Super-resolution enlarges a small picture by inventing the detail that was never recorded.
The analogy
Think about a friend describing a wedding you missed. They tell you the hall was big, the food was good, the bride wore red.
Your mind fills in a hall. It is a plausible hall. It is not the actual hall, because your friend never described the pillars or the ceiling.
Super-resolution does the same thing. The small picture is the description. The model fills in a plausible version of what was never captured.
That is worth holding on to. The extra detail is invented, not recovered. It was thrown away when the picture was made small, and it is not coming back.
Why plain enlarging looks bad
The simple way to enlarge is to guess each new pixel from its neighbours. Average them, smoothly.
The result is soft. Every edge that used to be crisp is now a gentle ramp. Nothing looks wrong exactly, and nothing looks sharp either.
That is because averaging is the safest answer. The model does not know whether a pixel is dark or bright. Splitting the difference is the least wrong single choice.
original: dark | bright a real edge
averaged: dark - grey - bright safe, and blurry
invented: dark | bright sharp, and maybe in the wrong placeThe trap in measuring this
Here comes the surprising part, and it explains a lot of confusing research.
The usual way to score an enlarged picture is to compare it pixel by pixel against the true one. Closer means better.
But that score rewards blurring. A soft, safe picture is on average nearer to the truth. A sharp picture whose edge is one pixel off scores worse. So a model trained to score well learns to be soft.
You will see this measured next. A deliberately blurred picture scores higher than a perfectly sharp one, shifted by a single pixel. To your eye, the sharp one is far better.
That single fact shaped fifteen years of research in this area.
The two kinds of tool
Faithful upscalers try to stay close to the truth. Safe, soft, and good when the picture is evidence.
Realistic upscalers produce convincing texture. Sharp, attractive, and willing to invent things that were never there.
You must pick based on what the picture is for. A wedding photo wants the second. A licence plate at a crime scene wants neither. Both are inventing, and a court should not see invented detail.
The honest limit
Ask a model to enlarge a face that is twelve pixels across. It will give you a face. That face is not the person.
Every detail is a guess drawn from the faces the model was trained on. Guesses drawn from a training set carry that set's biases. This has caused real public controversy.
If the picture is going to be used as evidence about a person, do not upscale it.
Where you have seen this
- Phone camera "digital zoom" beyond the lens's real range.
- Streaming boxes that upscale older shows to fit a bigger screen.
- Old family photographs restored to print size.
- Game settings that render at lower resolution and upscale for speed.
Remember this
- The detail is invented, never recovered.
- Pixel-comparison scores reward blur, which is why many upscalers look soft.
- Never upscale a picture that has to serve as evidence about a person.
What to learn next
- Watermarking and provenance — proving where an image came from, once it can be invented.
- Inpainting and outpainting — the other restoration task, with the same honesty problem.
- Convolutional neural networks — the backbone underneath every method here.
Developer — Code and libraries.
Setup
pip install numpy==1.26.4 pillow==11.0.0 scikit-image==0.25.2 scipy==1.14.1 torch==2.5.1The metric trap, measured
import numpy as np, math
from PIL import Image
from skimage.metrics import structural_similarity as ssim
rng = np.random.default_rng(0)
# Ground truth: smooth shading plus fine texture, the two things SR has to handle differently.
y, x = np.mgrid[0:256, 0:256]
hr = (110 + 60 * np.sin(x / 30) + 30 * np.sin(y / 7.0) + rng.normal(0, 4, (256, 256)))
hr = np.clip(hr, 0, 255).astype(np.uint8)
lr = Image.fromarray(hr).resize((64, 64), Image.BICUBIC) # the 4x downscale
def psnr(a, b): return 10 * math.log10(255.0 ** 2 / ((a.astype(float) - b.astype(float)) ** 2).mean())
print("upscaler PSNR (dB) SSIM")
for name, f in [("nearest", Image.NEAREST), ("bilinear", Image.BILINEAR),
("bicubic", Image.BICUBIC), ("lanczos", Image.LANCZOS)]:
up = np.array(lr.resize((256, 256), f))
print(f"{name:12s} {psnr(hr, up):11.2f} {ssim(hr, up, data_range=255):6.3f}")
print("\n--- the trade-off that makes SR benchmarks misleading ---")
bicubic = np.array(lr.resize((256, 256), Image.BICUBIC)).astype(float)
# candidate A: blur the bicubic result even more. Definitely worse to look at.
from scipy.ndimage import gaussian_filter
blurry = gaussian_filter(bicubic, 2.0)
# candidate B: the true image, shifted by one pixel. All the detail, slightly misplaced.
sharp_shifted = np.roll(hr.astype(float), 1, axis=0)
for name, cand in [("blurrier than bicubic", blurry), ("perfectly sharp, 1px off", sharp_shifted)]:
print(f"{name:26s} PSNR {psnr(hr, cand):6.2f} dB SSIM {ssim(hr, cand.astype(np.uint8), data_range=255):.3f} "
f"high-frequency energy {np.abs(np.diff(cand, axis=0)).mean():5.2f}")
print(f"{'the real image itself':26s} PSNR {'inf':>6} SSIM 1.000 "
f"high-frequency energy {np.abs(np.diff(hr.astype(float), axis=0)).mean():5.2f}")upscaler PSNR (dB) SSIM nearest 33.50 0.827 bilinear 36.03 0.890 bicubic 36.20 0.892 lanczos 36.20 0.892 --- the trade-off that makes SR benchmarks misleading --- blurrier than bicubic PSNR 35.90 dB SSIM 0.888 high-frequency energy 2.57 perfectly sharp, 1px off PSNR 31.66 dB SSIM 0.794 high-frequency energy 5.23 the real image itself PSNR inf SSIM 1.000 high-frequency energy 5.13
Reading the output
The interpolation ranking is what you would expect. Nearest is worst, bicubic and Lanczos are best and effectively tied. Nothing surprising, and this is your baseline: any learned upscaler must beat 36.20 dB here to be worth running.
Then the third block breaks the metric. A picture that has been deliberately blurred beyond bicubic scores 35.90 dB. A picture with every detail perfectly intact, shifted by one single row, scores 31.66 dB.
PSNR prefers the blur by more than 4 dB. So does SSIM, 0.888 against 0.794.
The high-frequency column shows what each candidate actually contains. The real image has 5.13. The shifted-sharp version has 5.23, essentially identical, because it is the real image. The blurred version has 2.57, half the detail.
So the metric prefers the candidate with half the texture. It rejects the one with all of it, slightly misplaced. Any model trained to minimise pixel error will learn this preference and produce soft images. That is not a flaw in the models. It is the loss function working correctly on a badly chosen objective.
What sharp models do instead
Add a term that does not compare pixels.
- Perceptual loss (Johnson et al., 2016): compare activations of a pretrained network rather than raw pixels. Two images with the same texture in slightly different positions look similar in feature space.
- Adversarial loss (SRGAN, Ledig et al., 2017): a discriminator scores whether the output looks like a real photograph.
- Both, weighted, plus a small pixel term for stability. That combination is ESRGAN and is still the recipe behind most practical upscalers.
The building block: sub-pixel convolution
Upsampling inside the network predicts extra channels at low resolution, then rearranges them. That beats interpolating, and it beats transposed convolution.
import torch, torch.nn as nn, math, time
torch.manual_seed(0)
# The sub-pixel shuffle: predict r*r channels at LOW resolution, then rearrange
# them into one channel at HIGH resolution. No interpolation, no transposed conv.
x = torch.arange(16).float().view(1, 4, 2, 2)
print("input ", tuple(x.shape))
print("output", tuple(nn.PixelShuffle(2)(x).shape), " <- 4 channels at 2x2 became 1 channel at 4x4")
print(nn.PixelShuffle(2)(x)[0, 0].int().tolist())
R = 4
def batch(n, size=64):
g = torch.arange(size).float()
yy, xx = torch.meshgrid(g, g, indexing="ij")
ph = torch.rand(n, 1, 1) * 6.28
hr = (0.5 + 0.25 * torch.sin(xx / 5 + ph) + 0.25 * torch.sin(yy / 7 + ph)).unsqueeze(1)
# A REALISTIC degradation, not a clean resize: downscale, then add sensor noise.
lr = torch.nn.functional.interpolate(hr, scale_factor=1 / R, mode="area")
lr = lr + 0.05 * torch.randn_like(lr)
return lr, hr
espcn = nn.Sequential(nn.Conv2d(1, 32, 5, padding=2), nn.Tanh(),
nn.Conv2d(32, 32, 3, padding=1), nn.Tanh(),
nn.Conv2d(32, R * R, 3, padding=1), nn.PixelShuffle(R))
opt = torch.optim.Adam(espcn.parameters(), lr=2e-3)
t0 = time.time()
for _ in range(800):
lr, hr = batch(16)
loss = ((espcn(lr) - hr) ** 2).mean()
opt.zero_grad(); loss.backward(); opt.step()
torch.manual_seed(99)
lr, hr = batch(64)
psnr = lambda a, b: 10 * math.log10(1.0 / ((a - b) ** 2).mean().item())
with torch.no_grad():
bic = torch.nn.functional.interpolate(lr, scale_factor=R, mode="bicubic", align_corners=False)
print(f"\nbicubic PSNR {psnr(bic, hr):.2f} dB")
print(f"tiny ESPCN PSNR {psnr(espcn(lr), hr):.2f} dB")
print(f"parameters {sum(p.numel() for p in espcn.parameters()):,}")
print(f"training seconds {time.time()-t0:.0f} (varies by machine)")input (1, 4, 2, 2) output (1, 1, 4, 4) <- 4 channels at 2x2 became 1 channel at 4x4 [[0, 4, 1, 5], [8, 12, 9, 13], [2, 6, 3, 7], [10, 14, 11, 15]] bicubic PSNR 26.98 dB tiny ESPCN PSNR 30.85 dB parameters 14,704 training seconds 9 (varies by machine)
Exact PSNR values depend on your PyTorch build. The ordering reproduces.
The shuffle pattern is worth reading carefully. Channel 0 supplies positions (0,0), (0,2), (2,0), (2,2), channel 1 supplies the odd columns, and so on. Every output pixel comes from exactly one input channel. That is why this has no checkerboard artefacts. Unlike transposed convolution, no output pixel sums an uneven number of contributions.
The learned model beat bicubic by 3.9 dB, and the reason matters. In the earlier script the degradation was a clean bicubic downscale. A tiny learned model loses to bicubic there, because there is nothing left to learn. Here the degradation also adds sensor noise, so the model can denoise as well as upscale. That is the difference between a toy benchmark and a real photograph. It is exactly why Real-ESRGAN's contribution was a degradation pipeline, not an architecture.
Common mistakes
Training on bicubic-downscaled data and deploying on real photos. Real low-resolution images carry compression artefacts, sensor noise, motion blur and prior resizing. A model trained only on clean downscales sharpens the artefacts confidently. This is the single most common reason an upscaler works on the benchmark and fails on your images.
Reporting PSNR only. See the first table. Report PSNR, SSIM, LPIPS and a no-reference score together, or say plainly that you optimised for one of them.
Upscaling before, rather than after, other processing. Upscale last. Denoise, colour-correct and crop on the small image. Everything is cheaper there, and the artefacts are not yet magnified.
Applying face-restoration models to non-faces. GFPGAN and similar models have a strong face prior. Run them on a hand or a plate of food and you may get a face.
Ignoring tile seams. Large images are upscaled in tiles for memory. Without overlap and blending, tile boundaries show as faint grid lines.
Try it yourself
Remove the lr = lr + 0.05 * torch.randn_like(lr) line and retrain. Bicubic will now beat the learned model. That one line is the whole argument for realistic degradation modelling. Running it both ways fixes the idea permanently.
What to learn next
- Watermarking and provenance — proving where an image came from, once it can be invented.
- Inpainting and outpainting — the other restoration task, with the same honesty problem.
- Convolutional neural networks — the backbone underneath every method here.
Researcher — Mathematics and papers.
The problem is one-to-many
Take a low-resolution $\mathbf{y} = (\mathbf{x} \otimes \mathbf{k}) \downarrow_s + \mathbf{n}$. Here $\mathbf{k}$ is a blur kernel, $\downarrow_s$ is decimation by factor $s$, and $\mathbf{n}$ is noise. Many high-resolution $\mathbf{x}$ are consistent with the same $\mathbf{y}$. The task is not inversion; it is sampling or selecting from a set.
Minimising $\mathbb{E}|\mathbf{x} - \hat{\mathbf{x}}|^2$ returns the conditional mean $\mathbb{E}[\mathbf{x} \mid \mathbf{y}]$. That is the average of all consistent images, and so it is blurred by construction. This is not an approximation error. It is the exact minimiser of the stated objective.
The perception-distortion trade-off
Blau and Michaeli (2018) prove this is a fundamental bound, not an engineering limitation. Define distortion as $\mathbb{E}[\Delta(\mathbf{x}, \hat{\mathbf{x}})]$, for any full-reference measure $\Delta$. Define perceptual quality as $d(p_{\hat{\mathbf{x}}}, p_{\mathbf{x}})$, a divergence between the output distribution and the distribution of natural images. Their result:
The perception-distortion function $P(D) = \min_{p_{\hat{\mathbf{x}}|\mathbf{y}}} d(p_{\mathbf{x}}, p_{\hat{\mathbf{x}}}) \;\text{ s.t. }\; \mathbb{E}[\Delta] \le D$ is monotonically non-increasing and convex.
Improving distortion past a point necessarily worsens perceptual quality, for every distortion measure. The PIRM 2018 challenge was designed around this, scoring methods on a perception-distortion plane rather than a single axis.
The practical reading is short. A paper claiming both the best PSNR and the best perceptual quality is doing one of two things. It is either evaluating where the bound is loose, or measuring something incorrectly.
Architecture lineage
| Year | Method | Contribution |
|---|---|---|
| 2014 | SRCNN | Three convolutions on a bicubic-upsampled input |
| 2016 | ESPCN | Sub-pixel convolution; process at low resolution, upsample last |
| 2016 | VDSR | Deep network learning the residual over bicubic |
| 2017 | SRGAN / SRResNet | Perceptual plus adversarial loss; the perceptual turn |
| 2017 | EDSR | Removes batch norm from residual blocks; large gains |
| 2018 | RCAN | Channel attention, very deep residual-in-residual |
| 2018 | ESRGAN | RRDB blocks, relativistic discriminator, pre-activation perceptual features |
| 2021 | SwinIR | Swin Transformer blocks for restoration |
| 2021 | Real-ESRGAN | High-order random degradation pipeline for real images |
| 2021 | BSRGAN | Randomly shuffled degradations, same motivation |
| 2023 onward | StableSR, DiffBIR, SUPIR | Diffusion priors for blind restoration |
EDSR's batch-norm removal is worth stating because it generalises to all restoration. Batch norm normalises away the per-image contrast and colour statistics. In classification that is a feature. In restoration it is destruction of the signal. Restoration networks use no normalisation, or layer norm.
Real-ESRGAN and BSRGAN are the practical turning point. Both argue the architecture was never the bottleneck. Models trained on bicubic downscales fail on real photographs. The degradation assumed at training time does not match reality. Their solution is a randomised composition of blur, resize, noise and JPEG compression, applied in several rounds. The NTIRE 2026 mobile real-world super-resolution challenge report confirms this remains standard. Entries follow Real-ESRGAN-style degradation simulation and BasicSR engineering practice as the stable baseline.
Diffusion-based restoration
Diffusion priors sit at the far perceptual end of the curve. StableSR, DiffBIR and SUPIR condition a pretrained text-to-image diffusion model on the degraded input. They produce textures no regression model can match. The cost is low fidelity, high compute, and hallucinated content.
The 2026 literature reports hybrid designs converging on this. They combine the structural reliability of GAN-based methods with the texture generation of diffusion models. Residual refinement holds the two ends together. Evaluation in these challenges leans on no-reference metrics: NIQE, NRQM, PI, CLIP-IQA, MUSIQ. The full-reference metrics penalise the behaviour these methods are designed for.
Evaluation, and what each metric actually rewards
- PSNR: pure distortion. Rewards the conditional mean. Correlates poorly with human judgement above about 30 dB.
- SSIM / MS-SSIM: local structure comparison. Better than PSNR, still a distortion measure and still on the wrong side of the bound.
- LPIPS (Zhang et al., 2018): distance in the features of a pretrained network, calibrated against human judgements. Better correlated, and it is still a reference metric, so it retains a distortion character.
- NIQE, PI, MUSIQ, CLIP-IQA: no-reference. They score whether the output looks like a natural image, with no ground truth involved. A model can score well here while inventing content freely, which is both the point and the danger.
The honest report is a table with at least one metric from each of the three groups. Add the operating point on the perception-distortion plane.
The evidential problem
Upscaling is generation, and generated faces are drawn from the training distribution. The 2020 PULSE incident is the widely cited example. A downsampled photograph of Barack Obama upscaled to a white face. The underlying cause is structural rather than a bug. The model samples a plausible face from its prior, and the prior reflects its training data.
This is a hard constraint on deployment. Super-resolution output must not be presented as recovered evidence about an individual, in forensic, medical or identification contexts. Menon et al. (2020) added a statement to PULSE's own paper acknowledging exactly this failure mode.
Papers
- Dong et al., Image Super-Resolution Using Deep Convolutional Networks (SRCNN), 2014 — arxiv.org/abs/1501.00092
- Shi et al., Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel CNN (ESPCN), 2016 — arxiv.org/abs/1609.05158
- Ledig et al., Photo-Realistic Single Image Super-Resolution Using a GAN (SRGAN), 2017 — arxiv.org/abs/1609.04802
- Lim et al., Enhanced Deep Residual Networks for Single Image Super-Resolution (EDSR), 2017 — arxiv.org/abs/1707.02921
- Blau and Michaeli, The Perception-Distortion Tradeoff, 2018 — arxiv.org/abs/1711.06077
- Wang et al., ESRGAN, 2018 — arxiv.org/abs/1809.00219
- Liang et al., SwinIR, 2021 — arxiv.org/abs/2108.10257
- Wang et al., Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data, 2021 — arxiv.org/abs/2107.10833
- Zhang et al., Designing a Practical Degradation Model for Deep Blind Image Super-Resolution (BSRGAN), 2021 — arxiv.org/abs/2103.14006
What to learn next
- Watermarking and provenance — proving where an image came from, once it can be invented.
- Inpainting and outpainting — the other restoration task, with the same honesty problem.
- Convolutional neural networks — the backbone underneath every method here.