Image Generation and Restoration
Latent diffusion
Latent diffusion compresses the image first and runs the whole denoising process on the small version, which is why image generators fit on ordinary graphics cards.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Latent diffusion shrinks the picture first, does all the hard work on the small version, then expands it back.
The analogy
Think about sending a long message to a friend on a slow connection. You do not send every letter. You send a short summary, and your friend fills in the rest.
Or think about a builder's plan for a house. The plan is a few sheets of paper. The house is thousands of bricks. Almost every decision gets made on the paper, because paper is cheap to change.
Latent diffusion does the same. It works on the plan, not the bricks.
Why anyone bothered
The earlier diffusion models worked on the actual pixels. A picture that fits comfortably on a phone screen already holds around three-quarters of a million numbers.
Every single denoising step touched all of them. Hundreds of steps, hundreds of times through the whole picture. Training one of these needed a room full of graphics cards, and even generating one image was slow.
So the researchers asked a sharp question. How much of a photograph is actually meaningful, and how much is repetition a computer could reconstruct?
The answer turned out to be: mostly repetition.
The two-part machine
photo -> [ encoder ] -> a small grid of numbers ("the latent")
|
all the diffusion happens here
(add noise, remove noise, repeat)
|
a new small grid of numbers
|
[ decoder ]
|
v
a new photoThe encoder and decoder are trained first, once, and then frozen. They never change again. Their only job is to translate between full-size pictures and small ones.
The denoising network never sees a pixel. It only ever handles the small version.
What this bought
Running a picture generator went from needing a research lab to needing a gaming laptop. That single change is why these tools reached ordinary people at all.
It also made the picture generator smaller. The network no longer needs to learn what skin texture or grass looks like. The decoder already knows.
The price you pay
Here is the part most explanations skip. The decoder is a lossy translator.
Whatever it cannot rebuild is gone before generation even begins. That is why these models struggle with tiny text in a picture. Fine jewellery and distant faces in a crowd fail the same way.
No amount of extra denoising steps fixes it. The ceiling was set by the encoder and decoder, not by the generator.
Where you have seen this
- Stable Diffusion, and nearly every open image tool built on it.
- Video generators, which use the same trick with an extra time dimension.
- Audio generators, which compress a sound clip the same way.
Remember this
- The image is compressed first, and all the slow work happens on the compressed version.
- The encoder and decoder are trained once and frozen.
- Whatever the decoder cannot rebuild is permanently lost, and that is the real quality ceiling.
What to learn next
- Classifier-free guidance — how the prompt actually steers the result.
- Autoencoders — the compression half of this system, from first principles.
- Text to image — the full pipeline these pieces assemble into.
Developer — Code and libraries.
Setup
pip install torch==2.5.1The arithmetic, and the ceiling
import torch, torch.nn as nn, math, time
# ---------- 1. the arithmetic that makes latent diffusion worth doing ----------
print("resolution pixel values latent values (f=8, 4ch) ratio self-attention pairs")
for res in (256, 512, 1024):
pix = res * res * 3
lat_side = res // 8
lat = lat_side * lat_side * 4
print(f"{res:>7} {pix:>12,} {lat:>24,} {pix/lat:5.1f}x "
f"{res*res:>12,} vs {lat_side*lat_side:,}")
# ---------- 2. a real autoencoder, and the quality ceiling it imposes ----------
def patterns(n, size=32, fine=False):
g = torch.arange(size).float()
yy, xx = torch.meshgrid(g, g, indexing="ij")
c = torch.randint(8, size - 8, (n, 2, 1, 1)).float()
if fine: # one-pixel-wide stripes: the "small text" case
phase = torch.randint(0, 2, (n, 1, 1)).float()
img = ((xx + phase) % 2)
else: # big smooth blobs: the easy case
img = (((xx - c[:, 0]) ** 2 + (yy - c[:, 1]) ** 2) < 36).float()
return (img * 2 - 1).unsqueeze(1)
class AE(nn.Module):
def __init__(self, z=4):
super().__init__()
self.enc = nn.Sequential(nn.Conv2d(1, 32, 3, 2, 1), nn.SiLU(), # 32 -> 16
nn.Conv2d(32, 64, 3, 2, 1), nn.SiLU(), # 16 -> 8
nn.Conv2d(64, z, 3, 1, 1)) # z channels at 8x8
self.dec = nn.Sequential(nn.Conv2d(z, 64, 3, 1, 1), nn.SiLU(),
nn.Upsample(scale_factor=2), nn.Conv2d(64, 32, 3, 1, 1), nn.SiLU(),
nn.Upsample(scale_factor=2), nn.Conv2d(32, 1, 3, 1, 1))
def forward(self, x): return self.dec(self.enc(x))
def psnr(a, b): # both in [-1, 1], so the peak-to-peak range is 2
return 10 * math.log10(4.0 / ((a - b) ** 2).mean().item())
torch.manual_seed(0)
ae = AE(); opt = torch.optim.Adam(ae.parameters(), lr=2e-3)
t0 = time.time()
for s in range(400):
x = torch.cat([patterns(24), patterns(8, fine=True)]) # trained on BOTH kinds
loss = ((ae(x) - x) ** 2).mean()
opt.zero_grad(); loss.backward(); opt.step()
with torch.no_grad():
easy, hard = patterns(64), patterns(64, fine=True)
print(f"\ncompression: 1x32x32 = {32*32} numbers -> 4x8x8 = {4*8*8} numbers ({32*32/(4*8*8):.0f}x)")
print(f"reconstruction PSNR, big smooth shapes : {psnr(ae(easy), easy):.1f} dB")
print(f"reconstruction PSNR, 1-pixel stripes : {psnr(ae(hard), hard):.1f} dB")
print(f"\ntraining seconds: {time.time()-t0:.0f} (varies by machine)")resolution pixel values latent values (f=8, 4ch) ratio self-attention pairs
256 196,608 4,096 48.0x 65,536 vs 1,024
512 786,432 16,384 48.0x 262,144 vs 4,096
1024 3,145,728 65,536 48.0x 1,048,576 vs 16,384
compression: 1x32x32 = 1024 numbers -> 4x8x8 = 256 numbers (4x)
reconstruction PSNR, big smooth shapes : 24.3 dB
reconstruction PSNR, 1-pixel stripes : 12.4 dB
training seconds: 9 (varies by machine)Exact PSNR values depend on your PyTorch build. The gap between the two rows is what reproduces.
Reading the output
The compression ratio is 48x and it does not change with resolution. Downsampling by 8 in each spatial direction gives 64x fewer positions. Moving from 3 colour channels to 4 latent channels multiplies part of that back: 64 * 3 / 4 = 48.
The attention column is the bigger win. Self-attention costs grow with the square of the token count. At 512 pixels, pixel-space attention runs over 262,144 positions and latent-space attention over 4,096. That is a ratio of about 4,100 to one in the attention operation alone. This is why pixel-space diffusion at high resolution was never going to happen.
The two PSNR rows are the whole point of this page. The same trained autoencoder returns 24.3 dB on big smooth shapes and 12.4 dB on one-pixel stripes. That is the difference between a recognisable reconstruction and a smear.
This is exactly why Stable Diffusion mangles small text, distant faces and fine patterned fabric. The generator may have produced a perfect latent. The decoder cannot render it.
Diagnosing the ceiling on a real model
Before blaming your prompt, your sampler or your LoRA, run the image through the autoencoder alone. No diffusion involved:
# needs diffusers, a model download, and preferably a GPU
from diffusers import AutoencoderKL
import torch
vae = AutoencoderKL.from_pretrained("stabilityai/sdxl-vae") # written against diffusers 0.40.0
with torch.no_grad():
latent = vae.encode(image_tensor).latent_dist.sample() * vae.config.scaling_factor
back = vae.decode(latent / vae.config.scaling_factor).sampleNo output block here on purpose. This needs a model download and hardware this page cannot assume. Inventing numbers would teach you to expect something that may not happen. Run it on your own images and compare back against the input. Whatever is lost there is your ceiling for that model.
Numbers worth memorising
Real Stable Diffusion autoencoder configurations, read from the published model files:
| Model | Latent channels | Downsample factor | Scaling factor |
|---|---|---|---|
| SD 1.5 | 4 | 8 | 0.18215 |
| SDXL | 4 | 8 | 0.13025 |
| SD 3.x / FLUX line | 16 | 8 | model-specific |
The scaling_factor exists because the raw latents do not have unit variance. The diffusion noise schedule assumes they do. Multiply on encode, divide on decode. Forgetting it produces washed-out or oversaturated images, and it is a common bug in hand-written pipelines.
The move from 4 to 16 latent channels in SD 3 answers the ceiling demonstrated above. More channels means a weaker compression, a better reconstruction, and a larger latent for the denoiser to handle.
Common mistakes
Blaming the sampler for lost detail. If the autoencoder round trip destroys it, no sampler and no step count recovers it. Test the round trip first.
Mixing a VAE from one model with a UNet from another. The latent spaces are not compatible unless the checkpoints share a lineage. Outputs come out as coloured noise.
Forgetting the scaling factor. See above.
Assuming the latent is interpretable. It is not a downscaled RGB image. Visualising the first three channels produces something that looks vaguely like one. Treat it as an opaque code.
Fine-tuning the VAE alongside the UNet. Moving the decoder invalidates every latent the UNet has learned. Keep it frozen unless you intend to retrain everything.
Try it yourself
Change AE(z=4) to AE(z=1) and rerun. Watch the smooth-shape PSNR fall too. Then try z=16 and see how much of the fine-stripe gap it closes. That sweep is the design decision the Stable Diffusion 3 authors faced. They moved from 4 channels to 16.
What to learn next
- Classifier-free guidance — how the prompt actually steers the result.
- Autoencoders — the compression half of this system, from first principles.
- Text to image — the full pipeline these pieces assemble into.
Researcher — Mathematics and papers.
The two-stage formulation
Rombach et al. (2022), High-Resolution Image Synthesis with Latent Diffusion Models (CVPR 2022), split generation into perceptual compression and semantic generation.
Stage one trains an autoencoder $\mathcal{E}, \mathcal{D}$ with $z = \mathcal{E}(x)$ and $\tilde{x} = \mathcal{D}(z)$. Here $x \in \mathbb{R}^{H \times W \times 3}$ and $z \in \mathbb{R}^{h \times w \times c}$, with $h = H/f$ and $w = W/f$. The objective combines a perceptual loss, a patch-based adversarial loss, and one of two regularisers:
- KL-reg: a small KL penalty toward $\mathcal{N}(0, I)$. The weight is low enough that the model is nearer a deterministic autoencoder than a VAE. This is what Stable Diffusion uses.
- VQ-reg: a vector quantisation layer inside the decoder.
Stage two trains a diffusion model on $z$ with the standard objective:
$$ L_{\text{LDM}} = \mathbb{E}{\mathcal{E}(x), \boldsymbol{\epsilon} \sim \mathcal{N}(0,1), t}\left[\left|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}\theta(z_t, t)\right|_2^2\right] $$
The stages are trained separately. The autoencoder is frozen for stage two, which is what makes it reusable across many diffusion models.
Choosing $f$
The paper's ablation over $f \in {1, 2, 4, 8, 16, 32}$ produces the central empirical result. Small $f$ leaves the diffusion model doing perceptual compression, which it does inefficiently. Large $f$ throws away information the diffusion model then cannot recover. $f = 4$ and $f = 8$ occupy the sweet spot, with $f = 8$ the deployed choice.
The framing worth remembering: the latent space should retain the semantic variation and discard the perceptual variation. That boundary is empirical, not principled. The whole 4-channel-to-16-channel migration admits it was drawn too aggressively.
Conditioning
Conditioning enters through cross-attention in the denoising U-Net. A domain-specific encoder $\tau_\theta$ maps the condition $y$ to $\tau_\theta(y) \in \mathbb{R}^{M \times d_\tau}$, and each attention layer computes:
$$ \mathrm{Attention}(Q, K, V) = \mathrm{softmax}!\left(\frac{QK^{\top}}{\sqrt{d}}\right) V $$
with $Q = W_Q^{(i)} \varphi_i(z_t)$, $K = W_K^{(i)} \tau_\theta(y)$, $V = W_V^{(i)} \tau_\theta(y)$, where $\varphi_i(z_t)$ is the flattened intermediate feature map. Image features query, condition features supply. The same mechanism carries text, semantic maps and layout. That is why one architecture covered every task in the paper.
The reconstruction ceiling, quantified
The autoencoder's reconstruction quality is an upper bound. It caps the generative model's fidelity in any metric the decoder affects. Published SD-VAE reconstruction figures on COCO validation sit near 24 to 26 dB PSNR, at $f=8$ with 4 channels. The round trip alone measurably degrades an image before generation begins.
Three consequences that matter in practice:
- FID measured on generated images includes the autoencoder's error. A reported FID for a latent diffusion model is not comparable to a pixel-space model's FID. The autoencoder's reconstruction FID has to be reported too.
- Small high-frequency structure is the systematic failure. Text glyphs, distant faces, thin lines, moiré fabric. These are the observable symptoms of the same cause.
- Adding latent channels is the direct fix and it is not free. Sixteen channels at $f=8$ is a 12x compression instead of 48x, so the denoiser handles four times the data.
Lineage
| Model | Backbone | Latent | Note |
|---|---|---|---|
| LDM / SD 1.x | U-Net, cross-attention | 4ch, $f=8$, KL-reg | The original |
| SDXL | Larger U-Net, two text encoders | 4ch, $f=8$ | Refiner stage, size conditioning |
| SD 3 | MMDiT | 16ch, $f=8$ | Rectified flow, transformer backbone |
| FLUX | Hybrid MMDiT | 16ch, $f=8$ | Single- and double-stream blocks |
Every entry keeps the two-stage latent structure. The backbone changed from convolutional to transformer, and the compression got gentler. The core decision has held since 2022, across images, video and audio. Diffusion runs on a learned compressed representation, not on pixels.
Papers
- Rombach et al., High-Resolution Image Synthesis with Latent Diffusion Models, 2022 — arxiv.org/abs/2112.10752
- Esser et al., Taming Transformers for High-Resolution Image Synthesis (VQGAN), 2021 — arxiv.org/abs/2012.09841
- Podell et al., SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis, 2023 — arxiv.org/abs/2307.01952
- Esser et al., Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, 2024 — arxiv.org/abs/2403.03206
- Blattmann et al., Stable Video Diffusion, 2023 — arxiv.org/abs/2311.15127
What to learn next
- Classifier-free guidance — how the prompt actually steers the result.
- Autoencoders — the compression half of this system, from first principles.
- Text to image — the full pipeline these pieces assemble into.