Diffusion models
A diffusion model makes an image by starting from pure noise and removing a little of it at a time, having first learned what noise looks like by adding it to real pictures on purpose.
- 22 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A diffusion model builds a picture by starting with pure static and removing a little of the mess at a time.
The analogy you have already lived
Picture an old photograph that has been sitting in a drawer, covered in dust. You brush a little off. More of the picture shows. You brush again, and again, and after many passes the photo is clear.
Now the clever bit. How would you teach someone to brush dust off photos well? You would take clean photos and sprinkle dust on them yourself, a little at a time. At every stage you write down exactly what you added.
Then you hand them a dusty photo and ask: how much dust is on this, and where? Because you sprinkled it, you know the right answer, so you can correct them.
That is a diffusion model, completely. The dust is called noise. The sprinkling is training. The brushing is generation.
Why it was invented
Before diffusion, the best image generators were GANs. Two networks are locked in a contest. One makes fakes, the other tries to catch them. The GANs lesson covers how they work.
GANs made sharp pictures and they were painful to train. The contest could collapse. One network could overpower the other and training would stall with no useful warning. Worse, a GAN could get lazy and cover a narrow slice of what it had seen. Good faces, but always the same kind of face.
Diffusion replaced the contest with a task that has a correct answer. The model is asked "what noise did I add here?". There is a right answer, because the noise was added on purpose. That single change made training stable and made the output cover far more variety.
The trade is honest: diffusion is much slower to generate with. You are about to see why.
How it works
There are two directions, and only one of them involves any learning.
Forward — destroying, no learning needed.
real photo -> a bit noisy -> noisier -> noisier -> pure static
step 0 step 100 step 400 step 700 step 1000You take a real picture and add a small amount of noise. Then a bit more. After enough steps, nothing of the original remains. This needs no model at all. It is a recipe.
Reverse — building, this is the learned part.
pure static -> slightly less -> less -> less -> a picture
step 1000 step 700 step 400 step 100 step 0Start with pure static. Ask the model: what noise is in this? Subtract a little of what it says. Ask again. Repeat, often twenty to a thousand times.
An image emerges that was never in the training set. It emerged from the static, guided at every step by a model that learned what "less noisy" looks like.
The idea worth pausing on
The training data creates itself.
Nobody labelled anything. You take clean images, mess them up yourself, and the mess you added is the answer key. That is why diffusion models could be trained on enormous image collections without an army of annotators.
If that feels like a trick, that is because it is one, and it is the good kind. Read the paragraph again. It is the reason this whole family of models exists.
Where the words come in
You typed "a cat sitting on a Mumbai balcony at sunset". Where does that enter?
At every denoising step, the model is given your text alongside the noisy image. So the question at each step changes. It is no longer "what noise is in this picture". It becomes "what noise is in this picture, if it is meant to become a cat on a balcony".
Your words do not describe a target. They steer a thousand small decisions, each nudging the static a little further toward one kind of picture.
Where you have already seen it
- Image generators. Every well-known text-to-image tool is a diffusion model or a close relative.
- The "remove object" tool in your phone's gallery. You circle a person, they vanish and the background fills in. That is diffusion filling the hole.
- Night mode and photo denoising. A dark, grainy photo cleaned up — the same denoising machinery, aimed at real noise.
- Upscaling. Turning a small blurry image into a large sharp one.
- Video and music generation. The same idea applied to frames or to sound.
What is honestly hard
It is slow. A language model produces one small piece of text each time it runs. A diffusion model may have to run dozens of times for a single image. Enormous effort has gone into cutting that number down, and it remains the main cost.
It repeats its training data more than you would like. Researchers have pulled specific training images back out of a trained model, more or less intact. That is a live legal and ethical problem, not a settled one.
It does not understand your sentence. Ask for "a red cube on top of a blue sphere" and you may get the colours swapped. Counting, spatial relations and text inside images are all genuinely weak. This has improved and is not solved.
Remember this
- Diffusion adds noise on purpose to make its own training data, then learns to take it away.
- Generation starts from pure static and walks backwards in many small steps.
- Your prompt steers each step; it is not a description handed over once at the start.
What to learn next
- GANs — the approach diffusion replaced, and why.
- Hallucination — the equivalent problem in text models.
- What is a neural network? — the denoiser underneath all of this.
Developer — Code and libraries.
We are going to train a real diffusion model. Not a toy analogy — the actual DDPM training loop and the actual sampling loop, in NumPy, on an eight-by-eight picture.
It runs in about six seconds on a laptop CPU. No GPU, no download, no framework.
Setup
pip install numpyThe whole thing
import numpy as np
T = 40 # number of noising steps
betas = np.linspace(1e-4, 0.2, T) # how much noise each step adds
alphas = 1.0 - betas
abar = np.cumprod(alphas) # lets us jump straight to any step
ring = np.array([
[0,0,1,1,1,1,0,0], [0,1,0,0,0,0,1,0], [1,0,0,0,0,0,0,1], [1,0,0,0,0,0,0,1],
[1,0,0,0,0,0,0,1], [1,0,0,0,0,0,0,1], [0,1,0,0,0,0,1,0], [0,0,1,1,1,1,0,0]],
dtype=float)
x0 = (ring * 2 - 1).ravel() # rescale to -1..+1, the usual convention
def picture(v):
chars = " .:-=+*#@"
g = v.reshape(8, 8)
g = (g - g.min()) / (g.max() - g.min() + 1e-9)
return ["".join(chars[int(c * (len(chars) - 1))] for c in row) for row in g]
def strip(title, labels, images):
print(title)
print(" ".join(f"{l:<8}" for l in labels))
for r in range(8):
print(" ".join(im[r] for im in images))
print()
# ---- forward process: pure noise addition, nothing is learned here ----
vrng, steps = np.random.default_rng(0), [0, 5, 15, 39]
strip("FORWARD - destroying the picture",
[f"t={t}" for t in steps],
[picture(np.sqrt(abar[t]) * x0 + np.sqrt(1 - abar[t]) * vrng.standard_normal(64))
for t in steps])
# ---- the denoiser: one hidden layer, trained to name the noise it sees ----
trng = np.random.default_rng(1)
H = 256
W1 = trng.normal(0, 0.1, (65, H)); b1 = np.zeros(H)
W2 = trng.normal(0, 0.1, (H, 64)); b2 = np.zeros(64)
def net(X): # X has shape (batch, 65)
Hh = np.tanh(X @ W1 + b1)
return Hh, Hh @ W2 + b2
B, lr = 64, 0.05
for _ in range(4000):
t = trng.integers(0, T, size=B)
eps = trng.standard_normal((B, 64)) # the noise we are about to hide
xt = np.sqrt(abar[t])[:, None] * x0 + np.sqrt(1 - abar[t])[:, None] * eps
X = np.hstack([xt, (t / T)[:, None]]) # the model must be told how noisy this is
Hh, pred = net(X)
G = 2 * (pred - eps) / (B * 64) # gradient of mean squared error
dH = (G @ W2.T) * (1 - Hh ** 2)
W2 -= lr * (Hh.T @ G); b2 -= lr * G.sum(0)
W1 -= lr * (X.T @ dH); b1 -= lr * dH.sum(0)
# ---- reverse process: start from noise, subtract the predicted noise, repeat ----
x, snaps = trng.standard_normal(64), {}
for t in range(T - 1, -1, -1):
_, eps_hat = net(np.hstack([x, [t / T]])[None, :])
mean = (x - betas[t] / np.sqrt(1 - abar[t]) * eps_hat[0]) / np.sqrt(alphas[t])
x = mean + (np.sqrt(betas[t]) * trng.standard_normal(64) if t > 0 else 0.0)
if t in (39, 25, 12, 0):
snaps[t] = x.copy()
strip("REVERSE - building it back out of noise",
[f"t={t}" for t in (39, 25, 12, 0)],
[picture(snaps[t]) for t in (39, 25, 12, 0)])FORWARD - destroying the picture t=0 t=5 t=15 t=39 #### ..###+.: :.=**#-= +--*:--= # # :*- .:*- =# =--*- --.=+-=: # # *.. ::# *: ::=:+ -+=#-=-+ # # *:.:...# =:-.-:.: =-:#-.-+ # # *. ::.* *:::+--- -++--:-: # # # .: ..* #==-=--+ +*::.: : # # .*:.::#: .# ::.*: +-+-*=-# #### ::*+** . :-+#*+ . -:==+-++ REVERSE - building it back out of noise t=39 t=25 t=12 t=0 :-:+*=- :.*#+#-- . ##** . #### #+=-.-#- .*::::=: :# .::#. .#. .#. +==.=:+* +:=: :.+ * : . :# # . .# *::+=.-- #..--. # * ....* # . ..* =---.*=* #.:- .:* # . ..* # . # ==+==-+# #.-. --# * :... * # .* =+-.++:+ *:.=.#. *....*. .# .*. .==-+*-- ::##+#:- .***# . #### .
Read the second strip left to right. t=39 is pure noise. By t=12 a ring is visible through the grain. At t=0 the ring is back, slightly rough, built entirely out of static by a network with 33,000 parameters.
The four ideas that make it work
The model predicts the noise, not the picture. This is the choice that surprises everyone. G = 2 * (pred - eps) compares the model's output against eps, the noise, not against x0.
Why? Because the noise is the same kind of thing at every timestep — always standard normal, always the same scale. The clean image is not: at t=39 almost nothing of it survives, so asking the model to output it directly makes the target vary wildly in difficulty. Predicting noise gives the network a target with stable statistics, and it trains far better as a result. Ho et al. reported this in 2020 and it has been the default ever since.
The model is told how noisy the input is. Look at np.hstack([xt, (t / T)[:, None]]) — the timestep is appended as an extra input. Without it, the model cannot know whether it is looking at a slightly grainy picture or near-total static, and those need completely different answers. Delete that column and re-run. The output collapses into grey mush.
One line jumps to any noise level. sqrt(abar[t]) * x0 + sqrt(1 - abar[t]) * eps produces x_t directly, without looping through t steps. This closed form is the reason training is cheap: each training example picks a random timestep and jumps straight to it.
The reverse step subtracts a scaled fraction, not everything. betas[t] / sqrt(1 - abar[t]) is far smaller than one. The model does not clean the image in a single go; it takes a small corrective step and re-examines. Removing all the predicted noise at once produces a blurry average of everything the model has seen.
Why there is still noise added on the way back
Look at x = mean + sqrt(betas[t]) * randn(...). We predict a cleaner image, then add a little fresh noise back in, at every step except the last.
That looks self-defeating and it is essential. Without it the reverse process is deterministic and collapses toward a single average output. The re-injected noise is what lets one model produce many different images from many different starting points. It is the same reason sampling beats greedy decoding in language models.
Why generation is slow
Training touched the network 4,000 times on batches of 64. Generation ran it 40 times sequentially, and each step depends on the one before, so none of it can be parallelised.
Real image models use 20 to 1,000 steps, with a network far larger than 33,000 parameters. That sequential chain is the entire performance story of diffusion, and every speed-up you have read about — DDIM, distillation, consistency models — is an attack on the number of steps.
Common mistakes
Training the model to output the clean image. It works badly for the reason above. Predict the noise. If you must have the clean image, derive it: x0_hat = (x_t - sqrt(1 - abar[t]) * eps_hat) / sqrt(abar[t]).
Not conditioning on the timestep. The most common bug when writing this from scratch. Symptom: training loss falls, generated output is featureless.
A noise schedule that ends too early. If abar[T-1] is not close to zero, the final training state still contains a trace of the image. At generation time you start from pure noise, which the model has never seen, and the first step is garbage. Print abar[-1] and check it is small.
Forgetting the -1..+1 rescaling. Diffusion assumes data centred at zero with roughly unit scale, because the noise is standard normal. Feeding 0..255 pixels means the signal drowns the noise at every timestep and the model learns nothing.
Expecting one image to teach generalisation. This model was trained on exactly one picture, so it can produce exactly one thing. That is correct behaviour and it makes the mechanism visible. Real models see hundreds of millions of images and interpolate between them.
Using a real model
For actual image generation, use the diffusers library.
pip install diffusers transformers accelerate torchfrom diffusers import StableDiffusionPipeline
import torch
pipe = StableDiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-2-1-base", torch_dtype=torch.float32)
image = pipe("a cat on a Mumbai balcony at sunset", num_inference_steps=25).images[0]
image.save("out.png")The two arguments worth knowing: num_inference_steps is the length of that sequential chain, and guidance_scale controls how hard the prompt pulls. Both are explained in the researcher section.
Try it yourself
Change ring into a different eight-by-eight shape — a cross, a diagonal, your initials — and re-run. Predict first: will 4,000 training steps be enough for a shape with more detail? Check whether you were right.
Then set T = 10 and re-run without changing anything else. The reverse process now has ten chances instead of forty, and each step must remove four times as much noise. Watch the quality drop. That is the accuracy-for-speed trade every fast sampler is negotiating.
Finally, delete the timestep column from X and see the mush for yourself. It is a five-second experiment and it makes the point permanent.
What to learn next
- GANs — the alternative approach and its failure modes.
- Temperature and sampling — randomness at generation time, in the text world.
- PyTorch basics — the framework to rewrite this in.
Researcher — Mathematics and papers.
The forward process
A fixed Markov chain that adds Gaussian noise on a variance-preserving schedule (Sohl-Dickstein et al., 2015; Ho et al., 2020):
q(x_t | x_{t-1}) = N( x_t ; sqrt(1 - beta_t) * x_{t-1} , beta_t * I )x_0is a data sample;x_tis its noised version at stept.beta_tin(0, 1)is the variance schedule, increasing int.Tis the total number of steps, classically 1000.
Define alpha_t = 1 - beta_t and abar_t = product over s = 1..t of alpha_s. Marginalising the chain gives the closed form that makes training tractable:
q(x_t | x_0) = N( x_t ; sqrt(abar_t) * x_0 , (1 - abar_t) * I )
equivalently x_t = sqrt(abar_t) * x_0 + sqrt(1 - abar_t) * eps, eps ~ N(0, I)The schedule is designed so abar_T ~= 0, making q(x_T) approximately N(0, I) regardless of the data. That is what allows sampling to start from pure noise.
Ho et al. used a linear beta schedule from 1e-4 to 0.02. Nichol & Dhariwal (2021) showed a cosine schedule destroys information more gradually at low resolution and improves log-likelihood — the linear schedule reaches near-pure noise too early, wasting steps.
The training objective
The variational bound decomposes into per-step KL terms. Ho et al.'s central practical result is that a reweighted, simplified form works better than the exact bound:
L_simple = E over t, x_0, eps of || eps - eps_theta( sqrt(abar_t) * x_0 + sqrt(1 - abar_t) * eps , t ) ||^2eps_theta(x_t, t)is the network, predicting the noise that was added.tis drawn uniformly from{1, ..., T}.
This is the loss implemented in the developer block. It is a reweighting of the ELBO that down-weights small t, which empirically improves sample quality at the cost of likelihood.
Parameterisation choices. Predicting eps is one of three equivalent options. Predicting x_0 directly is unstable at high t. The v-parameterisation (Salimans & Ho, 2022), v = sqrt(abar_t) * eps - sqrt(1 - abar_t) * x_0, is better conditioned across the whole schedule and is standard in distillation work.
The score connection
Diffusion is score-based generative modelling in disguise. The Stein score of the noised marginal relates to the noise prediction by:
grad_{x_t} log q(x_t) = - eps_theta(x_t, t) / sqrt(1 - abar_t)Song & Ermon (2019) arrived at the same family from denoising score matching and Langevin dynamics. Song et al. (2021) unified both under a stochastic differential equation:
forward: dx = f(x, t) dt + g(t) dw
reverse: dx = [ f(x, t) - g(t)^2 * grad_x log p_t(x) ] dt + g(t) dw_barfis the drift,gthe diffusion coefficient,wa Wiener process,w_barits time-reversal.
DDPM is the discretisation of the variance-preserving SDE. This framing is what allows arbitrary ODE and SDE solvers to be used as samplers, which is where most of the speed gains came from.
Sampling
Ancestral (DDPM):
x_{t-1} = ( x_t - (beta_t / sqrt(1 - abar_t)) * eps_theta(x_t, t) ) / sqrt(alpha_t) + sigma_t * zwith z ~ N(0, I) for t > 1 and z = 0 at t = 1. sigma_t^2 is set to beta_t or to the posterior variance beta_t * (1 - abar_{t-1}) / (1 - abar_t); the two bracket the true optimum and the difference is small.
DDIM (Song, Meng & Ermon, 2020, arXiv:2010.02502) constructs a non-Markovian forward process with the same marginals, giving:
x_{t-1} = sqrt(abar_{t-1}) * x0_hat + sqrt(1 - abar_{t-1} - sigma_t^2) * eps_theta(x_t, t) + sigma_t * z
where x0_hat = ( x_t - sqrt(1 - abar_t) * eps_theta(x_t, t) ) / sqrt(abar_t)At sigma_t = 0 sampling is deterministic and the trajectory is an ODE, which permits large step sizes. This is what reduced 1000 steps to 20–50 with acceptable quality, and it also makes the noise-to-image map invertible, enabling latent-space editing.
Higher-order solvers — DPM-Solver, DPM-Solver++, and the EDM framework of Karras et al. (2022, arXiv:2206.00364) — push this to 10–20 steps. Karras et al. is the most useful single paper for practitioners: it reparameterises the design space into preconditioning, noise distribution and sampler as independent choices, and much of the previous literature's disagreement dissolves under it.
Consistency models (Song et al., 2023, arXiv:2303.01469) train a network to map any point on the ODE trajectory directly to its endpoint, giving one- or two-step generation. Quality trails multi-step sampling but the gap has narrowed substantially.
Guidance
Classifier guidance (Dhariwal & Nichol, 2021, arXiv:2105.05233) adds the gradient of a separately trained noise-robust classifier to the score. It requires training that classifier on noisy inputs, which is a real cost.
Classifier-free guidance (Ho & Salimans, 2022, arXiv:2207.12598) removed that requirement and is now universal. Train one model with the conditioning randomly dropped, typically 10% of the time, then at sampling time extrapolate:
eps_guided = eps_theta(x_t, t, empty) + w * ( eps_theta(x_t, t, c) - eps_theta(x_t, t, empty) )cis the conditioning (a text embedding);emptyis the null condition.wis the guidance scale.w = 1recovers ordinary conditional sampling.
The mechanics matter for practice. It costs two forward passes per step, doubling inference cost. Raising w improves prompt adherence and reduces diversity, and past roughly 10–15 produces oversaturated, high-contrast artefacts. Guidance is a sampling-time bias, not a better model.
Latent diffusion
Rombach et al. (2022, arXiv:2112.10752) run the diffusion process in the latent space of a pretrained autoencoder rather than in pixel space. A 512x512x3 image becomes a 64x64x4 latent — roughly 48 times fewer values.
This is the change that made high-resolution text-to-image practical on consumer hardware, and it is why Stable Diffusion runs where pixel-space models did not. The cost is that the autoencoder's reconstruction error is a hard floor on output fidelity, which shows up as fine texture and small text being unrecoverable.
Architectures
The denoiser was originally a U-Net with residual blocks, self-attention at low resolutions, and timestep injection through adaptive normalisation. Cross-attention layers carry the text conditioning (Rombach et al.).
DiT (Peebles & Xie, 2023, arXiv:2212.09748) replaces the U-Net with a plain transformer over latent patches, conditioning through adaptive layer norm. It scales predictably with compute in a way the U-Net did not, and current large systems are transformer-based.
Flow matching, the current direction
Flow matching (Lipman et al., 2023, arXiv:2210.02747) and rectified flow (Liu et al., 2022, arXiv:2209.03003) train a velocity field along a straight interpolation between noise and data:
x_t = (1 - t) * x_0 + t * eps, target velocity v = eps - x_0Straight paths are cheaper to integrate than the curved trajectories of a diffusion SDE, so fewer solver steps are needed for the same quality. The training objective is a plain regression on velocity, with no noise schedule to design.
Esser et al. (2024, arXiv:2403.03206) scaled this with a transformer backbone for Stable Diffusion 3. Flow matching is currently displacing standard diffusion in new systems, and the two are close relatives rather than rivals — diffusion is recoverable as a particular choice of path.
Memorisation, stated plainly
Carlini et al. (2023, arXiv:2301.13188) extracted near-identical copies of training images from Stable Diffusion by generating many samples per caption and clustering. Extraction rate was low in absolute terms and demonstrably non-zero, and correlated strongly with how often an image was duplicated in the training set.
Two findings matter for anyone deploying these. Diffusion models memorise more than GANs at comparable quality. And deduplicating the training corpus is the most effective known mitigation, which is a data-engineering answer to a problem often framed as a modelling one.
Key references
- Sohl-Dickstein, J. et al. (2015). Deep Unsupervised Learning using Nonequilibrium Thermodynamics. arXiv:1503.03585
- Song, Y. & Ermon, S. (2019). Generative Modeling by Estimating Gradients of the Data Distribution. arXiv:1907.05600
- Ho, J., Jain, A. & Abbeel, P. (2020). Denoising Diffusion Probabilistic Models. arXiv:2006.11239
- Song, Y. et al. (2021). Score-Based Generative Modeling through Stochastic Differential Equations. arXiv:2011.13456
- Nichol, A. & Dhariwal, P. (2021). Improved Denoising Diffusion Probabilistic Models. arXiv:2102.09672
- Karras, T. et al. (2022). Elucidating the Design Space of Diffusion-Based Generative Models. arXiv:2206.00364
- Rombach, R. et al. (2022). High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752
- Ho, J. & Salimans, T. (2022). Classifier-Free Diffusion Guidance. arXiv:2207.12598
- Peebles, W. & Xie, S. (2023). Scalable Diffusion Models with Transformers. arXiv:2212.09748
- Lipman, Y. et al. (2023). Flow Matching for Generative Modeling. arXiv:2210.02747
- Carlini, N. et al. (2023). Extracting Training Data from Diffusion Models. arXiv:2301.13188
- Esser, P. et al. (2024). Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. arXiv:2403.03206
Open problems
Compositionality. Binding attributes to the right objects — "a red cube on a blue sphere" — remains unreliable. Cross-attention maps make the failure visible: attention for "red" spreads across both objects. Test-time attention-editing methods help and none solves it, which suggests the problem sits in the text encoder's representation rather than in the sampler.
Evaluation. FID is the standard metric and it is known to be poorly correlated with human judgement at the top of the range, sensitive to the Inception feature extractor's own biases, and gameable. CLIP score measures prompt adherence and is insensitive to image quality. There is no accepted single metric, and papers select whichever one favours their method.
Few-step generation without a quality floor. Distillation and consistency training reach one to four steps with a measurable gap to the multi-step teacher. Whether that gap is fundamental — whether some quality genuinely requires iterative refinement — is unresolved and is the most interesting theoretical question in the area.
What to learn next
- GANs — the adversarial alternative and its instabilities.
- Autoencoders — the compression step that latent diffusion depends on.
- Vision transformers — the backbone current systems use.