Generative AI

Diffusion models

A diffusion model makes an image by starting from pure noise and removing a little of it at a time, having first learned what noise looks like by adding it to real pictures on purpose.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it was invented
  4. How it works
  5. The idea worth pausing on
  6. Where the words come in
  7. Where you have already seen it
  8. What is honestly hard
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A diffusion model builds a picture by starting with pure static and removing a little of the mess at a time.

The analogy you have already lived

Picture an old photograph that has been sitting in a drawer, covered in dust. You brush a little off. More of the picture shows. You brush again, and again, and after many passes the photo is clear.

Now the clever bit. How would you teach someone to brush dust off photos well? You would take clean photos and sprinkle dust on them yourself, a little at a time. At every stage you write down exactly what you added.

Then you hand them a dusty photo and ask: how much dust is on this, and where? Because you sprinkled it, you know the right answer, so you can correct them.

That is a diffusion model, completely. The dust is called noise. The sprinkling is training. The brushing is generation.

Why it was invented

Before diffusion, the best image generators were GANs. Two networks are locked in a contest. One makes fakes, the other tries to catch them. The GANs lesson covers how they work.

GANs made sharp pictures and they were painful to train. The contest could collapse. One network could overpower the other and training would stall with no useful warning. Worse, a GAN could get lazy and cover a narrow slice of what it had seen. Good faces, but always the same kind of face.

Diffusion replaced the contest with a task that has a correct answer. The model is asked "what noise did I add here?". There is a right answer, because the noise was added on purpose. That single change made training stable and made the output cover far more variety.

The trade is honest: diffusion is much slower to generate with. You are about to see why.

How it works

There are two directions, and only one of them involves any learning.

Forward — destroying, no learning needed.

   real photo  ->  a bit noisy  ->  noisier  ->  noisier  ->  pure static
     step 0         step 100       step 400    step 700      step 1000

You take a real picture and add a small amount of noise. Then a bit more. After enough steps, nothing of the original remains. This needs no model at all. It is a recipe.

Reverse — building, this is the learned part.

   pure static  ->  slightly less  ->  less  ->  less  ->  a picture
     step 1000       step 700        step 400   step 100     step 0

Start with pure static. Ask the model: what noise is in this? Subtract a little of what it says. Ask again. Repeat, often twenty to a thousand times.

An image emerges that was never in the training set. It emerged from the static, guided at every step by a model that learned what "less noisy" looks like.

The idea worth pausing on

The training data creates itself.

Nobody labelled anything. You take clean images, mess them up yourself, and the mess you added is the answer key. That is why diffusion models could be trained on enormous image collections without an army of annotators.

If that feels like a trick, that is because it is one, and it is the good kind. Read the paragraph again. It is the reason this whole family of models exists.

Where the words come in

You typed "a cat sitting on a Mumbai balcony at sunset". Where does that enter?

At every denoising step, the model is given your text alongside the noisy image. So the question at each step changes. It is no longer "what noise is in this picture". It becomes "what noise is in this picture, if it is meant to become a cat on a balcony".

Your words do not describe a target. They steer a thousand small decisions, each nudging the static a little further toward one kind of picture.

Where you have already seen it

  • Image generators. Every well-known text-to-image tool is a diffusion model or a close relative.
  • The "remove object" tool in your phone's gallery. You circle a person, they vanish and the background fills in. That is diffusion filling the hole.
  • Night mode and photo denoising. A dark, grainy photo cleaned up — the same denoising machinery, aimed at real noise.
  • Upscaling. Turning a small blurry image into a large sharp one.
  • Video and music generation. The same idea applied to frames or to sound.

What is honestly hard

It is slow. A language model produces one small piece of text each time it runs. A diffusion model may have to run dozens of times for a single image. Enormous effort has gone into cutting that number down, and it remains the main cost.

It repeats its training data more than you would like. Researchers have pulled specific training images back out of a trained model, more or less intact. That is a live legal and ethical problem, not a settled one.

It does not understand your sentence. Ask for "a red cube on top of a blue sphere" and you may get the colours swapped. Counting, spatial relations and text inside images are all genuinely weak. This has improved and is not solved.

Remember this

  • Diffusion adds noise on purpose to make its own training data, then learns to take it away.
  • Generation starts from pure static and walks backwards in many small steps.
  • Your prompt steers each step; it is not a description handed over once at the start.

What to learn next

Developer — Code and libraries.

We are going to train a real diffusion model. Not a toy analogy — the actual DDPM training loop and the actual sampling loop, in NumPy, on an eight-by-eight picture.

It runs in about six seconds on a laptop CPU. No GPU, no download, no framework.

Setup

bash
pip install numpy

The whole thing

diffusion.py
import numpy as np

T      = 40                                   # number of noising steps
betas  = np.linspace(1e-4, 0.2, T)            # how much noise each step adds
alphas = 1.0 - betas
abar   = np.cumprod(alphas)                   # lets us jump straight to any step

ring = np.array([
    [0,0,1,1,1,1,0,0], [0,1,0,0,0,0,1,0], [1,0,0,0,0,0,0,1], [1,0,0,0,0,0,0,1],
    [1,0,0,0,0,0,0,1], [1,0,0,0,0,0,0,1], [0,1,0,0,0,0,1,0], [0,0,1,1,1,1,0,0]],
    dtype=float)
x0 = (ring * 2 - 1).ravel()                   # rescale to -1..+1, the usual convention

def picture(v):
    chars = " .:-=+*#@"
    g = v.reshape(8, 8)
    g = (g - g.min()) / (g.max() - g.min() + 1e-9)
    return ["".join(chars[int(c * (len(chars) - 1))] for c in row) for row in g]

def strip(title, labels, images):
    print(title)
    print("   ".join(f"{l:<8}" for l in labels))
    for r in range(8):
        print("   ".join(im[r] for im in images))
    print()

# ---- forward process: pure noise addition, nothing is learned here ----
vrng, steps = np.random.default_rng(0), [0, 5, 15, 39]
strip("FORWARD - destroying the picture",
      [f"t={t}" for t in steps],
      [picture(np.sqrt(abar[t]) * x0 + np.sqrt(1 - abar[t]) * vrng.standard_normal(64))
       for t in steps])

# ---- the denoiser: one hidden layer, trained to name the noise it sees ----
trng = np.random.default_rng(1)
H  = 256
W1 = trng.normal(0, 0.1, (65, H)); b1 = np.zeros(H)
W2 = trng.normal(0, 0.1, (H, 64)); b2 = np.zeros(64)

def net(X):                                   # X has shape (batch, 65)
    Hh = np.tanh(X @ W1 + b1)
    return Hh, Hh @ W2 + b2

B, lr = 64, 0.05
for _ in range(4000):
    t   = trng.integers(0, T, size=B)
    eps = trng.standard_normal((B, 64))       # the noise we are about to hide
    xt  = np.sqrt(abar[t])[:, None] * x0 + np.sqrt(1 - abar[t])[:, None] * eps
    X   = np.hstack([xt, (t / T)[:, None]])   # the model must be told how noisy this is
    Hh, pred = net(X)
    G   = 2 * (pred - eps) / (B * 64)         # gradient of mean squared error
    dH  = (G @ W2.T) * (1 - Hh ** 2)
    W2 -= lr * (Hh.T @ G); b2 -= lr * G.sum(0)
    W1 -= lr * (X.T @ dH); b1 -= lr * dH.sum(0)

# ---- reverse process: start from noise, subtract the predicted noise, repeat ----
x, snaps = trng.standard_normal(64), {}
for t in range(T - 1, -1, -1):
    _, eps_hat = net(np.hstack([x, [t / T]])[None, :])
    mean = (x - betas[t] / np.sqrt(1 - abar[t]) * eps_hat[0]) / np.sqrt(alphas[t])
    x = mean + (np.sqrt(betas[t]) * trng.standard_normal(64) if t > 0 else 0.0)
    if t in (39, 25, 12, 0):
        snaps[t] = x.copy()

strip("REVERSE - building it back out of noise",
      [f"t={t}" for t in (39, 25, 12, 0)],
      [picture(snaps[t]) for t in (39, 25, 12, 0)])
Output
FORWARD - destroying the picture
t=0        t=5        t=15       t=39
  ####     ..###+.:   :.=**#-=   +--*:--=
 #    #    :*- .:*-   =# =--*-   --.=+-=:
#      #   *..  ::#   *: ::=:+   -+=#-=-+
#      #   *:.:...#   =:-.-:.:   =-:#-.-+
#      #   *.  ::.*   *:::+---   -++--:-:
#      #   # .: ..*   #==-=--+   +*::.: :
 #    #    .*:.::#:   .# ::.*:   +-+-*=-#
  ####     ::*+** .   :-+#*+ .   -:==+-++

REVERSE - building it back out of noise
t=39       t=25       t=12       t=0
 :-:+*=-   :.*#+#--   . ##** .     ####
#+=-.-#-   .*::::=:   :# .::#.   .#.  .#.
+==.=:+*   +:=: :.+   * : . :#   # .   .#
*::+=.--   #..--. #   *  ....*   #  . ..*
=---.*=*   #.:- .:*   # .  ..*   # .    #
==+==-+#   #.-. --#   * :... *   #     .*
=+-.++:+    *:.=.#.    *....*.   .#   .*.
.==-+*--   ::##+#:-    .***# .     #### .

Read the second strip left to right. t=39 is pure noise. By t=12 a ring is visible through the grain. At t=0 the ring is back, slightly rough, built entirely out of static by a network with 33,000 parameters.

The four ideas that make it work

The model predicts the noise, not the picture. This is the choice that surprises everyone. G = 2 * (pred - eps) compares the model's output against eps, the noise, not against x0.

Why? Because the noise is the same kind of thing at every timestep — always standard normal, always the same scale. The clean image is not: at t=39 almost nothing of it survives, so asking the model to output it directly makes the target vary wildly in difficulty. Predicting noise gives the network a target with stable statistics, and it trains far better as a result. Ho et al. reported this in 2020 and it has been the default ever since.

The model is told how noisy the input is. Look at np.hstack([xt, (t / T)[:, None]]) — the timestep is appended as an extra input. Without it, the model cannot know whether it is looking at a slightly grainy picture or near-total static, and those need completely different answers. Delete that column and re-run. The output collapses into grey mush.

One line jumps to any noise level. sqrt(abar[t]) * x0 + sqrt(1 - abar[t]) * eps produces x_t directly, without looping through t steps. This closed form is the reason training is cheap: each training example picks a random timestep and jumps straight to it.

The reverse step subtracts a scaled fraction, not everything. betas[t] / sqrt(1 - abar[t]) is far smaller than one. The model does not clean the image in a single go; it takes a small corrective step and re-examines. Removing all the predicted noise at once produces a blurry average of everything the model has seen.

Why there is still noise added on the way back

Look at x = mean + sqrt(betas[t]) * randn(...). We predict a cleaner image, then add a little fresh noise back in, at every step except the last.

That looks self-defeating and it is essential. Without it the reverse process is deterministic and collapses toward a single average output. The re-injected noise is what lets one model produce many different images from many different starting points. It is the same reason sampling beats greedy decoding in language models.

Why generation is slow

Training touched the network 4,000 times on batches of 64. Generation ran it 40 times sequentially, and each step depends on the one before, so none of it can be parallelised.

Real image models use 20 to 1,000 steps, with a network far larger than 33,000 parameters. That sequential chain is the entire performance story of diffusion, and every speed-up you have read about — DDIM, distillation, consistency models — is an attack on the number of steps.

Common mistakes

Training the model to output the clean image. It works badly for the reason above. Predict the noise. If you must have the clean image, derive it: x0_hat = (x_t - sqrt(1 - abar[t]) * eps_hat) / sqrt(abar[t]).

Not conditioning on the timestep. The most common bug when writing this from scratch. Symptom: training loss falls, generated output is featureless.

A noise schedule that ends too early. If abar[T-1] is not close to zero, the final training state still contains a trace of the image. At generation time you start from pure noise, which the model has never seen, and the first step is garbage. Print abar[-1] and check it is small.

Forgetting the -1..+1 rescaling. Diffusion assumes data centred at zero with roughly unit scale, because the noise is standard normal. Feeding 0..255 pixels means the signal drowns the noise at every timestep and the model learns nothing.

Expecting one image to teach generalisation. This model was trained on exactly one picture, so it can produce exactly one thing. That is correct behaviour and it makes the mechanism visible. Real models see hundreds of millions of images and interpolate between them.

Using a real model

For actual image generation, use the diffusers library.

bash
pip install diffusers transformers accelerate torch
python
from diffusers import StableDiffusionPipeline
import torch

pipe = StableDiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-2-1-base", torch_dtype=torch.float32)
image = pipe("a cat on a Mumbai balcony at sunset", num_inference_steps=25).images[0]
image.save("out.png")

The two arguments worth knowing: num_inference_steps is the length of that sequential chain, and guidance_scale controls how hard the prompt pulls. Both are explained in the researcher section.

Try it yourself

Change ring into a different eight-by-eight shape — a cross, a diagonal, your initials — and re-run. Predict first: will 4,000 training steps be enough for a shape with more detail? Check whether you were right.

Then set T = 10 and re-run without changing anything else. The reverse process now has ten chances instead of forty, and each step must remove four times as much noise. Watch the quality drop. That is the accuracy-for-speed trade every fast sampler is negotiating.

Finally, delete the timestep column from X and see the mush for yourself. It is a five-second experiment and it makes the point permanent.

What to learn next

Researcher — Mathematics and papers.

The forward process

A fixed Markov chain that adds Gaussian noise on a variance-preserving schedule (Sohl-Dickstein et al., 2015; Ho et al., 2020):

text
q(x_t | x_{t-1}) = N( x_t ; sqrt(1 - beta_t) * x_{t-1} , beta_t * I )
  • x_0 is a data sample; x_t is its noised version at step t.
  • beta_t in (0, 1) is the variance schedule, increasing in t.
  • T is the total number of steps, classically 1000.

Define alpha_t = 1 - beta_t and abar_t = product over s = 1..t of alpha_s. Marginalising the chain gives the closed form that makes training tractable:

text
q(x_t | x_0) = N( x_t ; sqrt(abar_t) * x_0 , (1 - abar_t) * I )

equivalently   x_t = sqrt(abar_t) * x_0 + sqrt(1 - abar_t) * eps,   eps ~ N(0, I)

The schedule is designed so abar_T ~= 0, making q(x_T) approximately N(0, I) regardless of the data. That is what allows sampling to start from pure noise.

Ho et al. used a linear beta schedule from 1e-4 to 0.02. Nichol & Dhariwal (2021) showed a cosine schedule destroys information more gradually at low resolution and improves log-likelihood — the linear schedule reaches near-pure noise too early, wasting steps.

The training objective

The variational bound decomposes into per-step KL terms. Ho et al.'s central practical result is that a reweighted, simplified form works better than the exact bound:

text
L_simple = E over t, x_0, eps of  || eps - eps_theta( sqrt(abar_t) * x_0 + sqrt(1 - abar_t) * eps , t ) ||^2
  • eps_theta(x_t, t) is the network, predicting the noise that was added.
  • t is drawn uniformly from {1, ..., T}.

This is the loss implemented in the developer block. It is a reweighting of the ELBO that down-weights small t, which empirically improves sample quality at the cost of likelihood.

Parameterisation choices. Predicting eps is one of three equivalent options. Predicting x_0 directly is unstable at high t. The v-parameterisation (Salimans & Ho, 2022), v = sqrt(abar_t) * eps - sqrt(1 - abar_t) * x_0, is better conditioned across the whole schedule and is standard in distillation work.

The score connection

Diffusion is score-based generative modelling in disguise. The Stein score of the noised marginal relates to the noise prediction by:

text
grad_{x_t} log q(x_t)  =  - eps_theta(x_t, t) / sqrt(1 - abar_t)

Song & Ermon (2019) arrived at the same family from denoising score matching and Langevin dynamics. Song et al. (2021) unified both under a stochastic differential equation:

text
forward:   dx = f(x, t) dt + g(t) dw
reverse:   dx = [ f(x, t) - g(t)^2 * grad_x log p_t(x) ] dt + g(t) dw_bar
  • f is the drift, g the diffusion coefficient, w a Wiener process, w_bar its time-reversal.

DDPM is the discretisation of the variance-preserving SDE. This framing is what allows arbitrary ODE and SDE solvers to be used as samplers, which is where most of the speed gains came from.

Sampling

Ancestral (DDPM):

text
x_{t-1} = ( x_t - (beta_t / sqrt(1 - abar_t)) * eps_theta(x_t, t) ) / sqrt(alpha_t)  +  sigma_t * z

with z ~ N(0, I) for t > 1 and z = 0 at t = 1. sigma_t^2 is set to beta_t or to the posterior variance beta_t * (1 - abar_{t-1}) / (1 - abar_t); the two bracket the true optimum and the difference is small.

DDIM (Song, Meng & Ermon, 2020, arXiv:2010.02502) constructs a non-Markovian forward process with the same marginals, giving:

text
x_{t-1} = sqrt(abar_{t-1}) * x0_hat  +  sqrt(1 - abar_{t-1} - sigma_t^2) * eps_theta(x_t, t)  +  sigma_t * z

where  x0_hat = ( x_t - sqrt(1 - abar_t) * eps_theta(x_t, t) ) / sqrt(abar_t)

At sigma_t = 0 sampling is deterministic and the trajectory is an ODE, which permits large step sizes. This is what reduced 1000 steps to 20–50 with acceptable quality, and it also makes the noise-to-image map invertible, enabling latent-space editing.

Higher-order solvers — DPM-Solver, DPM-Solver++, and the EDM framework of Karras et al. (2022, arXiv:2206.00364) — push this to 10–20 steps. Karras et al. is the most useful single paper for practitioners: it reparameterises the design space into preconditioning, noise distribution and sampler as independent choices, and much of the previous literature's disagreement dissolves under it.

Consistency models (Song et al., 2023, arXiv:2303.01469) train a network to map any point on the ODE trajectory directly to its endpoint, giving one- or two-step generation. Quality trails multi-step sampling but the gap has narrowed substantially.

Guidance

Classifier guidance (Dhariwal & Nichol, 2021, arXiv:2105.05233) adds the gradient of a separately trained noise-robust classifier to the score. It requires training that classifier on noisy inputs, which is a real cost.

Classifier-free guidance (Ho & Salimans, 2022, arXiv:2207.12598) removed that requirement and is now universal. Train one model with the conditioning randomly dropped, typically 10% of the time, then at sampling time extrapolate:

text
eps_guided = eps_theta(x_t, t, empty)  +  w * ( eps_theta(x_t, t, c) - eps_theta(x_t, t, empty) )
  • c is the conditioning (a text embedding); empty is the null condition.
  • w is the guidance scale. w = 1 recovers ordinary conditional sampling.

The mechanics matter for practice. It costs two forward passes per step, doubling inference cost. Raising w improves prompt adherence and reduces diversity, and past roughly 10–15 produces oversaturated, high-contrast artefacts. Guidance is a sampling-time bias, not a better model.

Latent diffusion

Rombach et al. (2022, arXiv:2112.10752) run the diffusion process in the latent space of a pretrained autoencoder rather than in pixel space. A 512x512x3 image becomes a 64x64x4 latent — roughly 48 times fewer values.

This is the change that made high-resolution text-to-image practical on consumer hardware, and it is why Stable Diffusion runs where pixel-space models did not. The cost is that the autoencoder's reconstruction error is a hard floor on output fidelity, which shows up as fine texture and small text being unrecoverable.

Architectures

The denoiser was originally a U-Net with residual blocks, self-attention at low resolutions, and timestep injection through adaptive normalisation. Cross-attention layers carry the text conditioning (Rombach et al.).

DiT (Peebles & Xie, 2023, arXiv:2212.09748) replaces the U-Net with a plain transformer over latent patches, conditioning through adaptive layer norm. It scales predictably with compute in a way the U-Net did not, and current large systems are transformer-based.

Flow matching, the current direction

Flow matching (Lipman et al., 2023, arXiv:2210.02747) and rectified flow (Liu et al., 2022, arXiv:2209.03003) train a velocity field along a straight interpolation between noise and data:

text
x_t = (1 - t) * x_0 + t * eps,     target velocity  v = eps - x_0

Straight paths are cheaper to integrate than the curved trajectories of a diffusion SDE, so fewer solver steps are needed for the same quality. The training objective is a plain regression on velocity, with no noise schedule to design.

Esser et al. (2024, arXiv:2403.03206) scaled this with a transformer backbone for Stable Diffusion 3. Flow matching is currently displacing standard diffusion in new systems, and the two are close relatives rather than rivals — diffusion is recoverable as a particular choice of path.

Memorisation, stated plainly

Carlini et al. (2023, arXiv:2301.13188) extracted near-identical copies of training images from Stable Diffusion by generating many samples per caption and clustering. Extraction rate was low in absolute terms and demonstrably non-zero, and correlated strongly with how often an image was duplicated in the training set.

Two findings matter for anyone deploying these. Diffusion models memorise more than GANs at comparable quality. And deduplicating the training corpus is the most effective known mitigation, which is a data-engineering answer to a problem often framed as a modelling one.

Key references

  • Sohl-Dickstein, J. et al. (2015). Deep Unsupervised Learning using Nonequilibrium Thermodynamics. arXiv:1503.03585
  • Song, Y. & Ermon, S. (2019). Generative Modeling by Estimating Gradients of the Data Distribution. arXiv:1907.05600
  • Ho, J., Jain, A. & Abbeel, P. (2020). Denoising Diffusion Probabilistic Models. arXiv:2006.11239
  • Song, Y. et al. (2021). Score-Based Generative Modeling through Stochastic Differential Equations. arXiv:2011.13456
  • Nichol, A. & Dhariwal, P. (2021). Improved Denoising Diffusion Probabilistic Models. arXiv:2102.09672
  • Karras, T. et al. (2022). Elucidating the Design Space of Diffusion-Based Generative Models. arXiv:2206.00364
  • Rombach, R. et al. (2022). High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752
  • Ho, J. & Salimans, T. (2022). Classifier-Free Diffusion Guidance. arXiv:2207.12598
  • Peebles, W. & Xie, S. (2023). Scalable Diffusion Models with Transformers. arXiv:2212.09748
  • Lipman, Y. et al. (2023). Flow Matching for Generative Modeling. arXiv:2210.02747
  • Carlini, N. et al. (2023). Extracting Training Data from Diffusion Models. arXiv:2301.13188
  • Esser, P. et al. (2024). Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. arXiv:2403.03206

Open problems

Compositionality. Binding attributes to the right objects — "a red cube on a blue sphere" — remains unreliable. Cross-attention maps make the failure visible: attention for "red" spreads across both objects. Test-time attention-editing methods help and none solves it, which suggests the problem sits in the text encoder's representation rather than in the sampler.

Evaluation. FID is the standard metric and it is known to be poorly correlated with human judgement at the top of the range, sensitive to the Inception feature extractor's own biases, and gameable. CLIP score measures prompt adherence and is insensitive to image quality. There is no accepted single metric, and papers select whichever one favours their method.

Few-step generation without a quality floor. Distillation and consistency training reach one to four steps with a measurable gap to the multi-step teacher. Whether that gap is fundamental — whether some quality genuinely requires iterative refinement — is unresolved and is the most interesting theoretical question in the area.

What to learn next

  • GANs — the adversarial alternative and its instabilities.
  • Autoencoders — the compression step that latent diffusion depends on.
  • Vision transformers — the backbone current systems use.