Image Generation and Restoration

Diffusion samplers and step counts

The sampler decides how the model gets from noise to a picture, and a good one needs twenty steps where a naive one needs a thousand.

On this page 9
  1. The short answer
  2. The analogy
  3. Why there is a choice at all
  4. The two families
  5. The step count
  6. What to actually do
  7. Where you have seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

The sampler is your route from noise to a picture, and some routes need far fewer stops.

The analogy

Think about walking down a long staircase in the dark. You could take every single step, one at a time, and arrive safely. Slow, but safe.

Or you could take them two or three at a time, once you get a feel for the spacing. Faster, and fine, until you misjudge and stumble.

Now imagine a handrail that tells you where the next few steps are. Suddenly you can take four at a time in confidence.

Diffusion samplers are exactly this. The model was trained assuming a thousand small steps. A better sampler works out how to skip most of them safely.

Why there is a choice at all

Training taught the model one narrow skill: given a noisy picture, estimate the noise in it.

That skill says nothing about how to use it. Subtract all the noise at once? Subtract a little and repeat? Add some noise back between steps?

Those are separate decisions, made after training, and you can change them freely. That is why one downloaded model works with a dozen different samplers.

   pure noise
       |
   [ model ] -> "here is the noise"
       |
   the SAMPLER decides how big a step to take with that answer
       |
   slightly cleaner picture
       |
       +----> repeat, as many times as you choose

The two families

Random samplers add a little fresh noise back at every step. The same starting point and the same prompt give a slightly different picture each time.

Fixed samplers add nothing back. The same starting point always gives the same picture. That makes them predictable, which matters when you want to change a prompt and compare fairly.

Neither family is better. They fail in different ways, and you should know which one you are running.

The step count

More steps means smaller, safer moves. It also means more waiting, in direct proportion.

Here is what surprises people. Beyond a certain point, more steps changes nothing you can see. The picture stops improving and you keep paying.

Where that point sits depends on the sampler. A good one gets there in twenty steps. A naive one may need several hundred.

So the useful question is never "how many steps should I use". It is "how many steps does this sampler need before it stops improving".

What to actually do

Pick a sampler. Generate the same prompt at ten, twenty, thirty and fifty steps. Look at them side by side.

Find the point where you cannot tell the difference any more. Use that number for that model, and repeat the exercise when you change model.

That takes ten minutes and saves you from every argument on the internet about the correct step count.

Where you have seen this

  • The "sampling steps" and "sampler" dropdowns in image tools.
  • Fast preview modes that produce a rough picture in a second.
  • Video generators, where the step count directly sets how long a clip takes.

Remember this

  • The sampler is chosen after training, and one model works with many.
  • Random samplers vary between runs, fixed ones repeat exactly.
  • Beyond a certain step count you pay more and see nothing. Find that point yourself.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch==2.5.1 scipy==1.14.1

Measuring what step count actually buys

On real images this comparison is a matter of taste. On a distribution you know, it is a number. This trains one small diffusion model on a two-blob distribution. It then samples with a deterministic and a stochastic sampler, at seven step counts. Each result is scored against the real data with the Wasserstein-1 distance.

samplers.py
import torch, torch.nn as nn, time
from scipy.stats import wasserstein_distance
torch.manual_seed(0)

CENTRES = torch.tensor([[-2.0, 0.0], [2.0, 0.0]])
def data(n): return CENTRES[torch.randint(0, 2, (n,))] + 0.5 * torch.randn(n, 2)

T = 1000                                     # the model is TRAINED with 1000 steps
betas = torch.linspace(1e-4, 0.02, T)
abar = torch.cumprod(1 - betas, 0)

net = nn.Sequential(nn.Linear(3, 128), nn.SiLU(), nn.Linear(128, 128), nn.SiLU(), nn.Linear(128, 2))
def eps(x, t): return net(torch.cat([x, t[:, None].float() / T], -1))

opt = torch.optim.Adam(net.parameters(), lr=2e-3)
t0 = time.time()
for _ in range(4000):
    x0 = data(256); t = torch.randint(0, T, (256,)); n = torch.randn_like(x0)
    a = abar[t][:, None]
    loss = ((eps(a.sqrt() * x0 + (1 - a).sqrt() * n, t) - n) ** 2).mean()
    opt.zero_grad(); loss.backward(); opt.step()
print(f"trained in {time.time()-t0:.0f}s, loss {loss.item():.4f}")

@torch.no_grad()
def sample(n_steps, eta, n=4000):
    torch.manual_seed(7)
    ts = torch.linspace(T - 1, 0, n_steps).long()      # keep only n_steps of the 1000
    x = torch.randn(n, 2)
    for i, t in enumerate(ts):
        at = abar[t]
        at_prev = abar[ts[i + 1]] if i + 1 < len(ts) else torch.tensor(1.0)
        e = eps(x, torch.full((n,), t))
        x0 = (x - (1 - at).sqrt() * e) / at.sqrt()      # the model's guess at the clean sample
        sigma = eta * ((1 - at_prev) / (1 - at) * (1 - at / at_prev)).sqrt()
        x = at_prev.sqrt() * x0 + (1 - at_prev - sigma ** 2).clamp(min=0).sqrt() * e
        if i + 1 < len(ts): x = x + sigma * torch.randn_like(x)
    return x

ref = data(40000)[:, 0].numpy()
print("\nsteps   DDIM (eta=0)          DDPM-style (eta=1)")
print("        W1 dist   seconds     W1 dist   seconds")
for s in (5, 10, 25, 50, 100, 250, 1000):
    row = []
    for eta in (0.0, 1.0):
        t1 = time.time(); out = sample(s, eta); el = time.time() - t1
        row.append((wasserstein_distance(out[:, 0].numpy(), ref), el))
    print(f"{s:5d}   {row[0][0]:7.4f}   {row[0][1]:7.2f}     {row[1][0]:7.4f}   {row[1][1]:7.2f}")
Output
trained in 9s, loss 0.2692

steps   DDIM (eta=0)          DDPM-style (eta=1)
        W1 dist   seconds     W1 dist   seconds
    5    0.5240      0.03      0.4740      0.01
   10    0.2851      0.02      0.1839      0.02
   25    0.2295      0.04      0.0943      0.04
   50    0.2258      0.07      0.1133      0.07
  100    0.2251      0.14      0.1064      0.15
  250    0.2250      0.35      0.1056      0.35
 1000    0.2251      1.42      0.0962      1.62

The timings are from one laptop CPU and will differ on yours. The W1 values depend on the trained weights and so on your PyTorch build. The two patterns below are the reproducible content.

Reading the output

Both samplers stop improving at 25 steps. DDIM goes 0.2295 at 25 steps and 0.2251 at a thousand. That is a 40x compute increase for a change in the fourth decimal place. Everything between 25 and 1000 steps was wasted.

This is the finding that matters most, and it is why "add more steps" is bad advice.

The sampler choice matters more than the step count. At 25 steps the stochastic sampler scores 0.0943 against the deterministic 0.2295. That is better by a factor of two. No amount of extra deterministic steps closes that gap. Changing the sampler is a free improvement; changing the step count is a paid one.

The deterministic sampler converges to a slightly wrong answer. Its curve flattens at 0.225 and stays there. That is not a step-count problem, it is a bias. With a perfect noise estimator both routes reach the same distribution. With an imperfect one they do not. The noise injected by the stochastic sampler partly corrects errors made at earlier steps. Karras et al. (2022) analyse exactly this.

Five steps is not enough for either. 0.52 and 0.47 against 0.22 and 0.09. The very-few-step regime needs a different tool, covered below.

The eta parameter is the whole difference between the two columns. The same twelve lines of sampling code produce DDIM at eta=0 and DDPM-like behaviour at eta=1. Every sampler in every image tool is a variation on that update rule.

What the samplers in your dropdown actually are

Once the model gives you a noise estimate, generation is solving a differential equation backwards. Each named sampler is a different numerical solver.

Name in the UIWhat it isTypical steps
DDPMAncestral sampling, the original250 to 1000
DDIMThe deterministic version, eta=020 to 50
EulerFirst-order ODE solver20 to 30
Euler aEuler with noise re-injected, "ancestral"20 to 40
HeunSecond-order, two model calls per step15 to 25
DPM++ 2MSecond-order multistep, reuses the previous call15 to 25
DPM++ 2M SDEThe same, with a stochastic term15 to 30
UniPCPredictor-corrector, unified framework10 to 20
LCMFor consistency-distilled models only2 to 8

The diffusers documentation recommends DPM++ 2M SDE with Karras sigmas as a general-purpose default. It suggests Euler or Euler ancestral for anime-style outputs. Flow-matching schedulers go with models trained by flow matching. Distilled models need their own schedulers.

Swapping the scheduler in diffusers

python
# written against diffusers 0.40.0
from diffusers import DiffusionPipeline, DPMSolverMultistepScheduler
import torch

pipeline = DiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float16, device_map="cuda"
)
pipeline.scheduler = DPMSolverMultistepScheduler.from_config(
    pipeline.scheduler.config, algorithm_type="sde-dpmsolver++", use_karras_sigmas=True
)
image = pipeline("a cat wearing a jacket", num_inference_steps=20).images[0]

No output block here. This needs a GPU and a multi-gigabyte download, and any timing printed here would be invented. Run it yourself and time it on your hardware.

Two details in that snippet are worth noting. from_config rather than from_pretrained keeps the noise schedule the model was trained with while changing only the solver. And use_karras_sigmas=True reshapes the noise levels to cluster where structure forms. That is a different lever from the step count.

Timestep spacing, a real gotcha

The diffusers docs list three spacing strategies. leading gives evenly spaced steps. linspace includes the first and last steps and divides the rest evenly. trailing includes the last step and works backwards from the end. The documentation notes that trailing typically produces higher-quality images at low step counts.

At 50 steps you will not notice. At 5 steps, the spacing choice matters more than the sampler.

Common mistakes

Raising the step count to fix a bad image. The table above shows where that stops working. If 25 steps and 100 steps look the same, the problem is elsewhere.

Comparing samplers at a fixed step count. A second-order sampler calls the model twice per step. Heun at 20 steps costs the same as Euler at 40. Compare at equal model calls, not equal steps.

Changing the sampler and keeping the seed, then calling it an improvement. Different samplers follow different trajectories from the same seed, so the images differ in composition too. Compare across several seeds.

Using an ancestral sampler and expecting reproducibility. Euler a and any SDE variant inject fresh noise. Fixing the seed helps, but any change to the step count changes the whole result.

Applying a normal sampler to an LCM or Turbo model. Distilled models have their own schedulers and their own step ranges. Mismatch produces grey mush.

Try it yourself

Add eta=0.5 as a third column and rerun. Find where it lands relative to the two extremes. Then change torch.linspace(T-1, 0, n_steps) to a quadratic spacing and see how much the 5-step row improves.

What to learn next

Researcher — Mathematics and papers.

Sampling as solving a differential equation

Song et al. (2021) show the forward noising process is a stochastic differential equation. It has an associated deterministic probability flow ODE, with the same marginals at every $t$:

$$ \mathrm{d}\mathbf{x} = \left[\mathbf{f}(\mathbf{x}, t) - \tfrac{1}{2} g(t)^2 \nabla_{\mathbf{x}} \log p_t(\mathbf{x})\right]\mathrm{d}t $$

$\mathbf{f}$ is the drift, $g$ the diffusion coefficient, and $\nabla_{\mathbf{x}} \log p_t(\mathbf{x})$ the score, which the trained network supplies. Generation is integrating this backwards from $t = T$ to $t = 0$. Every named sampler is a numerical scheme for that integral. The number of function evaluations (NFE) is the honest cost axis.

DDIM and the eta parameter

Song et al. (2021), Denoising Diffusion Implicit Models, generalise DDPM to a family of non-Markovian processes sharing the same training objective:

$$ \mathbf{x}{t-1} = \sqrt{\bar{\alpha}{t-1}} \underbrace{\left(\frac{\mathbf{x}_t - \sqrt{1 - \bar{\alpha}t}\,\boldsymbol{\epsilon}\theta}{\sqrt{\bar{\alpha}t}}\right)}{\text{predicted } \mathbf{x}0} + \underbrace{\sqrt{1 - \bar{\alpha}{t-1} - \sigma_t^2}\, \boldsymbol{\epsilon}\theta}{\text{direction to } \mathbf{x}_t} + \sigma_t \boldsymbol{\epsilon} $$

with

$$ \sigma_t = \eta \sqrt{\frac{1 - \bar{\alpha}_{t-1}}{1 - \bar{\alpha}_t}} \sqrt{1 - \frac{\bar{\alpha}t}{\bar{\alpha}{t-1}}} $$

$\eta = 1$ recovers DDPM ancestral sampling; $\eta = 0$ is fully deterministic and is an Euler discretisation of the probability flow ODE. The key structural property is this. The objective does not depend on the forward process being Markovian. A model trained as DDPM can therefore be sampled as DDIM, with no retraining.

Determinism also gives DDIM an approximate inverse, called DDIM inversion. It maps a real image to the latent that reproduces it. That is the basis of most training-free editing methods.

Higher-order solvers

Lu et al. (2022), DPM-Solver, exploit the semi-linear structure of the probability flow ODE. Changing variables to $\lambda_t = \log(\alpha_t / \sigma_t)$, the half-log-SNR, the linear part integrates exactly and only a nonlinear term needs approximating:

$$ \mathbf{x}_t = \frac{\alpha_t}{\alpha_s}\mathbf{x}_s - \alpha_t \int_{\lambda_s}^{\lambda_t} e^{-\lambda} \hat{\boldsymbol{\epsilon}}\theta(\hat{\mathbf{x}}\lambda, \lambda) \, \mathrm{d}\lambda $$

Taylor expanding the integrand to order $k$ gives DPM-Solver-$k$. The exact treatment of the linear part buys the order-of-magnitude reduction in NFE. That is why these solvers dominate the field. DPM-Solver++ (2022) adapts this to the guided, $\mathbf{x}_0$-prediction setting where high guidance scales made the original unstable. UniPC (Zhao et al., 2023) adds a corrector step in a unified predictor-corrector framework.

Karras et al. (2022) separate the design into independent choices. Those are preconditioning, noise schedule, time discretisation, and solver. Their paper is Elucidating the Design Space of Diffusion-Based Generative Models, known as EDM. Their contributions that entered every codebase:

  • The Karras sigma schedule, $\sigma_i = \left(\sigma_{\max}^{1/\rho} + \frac{i}{N-1}\left(\sigma_{\min}^{1/\rho} - \sigma_{\max}^{1/\rho}\right)\right)^{\rho}$ with $\rho = 7$, which places steps where curvature is highest.
  • Heun's second-order method as the default deterministic solver.
  • An analysis of stochasticity. Added noise corrects accumulated discretisation and score-estimation error, the effect visible in the developer table above. Too much of it over-smooths.

Few-step generation

Below roughly ten steps, better solvers stop being enough and the model itself must change.

  • Progressive distillation (Salimans and Ho, 2022). A student learns to take one step where the teacher took two, halving NFE per round.
  • Consistency models (Song et al., 2023). Train $f_\theta(\mathbf{x}_t, t)$ to map any point on a trajectory directly to its origin. Self-consistency is enforced along the trajectory. One to four steps.
  • Latent Consistency Models (Luo et al., 2023). Consistency distillation applied in latent space. An LCM-LoRA converts existing checkpoints.
  • Adversarial Diffusion Distillation (Sauer et al., 2023, SDXL-Turbo): adds an adversarial loss to distillation, reaching one to four steps.
  • Rectified flow (Liu et al., 2023) and the reflow procedure straighten the ODE trajectories at training time. SD 3 therefore needs fewer steps by construction, not by distillation.

The consistent finding across all of these: quality per step improves, and sample diversity falls. Distilled few-step models are measurably less varied than their teachers. That is the same trade guidance makes, reached from a different direction.

Papers

What to learn next