Image Generation and Restoration
Diffusion samplers and step counts
The sampler decides how the model gets from noise to a picture, and a good one needs twenty steps where a naive one needs a thousand.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
The sampler is your route from noise to a picture, and some routes need far fewer stops.
The analogy
Think about walking down a long staircase in the dark. You could take every single step, one at a time, and arrive safely. Slow, but safe.
Or you could take them two or three at a time, once you get a feel for the spacing. Faster, and fine, until you misjudge and stumble.
Now imagine a handrail that tells you where the next few steps are. Suddenly you can take four at a time in confidence.
Diffusion samplers are exactly this. The model was trained assuming a thousand small steps. A better sampler works out how to skip most of them safely.
Why there is a choice at all
Training taught the model one narrow skill: given a noisy picture, estimate the noise in it.
That skill says nothing about how to use it. Subtract all the noise at once? Subtract a little and repeat? Add some noise back between steps?
Those are separate decisions, made after training, and you can change them freely. That is why one downloaded model works with a dozen different samplers.
pure noise
|
[ model ] -> "here is the noise"
|
the SAMPLER decides how big a step to take with that answer
|
slightly cleaner picture
|
+----> repeat, as many times as you chooseThe two families
Random samplers add a little fresh noise back at every step. The same starting point and the same prompt give a slightly different picture each time.
Fixed samplers add nothing back. The same starting point always gives the same picture. That makes them predictable, which matters when you want to change a prompt and compare fairly.
Neither family is better. They fail in different ways, and you should know which one you are running.
The step count
More steps means smaller, safer moves. It also means more waiting, in direct proportion.
Here is what surprises people. Beyond a certain point, more steps changes nothing you can see. The picture stops improving and you keep paying.
Where that point sits depends on the sampler. A good one gets there in twenty steps. A naive one may need several hundred.
So the useful question is never "how many steps should I use". It is "how many steps does this sampler need before it stops improving".
What to actually do
Pick a sampler. Generate the same prompt at ten, twenty, thirty and fifty steps. Look at them side by side.
Find the point where you cannot tell the difference any more. Use that number for that model, and repeat the exercise when you change model.
That takes ten minutes and saves you from every argument on the internet about the correct step count.
Where you have seen this
- The "sampling steps" and "sampler" dropdowns in image tools.
- Fast preview modes that produce a rough picture in a second.
- Video generators, where the step count directly sets how long a clip takes.
Remember this
- The sampler is chosen after training, and one model works with many.
- Random samplers vary between runs, fixed ones repeat exactly.
- Beyond a certain step count you pay more and see nothing. Find that point yourself.
What to learn next
- ControlNet — steering the same sampling loop with a picture instead of words.
- Classifier-free guidance — the other dial, and why it doubles your compute.
- Diffusion models — the process all of this is solving.
Developer — Code and libraries.
Setup
pip install torch==2.5.1 scipy==1.14.1Measuring what step count actually buys
On real images this comparison is a matter of taste. On a distribution you know, it is a number. This trains one small diffusion model on a two-blob distribution. It then samples with a deterministic and a stochastic sampler, at seven step counts. Each result is scored against the real data with the Wasserstein-1 distance.
import torch, torch.nn as nn, time
from scipy.stats import wasserstein_distance
torch.manual_seed(0)
CENTRES = torch.tensor([[-2.0, 0.0], [2.0, 0.0]])
def data(n): return CENTRES[torch.randint(0, 2, (n,))] + 0.5 * torch.randn(n, 2)
T = 1000 # the model is TRAINED with 1000 steps
betas = torch.linspace(1e-4, 0.02, T)
abar = torch.cumprod(1 - betas, 0)
net = nn.Sequential(nn.Linear(3, 128), nn.SiLU(), nn.Linear(128, 128), nn.SiLU(), nn.Linear(128, 2))
def eps(x, t): return net(torch.cat([x, t[:, None].float() / T], -1))
opt = torch.optim.Adam(net.parameters(), lr=2e-3)
t0 = time.time()
for _ in range(4000):
x0 = data(256); t = torch.randint(0, T, (256,)); n = torch.randn_like(x0)
a = abar[t][:, None]
loss = ((eps(a.sqrt() * x0 + (1 - a).sqrt() * n, t) - n) ** 2).mean()
opt.zero_grad(); loss.backward(); opt.step()
print(f"trained in {time.time()-t0:.0f}s, loss {loss.item():.4f}")
@torch.no_grad()
def sample(n_steps, eta, n=4000):
torch.manual_seed(7)
ts = torch.linspace(T - 1, 0, n_steps).long() # keep only n_steps of the 1000
x = torch.randn(n, 2)
for i, t in enumerate(ts):
at = abar[t]
at_prev = abar[ts[i + 1]] if i + 1 < len(ts) else torch.tensor(1.0)
e = eps(x, torch.full((n,), t))
x0 = (x - (1 - at).sqrt() * e) / at.sqrt() # the model's guess at the clean sample
sigma = eta * ((1 - at_prev) / (1 - at) * (1 - at / at_prev)).sqrt()
x = at_prev.sqrt() * x0 + (1 - at_prev - sigma ** 2).clamp(min=0).sqrt() * e
if i + 1 < len(ts): x = x + sigma * torch.randn_like(x)
return x
ref = data(40000)[:, 0].numpy()
print("\nsteps DDIM (eta=0) DDPM-style (eta=1)")
print(" W1 dist seconds W1 dist seconds")
for s in (5, 10, 25, 50, 100, 250, 1000):
row = []
for eta in (0.0, 1.0):
t1 = time.time(); out = sample(s, eta); el = time.time() - t1
row.append((wasserstein_distance(out[:, 0].numpy(), ref), el))
print(f"{s:5d} {row[0][0]:7.4f} {row[0][1]:7.2f} {row[1][0]:7.4f} {row[1][1]:7.2f}")trained in 9s, loss 0.2692
steps DDIM (eta=0) DDPM-style (eta=1)
W1 dist seconds W1 dist seconds
5 0.5240 0.03 0.4740 0.01
10 0.2851 0.02 0.1839 0.02
25 0.2295 0.04 0.0943 0.04
50 0.2258 0.07 0.1133 0.07
100 0.2251 0.14 0.1064 0.15
250 0.2250 0.35 0.1056 0.35
1000 0.2251 1.42 0.0962 1.62The timings are from one laptop CPU and will differ on yours. The W1 values depend on the trained weights and so on your PyTorch build. The two patterns below are the reproducible content.
Reading the output
Both samplers stop improving at 25 steps. DDIM goes 0.2295 at 25 steps and 0.2251 at a thousand. That is a 40x compute increase for a change in the fourth decimal place. Everything between 25 and 1000 steps was wasted.
This is the finding that matters most, and it is why "add more steps" is bad advice.
The sampler choice matters more than the step count. At 25 steps the stochastic sampler scores 0.0943 against the deterministic 0.2295. That is better by a factor of two. No amount of extra deterministic steps closes that gap. Changing the sampler is a free improvement; changing the step count is a paid one.
The deterministic sampler converges to a slightly wrong answer. Its curve flattens at 0.225 and stays there. That is not a step-count problem, it is a bias. With a perfect noise estimator both routes reach the same distribution. With an imperfect one they do not. The noise injected by the stochastic sampler partly corrects errors made at earlier steps. Karras et al. (2022) analyse exactly this.
Five steps is not enough for either. 0.52 and 0.47 against 0.22 and 0.09. The very-few-step regime needs a different tool, covered below.
The eta parameter is the whole difference between the two columns. The same twelve lines of sampling code produce DDIM at eta=0 and DDPM-like behaviour at eta=1. Every sampler in every image tool is a variation on that update rule.
What the samplers in your dropdown actually are
Once the model gives you a noise estimate, generation is solving a differential equation backwards. Each named sampler is a different numerical solver.
| Name in the UI | What it is | Typical steps |
|---|---|---|
| DDPM | Ancestral sampling, the original | 250 to 1000 |
| DDIM | The deterministic version, eta=0 | 20 to 50 |
| Euler | First-order ODE solver | 20 to 30 |
| Euler a | Euler with noise re-injected, "ancestral" | 20 to 40 |
| Heun | Second-order, two model calls per step | 15 to 25 |
| DPM++ 2M | Second-order multistep, reuses the previous call | 15 to 25 |
| DPM++ 2M SDE | The same, with a stochastic term | 15 to 30 |
| UniPC | Predictor-corrector, unified framework | 10 to 20 |
| LCM | For consistency-distilled models only | 2 to 8 |
The diffusers documentation recommends DPM++ 2M SDE with Karras sigmas as a general-purpose default. It suggests Euler or Euler ancestral for anime-style outputs. Flow-matching schedulers go with models trained by flow matching. Distilled models need their own schedulers.
Swapping the scheduler in diffusers
# written against diffusers 0.40.0
from diffusers import DiffusionPipeline, DPMSolverMultistepScheduler
import torch
pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float16, device_map="cuda"
)
pipeline.scheduler = DPMSolverMultistepScheduler.from_config(
pipeline.scheduler.config, algorithm_type="sde-dpmsolver++", use_karras_sigmas=True
)
image = pipeline("a cat wearing a jacket", num_inference_steps=20).images[0]No output block here. This needs a GPU and a multi-gigabyte download, and any timing printed here would be invented. Run it yourself and time it on your hardware.
Two details in that snippet are worth noting. from_config rather than from_pretrained keeps the noise schedule the model was trained with while changing only the solver. And use_karras_sigmas=True reshapes the noise levels to cluster where structure forms. That is a different lever from the step count.
Timestep spacing, a real gotcha
The diffusers docs list three spacing strategies. leading gives evenly spaced steps. linspace includes the first and last steps and divides the rest evenly. trailing includes the last step and works backwards from the end. The documentation notes that trailing typically produces higher-quality images at low step counts.
At 50 steps you will not notice. At 5 steps, the spacing choice matters more than the sampler.
Common mistakes
Raising the step count to fix a bad image. The table above shows where that stops working. If 25 steps and 100 steps look the same, the problem is elsewhere.
Comparing samplers at a fixed step count. A second-order sampler calls the model twice per step. Heun at 20 steps costs the same as Euler at 40. Compare at equal model calls, not equal steps.
Changing the sampler and keeping the seed, then calling it an improvement. Different samplers follow different trajectories from the same seed, so the images differ in composition too. Compare across several seeds.
Using an ancestral sampler and expecting reproducibility. Euler a and any SDE variant inject fresh noise. Fixing the seed helps, but any change to the step count changes the whole result.
Applying a normal sampler to an LCM or Turbo model. Distilled models have their own schedulers and their own step ranges. Mismatch produces grey mush.
Try it yourself
Add eta=0.5 as a third column and rerun. Find where it lands relative to the two extremes. Then change torch.linspace(T-1, 0, n_steps) to a quadratic spacing and see how much the 5-step row improves.
What to learn next
- ControlNet — steering the same sampling loop with a picture instead of words.
- Classifier-free guidance — the other dial, and why it doubles your compute.
- Diffusion models — the process all of this is solving.
Researcher — Mathematics and papers.
Sampling as solving a differential equation
Song et al. (2021) show the forward noising process is a stochastic differential equation. It has an associated deterministic probability flow ODE, with the same marginals at every $t$:
$$ \mathrm{d}\mathbf{x} = \left[\mathbf{f}(\mathbf{x}, t) - \tfrac{1}{2} g(t)^2 \nabla_{\mathbf{x}} \log p_t(\mathbf{x})\right]\mathrm{d}t $$
$\mathbf{f}$ is the drift, $g$ the diffusion coefficient, and $\nabla_{\mathbf{x}} \log p_t(\mathbf{x})$ the score, which the trained network supplies. Generation is integrating this backwards from $t = T$ to $t = 0$. Every named sampler is a numerical scheme for that integral. The number of function evaluations (NFE) is the honest cost axis.
DDIM and the eta parameter
Song et al. (2021), Denoising Diffusion Implicit Models, generalise DDPM to a family of non-Markovian processes sharing the same training objective:
$$ \mathbf{x}{t-1} = \sqrt{\bar{\alpha}{t-1}} \underbrace{\left(\frac{\mathbf{x}_t - \sqrt{1 - \bar{\alpha}t}\,\boldsymbol{\epsilon}\theta}{\sqrt{\bar{\alpha}t}}\right)}{\text{predicted } \mathbf{x}0} + \underbrace{\sqrt{1 - \bar{\alpha}{t-1} - \sigma_t^2}\, \boldsymbol{\epsilon}\theta}{\text{direction to } \mathbf{x}_t} + \sigma_t \boldsymbol{\epsilon} $$
with
$$ \sigma_t = \eta \sqrt{\frac{1 - \bar{\alpha}_{t-1}}{1 - \bar{\alpha}_t}} \sqrt{1 - \frac{\bar{\alpha}t}{\bar{\alpha}{t-1}}} $$
$\eta = 1$ recovers DDPM ancestral sampling; $\eta = 0$ is fully deterministic and is an Euler discretisation of the probability flow ODE. The key structural property is this. The objective does not depend on the forward process being Markovian. A model trained as DDPM can therefore be sampled as DDIM, with no retraining.
Determinism also gives DDIM an approximate inverse, called DDIM inversion. It maps a real image to the latent that reproduces it. That is the basis of most training-free editing methods.
Higher-order solvers
Lu et al. (2022), DPM-Solver, exploit the semi-linear structure of the probability flow ODE. Changing variables to $\lambda_t = \log(\alpha_t / \sigma_t)$, the half-log-SNR, the linear part integrates exactly and only a nonlinear term needs approximating:
$$ \mathbf{x}_t = \frac{\alpha_t}{\alpha_s}\mathbf{x}_s - \alpha_t \int_{\lambda_s}^{\lambda_t} e^{-\lambda} \hat{\boldsymbol{\epsilon}}\theta(\hat{\mathbf{x}}\lambda, \lambda) \, \mathrm{d}\lambda $$
Taylor expanding the integrand to order $k$ gives DPM-Solver-$k$. The exact treatment of the linear part buys the order-of-magnitude reduction in NFE. That is why these solvers dominate the field. DPM-Solver++ (2022) adapts this to the guided, $\mathbf{x}_0$-prediction setting where high guidance scales made the original unstable. UniPC (Zhao et al., 2023) adds a corrector step in a unified predictor-corrector framework.
Karras et al. (2022) separate the design into independent choices. Those are preconditioning, noise schedule, time discretisation, and solver. Their paper is Elucidating the Design Space of Diffusion-Based Generative Models, known as EDM. Their contributions that entered every codebase:
- The Karras sigma schedule, $\sigma_i = \left(\sigma_{\max}^{1/\rho} + \frac{i}{N-1}\left(\sigma_{\min}^{1/\rho} - \sigma_{\max}^{1/\rho}\right)\right)^{\rho}$ with $\rho = 7$, which places steps where curvature is highest.
- Heun's second-order method as the default deterministic solver.
- An analysis of stochasticity. Added noise corrects accumulated discretisation and score-estimation error, the effect visible in the developer table above. Too much of it over-smooths.
Few-step generation
Below roughly ten steps, better solvers stop being enough and the model itself must change.
- Progressive distillation (Salimans and Ho, 2022). A student learns to take one step where the teacher took two, halving NFE per round.
- Consistency models (Song et al., 2023). Train $f_\theta(\mathbf{x}_t, t)$ to map any point on a trajectory directly to its origin. Self-consistency is enforced along the trajectory. One to four steps.
- Latent Consistency Models (Luo et al., 2023). Consistency distillation applied in latent space. An LCM-LoRA converts existing checkpoints.
- Adversarial Diffusion Distillation (Sauer et al., 2023, SDXL-Turbo): adds an adversarial loss to distillation, reaching one to four steps.
- Rectified flow (Liu et al., 2023) and the reflow procedure straighten the ODE trajectories at training time. SD 3 therefore needs fewer steps by construction, not by distillation.
The consistent finding across all of these: quality per step improves, and sample diversity falls. Distilled few-step models are measurably less varied than their teachers. That is the same trade guidance makes, reached from a different direction.
Papers
- Song et al., Denoising Diffusion Implicit Models, 2021 — arxiv.org/abs/2010.02502
- Song et al., Score-Based Generative Modeling through SDEs, 2021 — arxiv.org/abs/2011.13456
- Lu et al., DPM-Solver, 2022 — arxiv.org/abs/2206.00927
- Lu et al., DPM-Solver++, 2022 — arxiv.org/abs/2211.01095
- Karras et al., Elucidating the Design Space of Diffusion-Based Generative Models, 2022 — arxiv.org/abs/2206.00364
- Zhao et al., UniPC, 2023 — arxiv.org/abs/2302.04867
- Song et al., Consistency Models, 2023 — arxiv.org/abs/2303.01469
- Liu et al., Flow Straight and Fast (Rectified Flow), 2023 — arxiv.org/abs/2209.03003
- Sauer et al., Adversarial Diffusion Distillation, 2023 — arxiv.org/abs/2311.17042
What to learn next
- ControlNet — steering the same sampling loop with a picture instead of words.
- Classifier-free guidance — the other dial, and why it doubles your compute.
- Diffusion models — the process all of this is solving.