Image Generation and Restoration
Classifier-free guidance
Guidance runs the model twice, with and without your prompt, then exaggerates the difference so the picture obeys the prompt more strongly.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Guidance asks the model twice, with and without your prompt, then exaggerates the difference.
The analogy
Imagine two cooks in a kitchen. You tell the first one "make something with lots of garlic". You tell the second one nothing at all.
Both cook. Now compare the two dishes. Whatever the first dish has more of, that is the garlic effect.
Now here is the move. Take the second dish and push it in that direction, further than the first cook went. More garlic than the garlic cook made.
That is guidance. Not "follow the prompt". More like "follow the prompt, then keep going".
Why this was needed
Early picture models had a real problem. Ask for "a red bus on a rainy street". You would get a street and rain. Maybe a bus somewhere, in some colour.
The prompt was a suggestion. The model kept drifting toward the average of everything it had ever seen.
The first fix used a second network, a separate classifier that could recognise buses, to nudge the picture. It worked, but it meant training and maintaining another model for every set of labels.
Classifier-free guidance removes that second model. During training, the prompt is dropped at random, maybe one time in ten. The model therefore learns both jobs: making pictures with a prompt, and making pictures without one.
At generation time you run it both ways and take the difference.
noisy picture -> [ model, with your prompt ] -> guess A
noisy picture -> [ model, with no prompt ] -> guess B
the difference (A minus B) is "what the prompt is doing"
final guess = B + strength x (A minus B)Set the strength to one and you get the ordinary prompted result. Set it higher and the prompt's influence is exaggerated.
The dial you have already used
Every image tool exposes this. It is called guidance scale, CFG scale, or prompt strength.
Turn it up and pictures match the prompt more closely, look sharper, and become more predictable.
Turn it up further and they go wrong. Colours burn out. Contrast goes harsh. Everything starts looking the same as everything else.
That last part surprises people. High guidance does not mean high quality. It means low variety.
The trade you are making
Low guidance: varied, surprising, often ignores half your prompt.
High guidance: obedient, sharp, and increasingly all the same picture.
Somewhere in the middle is the setting for your task. There is no universal right answer, and the good range differs between models.
What it costs
Two runs of the model per step instead of one. So guidance roughly doubles your generation time.
That is a real price. It is why fast models work so hard to build the guidance effect into one pass.
Where you have seen this
- The "guidance scale" slider in every image generator.
- Burnt, oversaturated AI images, which usually mean the dial was too high.
- Video and audio generators, which use exactly the same mechanism.
Remember this
- The model is run twice, with and without your prompt.
- The difference between the two runs is amplified, not only followed.
- Higher guidance buys obedience and costs variety.
What to learn next
- Diffusion samplers and step counts — the other dial that decides what your picture looks like.
- Prompt engineering — writing the condition that guidance amplifies.
- Temperature and sampling — the same obedience-versus-variety trade, in text.
Developer — Code and libraries.
Setup
pip install torch==2.5.1Guidance on a two-dimensional problem you can measure
Real image models make this hard to study, because "did the guidance help" becomes a judgement call. On a toy distribution where the correct answer is known, every effect becomes a number.
Two classes, each a Gaussian blob. Train a conditional denoiser with 10 percent label dropout. Then sample class 0 at several guidance scales. Measure how well the samples match the real class 0 distribution.
import torch, torch.nn as nn, math, time
torch.manual_seed(0)
CENTRES = torch.tensor([[-2.0, 0.0], [2.0, 0.0]])
NULL = 2 # the "no condition" token, id 2
def sample_data(n):
y = torch.randint(0, 2, (n,))
return CENTRES[y] + 0.5 * torch.randn(n, 2), y
T = 100
betas = torch.linspace(1e-4, 0.06, T)
alphas = 1 - betas
abar = torch.cumprod(alphas, 0)
class Eps(nn.Module):
def __init__(self, h=128):
super().__init__()
self.emb = nn.Embedding(3, h) # class 0, class 1, and the null token
self.net = nn.Sequential(nn.Linear(2 + 1 + h, h), nn.SiLU(),
nn.Linear(h, h), nn.SiLU(), nn.Linear(h, 2))
def forward(self, x, t, y):
return self.net(torch.cat([x, t[:, None].float() / T, self.emb(y)], -1))
net = Eps(); opt = torch.optim.Adam(net.parameters(), lr=2e-3)
t0 = time.time()
for step in range(3000):
x0, y = sample_data(256)
y = torch.where(torch.rand(256) < 0.1, torch.full_like(y, NULL), y) # 10% label dropout
t = torch.randint(0, T, (256,))
noise = torch.randn_like(x0)
a = abar[t][:, None]
xt = a.sqrt() * x0 + (1 - a).sqrt() * noise
loss = ((net(xt, t, y) - noise) ** 2).mean()
opt.zero_grad(); loss.backward(); opt.step()
print(f"trained in {time.time()-t0:.0f}s, final loss {loss.item():.4f}")
@torch.no_grad()
def sample(n, cls, w):
x = torch.randn(n, 2)
y = torch.full((n,), cls); null = torch.full((n,), NULL)
for i in reversed(range(T)):
t = torch.full((n,), i)
e_c, e_u = net(x, t, y), net(x, t, null)
eps = e_u + w * (e_c - e_u) # the whole of classifier-free guidance
mean = (x - betas[i] / (1 - abar[i]).sqrt() * eps) / alphas[i].sqrt()
x = mean + (betas[i].sqrt() * torch.randn_like(x) if i else 0)
return x
print("\n w on the right class mean distance from its centre spread")
for w in (0.0, 1.0, 2.0, 5.0, 10.0):
s = sample(2000, 0, w)
right = ((s - CENTRES[0]).norm(dim=1) < (s - CENTRES[1]).norm(dim=1)).float().mean()
print(f"{w:5.1f} {right:16.1%} {(s - CENTRES[0]).norm(dim=1).mean():25.3f} {s.std(0).mean():13.3f}")
real, _ = sample_data(200000)
r0 = real[:, 0] < 0
print(f"\nreal class-0 data: mean distance {(real[r0]-CENTRES[0]).norm(dim=1).mean():.3f}, "
f"spread {real[r0].std(0).mean():.3f}")trained in 9s, final loss 0.2975 w on the right class mean distance from its centre spread 0.0 49.2% 2.395 1.301 1.0 99.9% 0.617 0.495 2.0 100.0% 0.535 0.424 5.0 100.0% 0.531 0.346 10.0 100.0% 1.048 1.540 real class-0 data: mean distance 0.628, spread 0.501
Exact figures vary with your PyTorch build. Every trend below reproduces.
Reading the output, row by row
w = 0 is the unconditional model. 49.2% on the right class is a coin flip. The spread of 1.301 reflects samples landing on both blobs. The condition has been switched off entirely, which is what w = 0 means in this formula.
w = 1 is plain conditional sampling, with no guidance at all. Look at how well it matches. Mean distance 0.617 against the real 0.628. Spread 0.495 against the real 0.501. This is the honest baseline, and it is the best distributional match on the whole table.
w = 2 and w = 5 are more obedient and less truthful. Class purity is 100%. But the spread has fallen to 0.424 and then 0.346, below the real 0.501. The samples are being squeezed toward the middle of the class. In image terms: more prototypical, less varied.
w = 10 breaks. Mean distance jumps to 1.048 and spread to 1.540. The samples have been pushed past the data and out the other side. This is the numerical version of the burnt, oversaturated look you get from a guidance scale of 25.
The important honest point: guidance does not improve the model's match to the data distribution. It trades distributional accuracy for conditional obedience. That trade is usually worth making for a picture a human will judge. It is usually wrong when you need samples that represent the data.
What the formula looks like in a real pipeline
The single line that matters:
eps = eps_uncond + guidance_scale * (eps_cond - eps_uncond)In production the two forward passes are batched together:
latent_in = torch.cat([latents] * 2) # one copy per condition
embeds = torch.cat([negative_embeds, prompt_embeds]) # negative first, by convention
noise_pred = unet(latent_in, t, encoder_hidden_states=embeds).sample
uncond, cond = noise_pred.chunk(2)
noise_pred = uncond + guidance_scale * (cond - uncond)That is why guidance costs 2x compute rather than 2x wall-clock on a GPU. The batch dimension absorbs it, up to your memory limit.
Negative prompts are the same mechanism
Nothing in the formula requires the second run to be unconditional. Replace the empty prompt with "blurry, watermark, extra fingers" and you get:
eps = eps_negative + scale * (eps_positive - eps_negative)The result is pushed away from the negative prompt and toward the positive one. Negative prompts are not a separate feature. They are the unconditional slot, filled in.
Practical ranges
Guidance scales are not transferable between model families, and quoting a number without the model is meaningless. Published defaults from the model cards and pipeline defaults:
| Model family | Typical range |
|---|---|
| Stable Diffusion 1.5 / 2.1 | 7 to 9 |
| SDXL | 5 to 8 |
| SD 3.5 | 4 to 5 |
| FLUX.1 dev (distilled guidance) | 3 to 4 |
| Turbo / LCM / distilled models | 0 to 2 |
Distilled models are the interesting case. They have the guidance effect baked into the weights. Applying guidance on top of that double-counts and produces artefacts.
Common mistakes
Raising guidance to fix a prompt the model does not understand. If the concept is not in the model, amplifying its absence does nothing except burn the colours. Change the prompt or the model.
Assuming higher guidance is more accurate. The table above shows it moving away from the data. It moves toward the prototype of the class, which is a different thing.
Applying guidance to a distilled model. See above.
Comparing FID across guidance scales without saying so. Guidance improves human preference scores and worsens FID, because FID measures distributional match. Both numbers are correct and they disagree, which is exactly what the table predicts.
Forgetting that the negative prompt costs nothing extra. You are already paying for the second forward pass. An empty negative prompt wastes it.
Try it yourself
Change the label dropout from 0.1 to 0.0, retrain, and rerun the guidance sweep. With no dropout the model never learned the unconditional job. eps_uncond is then nonsense, and guidance produces garbage at every scale. That one experiment shows why the dropout is the whole trick.
What to learn next
- Diffusion samplers and step counts — the other dial that decides what your picture looks like.
- Prompt engineering — writing the condition that guidance amplifies.
- Temperature and sampling — the same obedience-versus-variety trade, in text.
Researcher — Mathematics and papers.
Classifier guidance, first
Dhariwal and Nichol (2021) introduce guidance by adding the gradient of a separately trained classifier $p_\phi(y \mid \mathbf{x}_t)$ to the score:
$$ \hat{\boldsymbol{\epsilon}}(\mathbf{x}t, t, y) = \boldsymbol{\epsilon}\theta(\mathbf{x}_t, t) - s\sqrt{1 - \bar{\alpha}t}\, \nabla{\mathbf{x}t} \log p\phi(y \mid \mathbf{x}_t) $$
$s$ is the guidance scale. This follows from $\nabla \log p(\mathbf{x}_t \mid y) = \nabla \log p(\mathbf{x}_t) + \nabla \log p(y \mid \mathbf{x}_t)$, with $s$ inserted as a free exponent on the likelihood term. Two costs follow. A classifier must be trained on noisy images at every timestep. The label set is fixed at classifier-training time.
Classifier-free guidance
Ho and Salimans (2022), Classifier-Free Diffusion Guidance, eliminate the classifier. Train a single conditional model $\boldsymbol{\epsilon}_\theta(\mathbf{x}t, t, \mathbf{c})$. The condition $\mathbf{c}$ is replaced with a learned null token $\varnothing$ with probability $p{\text{uncond}}$, typically 0.1 to 0.2. At sampling:
$$ \tilde{\boldsymbol{\epsilon}}_\theta(\mathbf{x}t, t, \mathbf{c}) = \boldsymbol{\epsilon}\theta(\mathbf{x}t, t, \varnothing) + w\left(\boldsymbol{\epsilon}\theta(\mathbf{x}t, t, \mathbf{c}) - \boldsymbol{\epsilon}\theta(\mathbf{x}_t, t, \varnothing)\right) $$
Two conventions for $w$ circulate and confusing them is a common source of bugs. In the form above (used in Stable Diffusion), $w = 1$ is plain conditional sampling and $w = 0$ is unconditional. The original paper writes $(1+w)\boldsymbol{\epsilon}c - w\boldsymbol{\epsilon}\varnothing$, in which $w = 0$ is plain conditional. Always check which one a codebase means.
Why it samples from a sharpened distribution
The implied score is:
$$ \nabla_{\mathbf{x}_t} \log \tilde{p}(\mathbf{x}t \mid \mathbf{c}) = \nabla{\mathbf{x}_t} \log p(\mathbf{x}t) + w \nabla{\mathbf{x}_t} \log p(\mathbf{c} \mid \mathbf{x}_t) $$
which corresponds to sampling from $p(\mathbf{x}_t) \, p(\mathbf{c} \mid \mathbf{x}_t)^{w}$. Raising the likelihood to a power greater than one concentrates mass where the condition is most strongly satisfied. It is a temperature on the conditional term. The resulting distribution is not the true conditional for any $w \ne 1$.
Karras et al. (2024), Guiding a Diffusion Model with a Bad Version of Itself, sharpen this analysis. Much of what CFG does at high scale is truncation toward high-density regions. It conflates two effects: improving image quality, and reducing variation. They propose autoguidance. It replaces the unconditional model with a smaller or less-trained version of the same conditional model. Quality improves without the same diversity collapse. This is the cleanest current account of what CFG is doing and what it costs.
The measured trade-off
The finding that reproduces everywhere: FID is U-shaped in $w$, Inception Score and human preference are monotonic. Guidance improves per-sample fidelity and prompt adherence while shrinking coverage, and FID is sensitive to both.
Reporting practice follows from this. A guidance sweep with FID and CLIP-score plotted against each other, the standard Pareto curve, is the honest presentation. A single FID at an unstated guidance scale is not comparable to anything.
Known failure modes and their fixes
Oversaturation and burnt contrast at high $w$. Caused by the guided prediction leaving the valid range. Lin et al. (2023) prescribe rescaling the guided output, toward the standard deviation of the conditional prediction. Their paper is Common Diffusion Noise Schedules and Sample Steps are Flawed. diffusers exposes the fix as guidance_rescale, and 0.7 is a common value.
Zero terminal SNR. The same paper shows most schedules do not reach pure noise at the final step. The model always sees a residual mean brightness. It therefore cannot produce very dark or very bright images. Fixes are rescale_betas_zero_snr=True with timestep_spacing="trailing", and a model trained with $\mathbf{v}$-prediction.
Doubled compute. Addressed by distillation. Meng et al. (2023), On Distillation of Guided Diffusion Models, train a student to match the guided output in one pass. FLUX.1 dev ships a guidance-distilled model, which is why applying external CFG to it is a mistake.
Interval guidance. Kynkäänniemi et al. (2024) show guidance is harmful at very high and very low noise levels. It is beneficial only in the middle. Restricting CFG to a timestep interval improves FID substantially, at the same prompt adherence.
Papers
- Dhariwal and Nichol, Diffusion Models Beat GANs on Image Synthesis, 2021 — arxiv.org/abs/2105.05233
- Ho and Salimans, Classifier-Free Diffusion Guidance, 2022 — arxiv.org/abs/2207.12598
- Lin et al., Common Diffusion Noise Schedules and Sample Steps are Flawed, 2023 — arxiv.org/abs/2305.08891
- Meng et al., On Distillation of Guided Diffusion Models, 2023 — arxiv.org/abs/2210.03142
- Kynkäänniemi et al., Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models, 2024 — arxiv.org/abs/2404.07724
- Karras et al., Guiding a Diffusion Model with a Bad Version of Itself, 2024 — arxiv.org/abs/2406.02507
What to learn next
- Diffusion samplers and step counts — the other dial that decides what your picture looks like.
- Prompt engineering — writing the condition that guidance amplifies.
- Temperature and sampling — the same obedience-versus-variety trade, in text.