Image Generation and Restoration

Classifier-free guidance

Guidance runs the model twice, with and without your prompt, then exaggerates the difference so the picture obeys the prompt more strongly.

On this page 9
  1. The short answer
  2. The analogy
  3. Why this was needed
  4. The dial you have already used
  5. The trade you are making
  6. What it costs
  7. Where you have seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Guidance asks the model twice, with and without your prompt, then exaggerates the difference.

The analogy

Imagine two cooks in a kitchen. You tell the first one "make something with lots of garlic". You tell the second one nothing at all.

Both cook. Now compare the two dishes. Whatever the first dish has more of, that is the garlic effect.

Now here is the move. Take the second dish and push it in that direction, further than the first cook went. More garlic than the garlic cook made.

That is guidance. Not "follow the prompt". More like "follow the prompt, then keep going".

Why this was needed

Early picture models had a real problem. Ask for "a red bus on a rainy street". You would get a street and rain. Maybe a bus somewhere, in some colour.

The prompt was a suggestion. The model kept drifting toward the average of everything it had ever seen.

The first fix used a second network, a separate classifier that could recognise buses, to nudge the picture. It worked, but it meant training and maintaining another model for every set of labels.

Classifier-free guidance removes that second model. During training, the prompt is dropped at random, maybe one time in ten. The model therefore learns both jobs: making pictures with a prompt, and making pictures without one.

At generation time you run it both ways and take the difference.

noisy picture -> [ model, with your prompt    ] -> guess A
noisy picture -> [ model, with no prompt      ] -> guess B

        the difference (A minus B) is "what the prompt is doing"

  final guess  =  B  +  strength x (A minus B)

Set the strength to one and you get the ordinary prompted result. Set it higher and the prompt's influence is exaggerated.

The dial you have already used

Every image tool exposes this. It is called guidance scale, CFG scale, or prompt strength.

Turn it up and pictures match the prompt more closely, look sharper, and become more predictable.

Turn it up further and they go wrong. Colours burn out. Contrast goes harsh. Everything starts looking the same as everything else.

That last part surprises people. High guidance does not mean high quality. It means low variety.

The trade you are making

Low guidance: varied, surprising, often ignores half your prompt.

High guidance: obedient, sharp, and increasingly all the same picture.

Somewhere in the middle is the setting for your task. There is no universal right answer, and the good range differs between models.

What it costs

Two runs of the model per step instead of one. So guidance roughly doubles your generation time.

That is a real price. It is why fast models work so hard to build the guidance effect into one pass.

Where you have seen this

  • The "guidance scale" slider in every image generator.
  • Burnt, oversaturated AI images, which usually mean the dial was too high.
  • Video and audio generators, which use exactly the same mechanism.

Remember this

  • The model is run twice, with and without your prompt.
  • The difference between the two runs is amplified, not only followed.
  • Higher guidance buys obedience and costs variety.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch==2.5.1

Guidance on a two-dimensional problem you can measure

Real image models make this hard to study, because "did the guidance help" becomes a judgement call. On a toy distribution where the correct answer is known, every effect becomes a number.

Two classes, each a Gaussian blob. Train a conditional denoiser with 10 percent label dropout. Then sample class 0 at several guidance scales. Measure how well the samples match the real class 0 distribution.

cfg.py
import torch, torch.nn as nn, math, time
torch.manual_seed(0)

CENTRES = torch.tensor([[-2.0, 0.0], [2.0, 0.0]])
NULL = 2                                   # the "no condition" token, id 2

def sample_data(n):
    y = torch.randint(0, 2, (n,))
    return CENTRES[y] + 0.5 * torch.randn(n, 2), y

T = 100
betas = torch.linspace(1e-4, 0.06, T)
alphas = 1 - betas
abar = torch.cumprod(alphas, 0)

class Eps(nn.Module):
    def __init__(self, h=128):
        super().__init__()
        self.emb = nn.Embedding(3, h)      # class 0, class 1, and the null token
        self.net = nn.Sequential(nn.Linear(2 + 1 + h, h), nn.SiLU(),
                                 nn.Linear(h, h), nn.SiLU(), nn.Linear(h, 2))
    def forward(self, x, t, y):
        return self.net(torch.cat([x, t[:, None].float() / T, self.emb(y)], -1))

net = Eps(); opt = torch.optim.Adam(net.parameters(), lr=2e-3)
t0 = time.time()
for step in range(3000):
    x0, y = sample_data(256)
    y = torch.where(torch.rand(256) < 0.1, torch.full_like(y, NULL), y)   # 10% label dropout
    t = torch.randint(0, T, (256,))
    noise = torch.randn_like(x0)
    a = abar[t][:, None]
    xt = a.sqrt() * x0 + (1 - a).sqrt() * noise
    loss = ((net(xt, t, y) - noise) ** 2).mean()
    opt.zero_grad(); loss.backward(); opt.step()
print(f"trained in {time.time()-t0:.0f}s, final loss {loss.item():.4f}")

@torch.no_grad()
def sample(n, cls, w):
    x = torch.randn(n, 2)
    y = torch.full((n,), cls); null = torch.full((n,), NULL)
    for i in reversed(range(T)):
        t = torch.full((n,), i)
        e_c, e_u = net(x, t, y), net(x, t, null)
        eps = e_u + w * (e_c - e_u)            # the whole of classifier-free guidance
        mean = (x - betas[i] / (1 - abar[i]).sqrt() * eps) / alphas[i].sqrt()
        x = mean + (betas[i].sqrt() * torch.randn_like(x) if i else 0)
    return x

print("\n  w   on the right class   mean distance from its centre   spread")
for w in (0.0, 1.0, 2.0, 5.0, 10.0):
    s = sample(2000, 0, w)
    right = ((s - CENTRES[0]).norm(dim=1) < (s - CENTRES[1]).norm(dim=1)).float().mean()
    print(f"{w:5.1f} {right:16.1%} {(s - CENTRES[0]).norm(dim=1).mean():25.3f} {s.std(0).mean():13.3f}")
real, _ = sample_data(200000)
r0 = real[:, 0] < 0
print(f"\nreal class-0 data: mean distance {(real[r0]-CENTRES[0]).norm(dim=1).mean():.3f}, "
      f"spread {real[r0].std(0).mean():.3f}")
Output
trained in 9s, final loss 0.2975

  w   on the right class   mean distance from its centre   spread
  0.0            49.2%                     2.395         1.301
  1.0            99.9%                     0.617         0.495
  2.0           100.0%                     0.535         0.424
  5.0           100.0%                     0.531         0.346
 10.0           100.0%                     1.048         1.540

real class-0 data: mean distance 0.628, spread 0.501

Exact figures vary with your PyTorch build. Every trend below reproduces.

Reading the output, row by row

w = 0 is the unconditional model. 49.2% on the right class is a coin flip. The spread of 1.301 reflects samples landing on both blobs. The condition has been switched off entirely, which is what w = 0 means in this formula.

w = 1 is plain conditional sampling, with no guidance at all. Look at how well it matches. Mean distance 0.617 against the real 0.628. Spread 0.495 against the real 0.501. This is the honest baseline, and it is the best distributional match on the whole table.

w = 2 and w = 5 are more obedient and less truthful. Class purity is 100%. But the spread has fallen to 0.424 and then 0.346, below the real 0.501. The samples are being squeezed toward the middle of the class. In image terms: more prototypical, less varied.

w = 10 breaks. Mean distance jumps to 1.048 and spread to 1.540. The samples have been pushed past the data and out the other side. This is the numerical version of the burnt, oversaturated look you get from a guidance scale of 25.

The important honest point: guidance does not improve the model's match to the data distribution. It trades distributional accuracy for conditional obedience. That trade is usually worth making for a picture a human will judge. It is usually wrong when you need samples that represent the data.

What the formula looks like in a real pipeline

The single line that matters:

python
eps = eps_uncond + guidance_scale * (eps_cond - eps_uncond)

In production the two forward passes are batched together:

python
latent_in = torch.cat([latents] * 2)                       # one copy per condition
embeds = torch.cat([negative_embeds, prompt_embeds])       # negative first, by convention
noise_pred = unet(latent_in, t, encoder_hidden_states=embeds).sample
uncond, cond = noise_pred.chunk(2)
noise_pred = uncond + guidance_scale * (cond - uncond)

That is why guidance costs 2x compute rather than 2x wall-clock on a GPU. The batch dimension absorbs it, up to your memory limit.

Negative prompts are the same mechanism

Nothing in the formula requires the second run to be unconditional. Replace the empty prompt with "blurry, watermark, extra fingers" and you get:

python
eps = eps_negative + scale * (eps_positive - eps_negative)

The result is pushed away from the negative prompt and toward the positive one. Negative prompts are not a separate feature. They are the unconditional slot, filled in.

Practical ranges

Guidance scales are not transferable between model families, and quoting a number without the model is meaningless. Published defaults from the model cards and pipeline defaults:

Model familyTypical range
Stable Diffusion 1.5 / 2.17 to 9
SDXL5 to 8
SD 3.54 to 5
FLUX.1 dev (distilled guidance)3 to 4
Turbo / LCM / distilled models0 to 2

Distilled models are the interesting case. They have the guidance effect baked into the weights. Applying guidance on top of that double-counts and produces artefacts.

Common mistakes

Raising guidance to fix a prompt the model does not understand. If the concept is not in the model, amplifying its absence does nothing except burn the colours. Change the prompt or the model.

Assuming higher guidance is more accurate. The table above shows it moving away from the data. It moves toward the prototype of the class, which is a different thing.

Applying guidance to a distilled model. See above.

Comparing FID across guidance scales without saying so. Guidance improves human preference scores and worsens FID, because FID measures distributional match. Both numbers are correct and they disagree, which is exactly what the table predicts.

Forgetting that the negative prompt costs nothing extra. You are already paying for the second forward pass. An empty negative prompt wastes it.

Try it yourself

Change the label dropout from 0.1 to 0.0, retrain, and rerun the guidance sweep. With no dropout the model never learned the unconditional job. eps_uncond is then nonsense, and guidance produces garbage at every scale. That one experiment shows why the dropout is the whole trick.

What to learn next

Researcher — Mathematics and papers.

Classifier guidance, first

Dhariwal and Nichol (2021) introduce guidance by adding the gradient of a separately trained classifier $p_\phi(y \mid \mathbf{x}_t)$ to the score:

$$ \hat{\boldsymbol{\epsilon}}(\mathbf{x}t, t, y) = \boldsymbol{\epsilon}\theta(\mathbf{x}_t, t) - s\sqrt{1 - \bar{\alpha}t}\, \nabla{\mathbf{x}t} \log p\phi(y \mid \mathbf{x}_t) $$

$s$ is the guidance scale. This follows from $\nabla \log p(\mathbf{x}_t \mid y) = \nabla \log p(\mathbf{x}_t) + \nabla \log p(y \mid \mathbf{x}_t)$, with $s$ inserted as a free exponent on the likelihood term. Two costs follow. A classifier must be trained on noisy images at every timestep. The label set is fixed at classifier-training time.

Classifier-free guidance

Ho and Salimans (2022), Classifier-Free Diffusion Guidance, eliminate the classifier. Train a single conditional model $\boldsymbol{\epsilon}_\theta(\mathbf{x}t, t, \mathbf{c})$. The condition $\mathbf{c}$ is replaced with a learned null token $\varnothing$ with probability $p{\text{uncond}}$, typically 0.1 to 0.2. At sampling:

$$ \tilde{\boldsymbol{\epsilon}}_\theta(\mathbf{x}t, t, \mathbf{c}) = \boldsymbol{\epsilon}\theta(\mathbf{x}t, t, \varnothing) + w\left(\boldsymbol{\epsilon}\theta(\mathbf{x}t, t, \mathbf{c}) - \boldsymbol{\epsilon}\theta(\mathbf{x}_t, t, \varnothing)\right) $$

Two conventions for $w$ circulate and confusing them is a common source of bugs. In the form above (used in Stable Diffusion), $w = 1$ is plain conditional sampling and $w = 0$ is unconditional. The original paper writes $(1+w)\boldsymbol{\epsilon}c - w\boldsymbol{\epsilon}\varnothing$, in which $w = 0$ is plain conditional. Always check which one a codebase means.

Why it samples from a sharpened distribution

The implied score is:

$$ \nabla_{\mathbf{x}_t} \log \tilde{p}(\mathbf{x}t \mid \mathbf{c}) = \nabla{\mathbf{x}_t} \log p(\mathbf{x}t) + w \nabla{\mathbf{x}_t} \log p(\mathbf{c} \mid \mathbf{x}_t) $$

which corresponds to sampling from $p(\mathbf{x}_t) \, p(\mathbf{c} \mid \mathbf{x}_t)^{w}$. Raising the likelihood to a power greater than one concentrates mass where the condition is most strongly satisfied. It is a temperature on the conditional term. The resulting distribution is not the true conditional for any $w \ne 1$.

Karras et al. (2024), Guiding a Diffusion Model with a Bad Version of Itself, sharpen this analysis. Much of what CFG does at high scale is truncation toward high-density regions. It conflates two effects: improving image quality, and reducing variation. They propose autoguidance. It replaces the unconditional model with a smaller or less-trained version of the same conditional model. Quality improves without the same diversity collapse. This is the cleanest current account of what CFG is doing and what it costs.

The measured trade-off

The finding that reproduces everywhere: FID is U-shaped in $w$, Inception Score and human preference are monotonic. Guidance improves per-sample fidelity and prompt adherence while shrinking coverage, and FID is sensitive to both.

Reporting practice follows from this. A guidance sweep with FID and CLIP-score plotted against each other, the standard Pareto curve, is the honest presentation. A single FID at an unstated guidance scale is not comparable to anything.

Known failure modes and their fixes

Oversaturation and burnt contrast at high $w$. Caused by the guided prediction leaving the valid range. Lin et al. (2023) prescribe rescaling the guided output, toward the standard deviation of the conditional prediction. Their paper is Common Diffusion Noise Schedules and Sample Steps are Flawed. diffusers exposes the fix as guidance_rescale, and 0.7 is a common value.

Zero terminal SNR. The same paper shows most schedules do not reach pure noise at the final step. The model always sees a residual mean brightness. It therefore cannot produce very dark or very bright images. Fixes are rescale_betas_zero_snr=True with timestep_spacing="trailing", and a model trained with $\mathbf{v}$-prediction.

Doubled compute. Addressed by distillation. Meng et al. (2023), On Distillation of Guided Diffusion Models, train a student to match the guided output in one pass. FLUX.1 dev ships a guidance-distilled model, which is why applying external CFG to it is a mistake.

Interval guidance. Kynkäänniemi et al. (2024) show guidance is harmful at very high and very low noise levels. It is beneficial only in the middle. Restricting CFG to a timestep interval improves FID substantially, at the same prompt adherence.

Papers

What to learn next