Image Generation and Restoration
ControlNet
ControlNet bolts a trainable copy of the image model onto the frozen original, joined by layers initialised to zero so it changes nothing until it has learned something.
- 14 min read
- 3 reading levels
- Updated
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
ControlNet lets you hand the model a rough drawing and say "the picture must follow these lines".
The analogy
Think about tracing paper laid over a drawing. The lines underneath show through. Whatever you paint on top, the shape stays where the original artist put it.
Words alone cannot do this. Describe a person's exact pose in a sentence and try it. Arm slightly bent, head turned three quarters, weight on the left leg. You will not get what you pictured.
A stick figure takes two seconds and says it perfectly.
What it takes as input
ControlNet accepts a picture that describes structure rather than appearance:
- An edge map, which is the outlines of a photo with everything else stripped away.
- A pose skeleton, which is a stick figure of where the joints are.
- A depth map, a grey picture where near things are bright and far things are dark.
- A scribble, which is exactly what it sounds like.
- A segmentation map, where each region is coloured by what it contains.
your prompt: "a woman in a red saree, evening light"
+
your control image: a stick figure in a specific pose
|
v
[ image model + ControlNet ]
|
v
a woman in a red saree, in exactly that poseThe clever part
Attaching new machinery to a trained model usually damages it. The new pieces start out random, and random noise flowing into a careful system wrecks it.
ControlNet gets around this in a way worth admiring.
The connection between the new branch and the original model runs through layers whose weights start at exactly zero. Zero multiplied by anything is zero. So on the very first training step, the new branch contributes nothing at all.
The model behaves exactly as it did before. Then, over training, those zero weights slowly become non-zero, and the control branch fades in.
Nothing is ever broken, because nothing is ever suddenly changed.
Why the original stays frozen
The main image model is not trained at all during this. Its weights are locked.
Two reasons. It keeps everything the big model already knows. And you can train a ControlNet on an ordinary graphics card, not a data centre.
It also means ControlNets are swappable. One edge ControlNet, one pose ControlNet, one depth ControlNet. All sit beside the same frozen base model, loaded as needed.
The dial that matters
There is a strength setting. Low, and the control is a hint the model may ignore. High, and the control is a rule the model must obey, even when obeying makes the picture worse.
Both extremes are wrong for most work. Middle settings usually win.
The honest catch
ControlNet is faithful, and that includes being faithful to your mistakes.
Suppose your edge detector picked up noise, grain or a compression artefact. ControlNet will draw those lines carefully into the final picture. Preparing a clean control image matters as much as the prompt does. You will see that measured next.
Where you have seen this
- Interior design tools that keep the room layout and change the furniture.
- Character tools that hold a pose fixed and change the outfit.
- Product photo tools that keep the product shape and change the background.
- Colouring and line-art tools built on a sketch.
Remember this
- ControlNet adds structure control that words cannot express.
- The joining layers start at zero, so the base model is never disturbed.
- The original model stays frozen, which makes ControlNets small and swappable.
What to learn next
- Inpainting and outpainting — the other way to constrain a generated image.
- LoRA — the other way to extend a frozen model cheaply.
- OpenCV — the toolkit that produces most control images.
Developer — Code and libraries.
Setup
pip install torch==2.5.1 opencv-python==4.10.0The zero convolution, verified rather than described
The claim has two halves. An untrained ControlNet changes the base model's output by exactly nothing. It still receives gradients, so it can learn. Both halves are checkable in a few lines.
import torch, torch.nn as nn, copy
torch.manual_seed(0)
def trunk_block(): return nn.Sequential(nn.Conv2d(4, 16, 3, padding=1), nn.SiLU(),
nn.Conv2d(16, 4, 3, padding=1))
class ControlNet(nn.Module):
def __init__(self, frozen):
super().__init__()
self.frozen = frozen
for p in self.frozen.parameters():
p.requires_grad_(False) # the base model never moves
self.copy = copy.deepcopy(frozen) # trainable copy of the same block
for p in self.copy.parameters():
p.requires_grad_(True)
self.hint = nn.Conv2d(1, 4, 3, padding=1) # turns the control image into features
self.zero = nn.Conv2d(4, 4, 1) # the "zero convolution"
nn.init.zeros_(self.zero.weight); nn.init.zeros_(self.zero.bias)
def forward(self, x, control):
return self.frozen(x) + self.zero(self.copy(x + self.hint(control)))
base = trunk_block()
cn = ControlNet(base)
x, c = torch.randn(2, 4, 8, 8), torch.randn(2, 1, 8, 8)
with torch.no_grad():
diff = (cn(x, c) - base(x)).abs().max().item()
print(f"step 0, largest difference from the frozen model: {diff:.1e}")
print("-> attaching an untrained ControlNet changes NOTHING. That is the whole trick.\n")
loss = ((cn(x, c) - torch.randn_like(x)) ** 2).mean()
loss.backward()
print(f"gradient norm on the zero conv: {cn.zero.weight.grad.norm():.4f} (non-zero, so it can learn)")
print(f"gradient norm on the frozen base: "
f"{sum(p.grad is not None for p in cn.frozen.parameters())} parameters received a gradient")
target = torch.randn(2, 4, 8, 8)
opt = torch.optim.Adam([p for p in cn.parameters() if p.requires_grad], lr=1e-2)
print("\nstep loss contribution of the control branch")
for s in range(201):
out = cn(x, c)
l = ((out - target) ** 2).mean()
opt.zero_grad(); l.backward(); opt.step()
if s % 50 == 0:
with torch.no_grad():
share = (cn(x, c) - base(x)).abs().mean() / cn(x, c).abs().mean()
print(f"{s:4d} {l.item():.4f} {share:26.1%}")step 0, largest difference from the frozen model: 0.0e+00 -> attaching an untrained ControlNet changes NOTHING. That is the whole trick. gradient norm on the zero conv: 0.0436 (non-zero, so it can learn) gradient norm on the frozen base: 0 parameters received a gradient step loss contribution of the control branch 0 0.9353 7.8% 50 0.0331 101.1% 100 0.0002 101.4% 150 0.0000 101.4% 200 0.0000 101.4%
Reading the output
0.0e+00, not "small". The difference is exactly zero, in floating point. The zero convolution multiplies its input by a matrix of zeros. Fine-tuning approaches that add randomly initialised layers cannot claim this. That is why they degrade the base model early in training.
The gradient on the zero conv is non-zero. This is the part that catches people out. If the weights are zero, how does anything learn? The gradient of a convolution with respect to its weights is the input, not the weights. So a zero-weight layer receives a real gradient on step one and starts moving immediately. Only its output is zero, not its gradient.
Zero parameters in the frozen trunk received a gradient. requires_grad_(False) did its job. That is the memory saving: no optimiser state and no stored activations for the base model.
The control branch fades in. From 7.8% of the output at step 0 to over 100% by step 50. It ends above 100 percent for a reason. The branch has learned to cancel part of the frozen output as well as add to it.
Preparing the control image is half the job
ControlNet copies your control image faithfully, including the parts you did not mean.
import cv2, numpy as np
print("opencv", cv2.__version__)
# A synthetic "building": two rectangles, a triangular roof, and mild sensor noise.
img = np.full((128, 128), 40, np.uint8)
cv2.rectangle(img, (20, 50), (60, 110), 200, -1)
cv2.rectangle(img, (70, 30), (108, 110), 140, -1)
cv2.fillPoly(img, [np.array([[70, 30], [89, 12], [108, 30]])], 90)
rng = np.random.default_rng(0)
img = np.clip(img.astype(np.int16) + rng.normal(0, 18, img.shape), 0, 255).astype(np.uint8)
print("\nlow high edge pixels share of the image")
for lo, hi in [(50, 150), (100, 200), (150, 250), (200, 250)]:
e = cv2.Canny(img, lo, hi)
print(f"{lo:4d} {hi:4d} {int((e > 0).sum()):13d} {(e > 0).mean():18.2%}")
blur = cv2.GaussianBlur(img, (5, 5), 0)
e_raw, e_blur = cv2.Canny(img, 100, 200), cv2.Canny(blur, 100, 200)
print(f"\nwithout blurring first: {int((e_raw > 0).sum()):4d} edge pixels")
print(f"with a 5x5 blur first : {int((e_blur > 0).sum()):4d} edge pixels")
print("the extra edges in the first case are noise, and a ControlNet will faithfully draw them")opencv 4.10.0 low high edge pixels share of the image 50 150 5519 33.69% 100 200 3326 20.30% 150 250 935 5.71% 200 250 668 4.08% without blurring first: 3326 edge pixels with a 5x5 blur first : 503 edge pixels the extra edges in the first case are noise, and a ControlNet will faithfully draw them
A third of the image became "edges" at the loosest threshold. After a five-by-five blur, the same thresholds give 503 edge pixels, close to the true structure. The other 2823 were grain.
If your ControlNet output looks cluttered and scratchy, this is almost always why. Blur before you detect, and look at the control image before you generate.
Using it in diffusers
# written against diffusers 0.40.0
import torch
from diffusers import FluxControlNetPipeline, FluxControlNetModel
controlnet = FluxControlNetModel.from_pretrained(
"InstantX/FLUX.1-dev-Controlnet-Canny", dtype=torch.bfloat16
)
pipeline = FluxControlNetPipeline.from_pretrained(
"black-forest-labs/FLUX.1-dev", controlnet=controlnet, dtype=torch.bfloat16
).to("cuda")
image = pipeline(
"a photorealistic overhead image of a cat in a pool floatie",
control_image=canny_image,
controlnet_conditioning_scale=0.5,
num_inference_steps=50,
guidance_scale=3.5,
).images[0]No output block, deliberately: this needs a GPU and a multi-gigabyte download. Printing an invented image description or timing here would teach you to expect something that may not happen.
The parameter names are the load-bearing part. controlnet_conditioning_scale sets the strength. control_image is the structure image, distinct from image. In the img2img and inpaint pipelines, image is the picture being edited. Passing a control image as image is a common and confusing mistake.
For multiple controls, pass a list of ControlNet models and a matching list of scales:
pipeline(prompt, image=[canny_image, depth_image],
controlnet_conditioning_scale=[0.5, 0.5])Common mistakes
Conditioning scale at 1.0 by default. The diffusers examples use 0.5 for a reason. At 1.0 the control dominates and the prompt has little room left.
Mismatched ControlNet and base model. A ControlNet trained for SD 1.5 will not load against SDXL. An SDXL one will not load against FLUX. The block structure differs.
Feeding a raw photo as an edge control. The preprocessor is not optional. An edge ControlNet expects an edge map, not a photograph.
Forgetting to resize the control image. It must match the generation resolution. Silent aspect-ratio distortion is a frequent cause of "the pose is slightly off".
Stacking four ControlNets at full strength. They fight, and the result satisfies none of them. Mask them so they cover different regions, and lower each scale.
Try it yourself
Take the canny script and add a cv2.medianBlur variant alongside the Gaussian. Compare edge counts. Median blur removes salt-and-pepper noise while keeping true edges sharp. On grainy photographs it often beats Gaussian for this job.
What to learn next
- Inpainting and outpainting — the other way to constrain a generated image.
- LoRA — the other way to extend a frozen model cheaply.
- OpenCV — the toolkit that produces most control images.
Researcher — Mathematics and papers.
Architecture
Zhang, Rao and Agrawala (2023) published Adding Conditional Control to Text-to-Image Diffusion Models at ICCV 2023. They define ControlNet for a neural block $\mathcal{F}(\cdot; \Theta)$ with frozen parameters $\Theta$:
$$ \mathbf{y}c = \mathcal{F}(\mathbf{x}; \Theta) + \mathcal{Z}\big(\mathcal{F}(\mathbf{x} + \mathcal{Z}(\mathbf{c}; \Theta{z1}); \Theta_c); \Theta_{z2}\big) $$
$\mathbf{c}$ is the conditioning image, and $\Theta_c$ a trainable copy of $\Theta$. $\mathcal{Z}(\cdot; \Theta_{z})$ is a $1 \times 1$ convolution with weights and bias initialised to zero. At initialisation both zero convolutions output zero, so $\mathbf{y}_c = \mathcal{F}(\mathbf{x}; \Theta)$ exactly.
The gradient argument is the part worth stating carefully. For $y = wx + b$ with $w = b = 0$:
$$ \frac{\partial y}{\partial w} = x, \qquad \frac{\partial y}{\partial x} = w = 0, \qquad \frac{\partial y}{\partial b} = 1 $$
So $\partial y / \partial w \ne 0$ provided $x \ne 0$, and $w$ moves off zero on the first step. Once $w \ne 0$, $\partial y / \partial x$ becomes non-zero and the preceding layers begin to train. The paper calls the resulting behaviour "sudden convergence". The model ignores the control for thousands of steps, then follows it abruptly.
In Stable Diffusion, the trainable copy covers the 12 encoder blocks and the middle block. The decoder is left to the frozen model, with ControlNet outputs added to the skip connections. For SD 1.5 this is roughly 360 million additional parameters against 860 million frozen.
Training
Text prompts are replaced with the empty string 50 percent of the time during training. This forces the model to extract semantics from the control image rather than leaning on the caption. It is what makes guess_mode work: generation from the control alone, with no prompt.
The published models were trained on 50 thousand to 3 million condition-image pairs. Those pairs came from running an off-the-shelf detector over an image corpus. Small conditions trained on a single consumer GPU, which was the practical breakthrough.
The variant family
| Method | Year | Trade-off |
|---|---|---|
| ControlNet | 2023 | Full trainable encoder copy; highest fidelity, largest |
| T2I-Adapter | 2023 | Small feature-extraction network, around 77M parameters; lighter, less precise |
| ControlNet-LITE / small | 2023 | Reduced-channel copies; a fraction of the size |
| ControlNet-XS | 2023 | Argues the copy need not mirror the base at all; roughly 1 percent of the parameters |
| Uni-ControlNet / UniControl | 2023 | One adapter across many condition types |
| ControlNet++ | 2024 | Adds a cycle-consistency reward: run a detector on the output and match it back to the input control |
ControlNet++ (Li et al., 2024) is the most useful correction to the original. It observes that ControlNet is trained with a diffusion loss only. Nothing rewards the output for actually matching the control. Their reward re-detects the condition from the generated image and compares it to the input. Measured controllability improves substantially.
The lesson generalises past ControlNet. If you have a detector for the property you are conditioning on, use it as an evaluation metric. Consider it as a training signal too.
Evaluation
Controllability is measurable, and should be measured rather than eyeballed:
- Condition reconstruction. Run the same preprocessor on the generated image and compare against the input control. F1 for edges, mAP or PCK for pose, RMSE for depth, mIoU for segmentation.
- Prompt adherence. CLIP score against the text prompt.
- Image quality. FID against the target distribution.
These three trade against one another as controlnet_conditioning_scale rises: condition reconstruction improves, prompt adherence and FID degrade. Report the curve, not a point.
Beyond the U-Net
ControlNet was designed against the U-Net block structure. Transformer backbones (SD 3, FLUX) needed a redesign, since there is no encoder-decoder skip structure to inject into. Current FLUX and SD 3 ControlNets copy a subset of the transformer blocks. Their outputs are added into the residual stream, through the same zero-initialised connection.
The zero-initialisation idea has outlived the architecture it was invented for. It now appears wherever a trained model must be extended without disturbance. Zero-initialised gating also appears in LLaMA-Adapter and in several parameter-efficient fine-tuning methods, for the same reason.
Papers
- Zhang et al., Adding Conditional Control to Text-to-Image Diffusion Models, 2023 — arxiv.org/abs/2302.05543
- Mou et al., T2I-Adapter, 2023 — arxiv.org/abs/2302.08453
- Zavadski et al., ControlNet-XS, 2023 — arxiv.org/abs/2312.06573
- Zhao et al., Uni-ControlNet, 2023 — arxiv.org/abs/2305.16322
- Li et al., ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback, 2024 — arxiv.org/abs/2404.07987
What to learn next
- Inpainting and outpainting — the other way to constrain a generated image.
- LoRA — the other way to extend a frozen model cheaply.
- OpenCV — the toolkit that produces most control images.