Image Generation and Restoration

StyleGAN and latent space

StyleGAN feeds the random input in as a style at every layer, which turns the space of random numbers into a map you can steer.

On this page 9
  1. The short answer
  2. The analogy
  3. What StyleGAN changed
  4. Latent space, without the jargon
  5. The catch nobody mentions first
  6. The truncation dial
  7. Where you have seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

StyleGAN builds a picture in stages, and lets you set the style of each stage separately.

The analogy

Think about ordering a shirt from a tailor. You do not hand over one number and hope. You settle the cut, then the fabric, then the collar, then the buttons.

Each decision is separate. You can change the buttons without redoing the cut.

Older image generators worked the other way. One list of random numbers went in at the bottom, and everything that came out was tangled together. Changing one thing changed all of it.

What StyleGAN changed

StyleGAN feeds the random numbers in again at every layer. They arrive as a style setting, not as a starting picture.

The early layers control coarse things: pose, face shape, hair length. The middle layers control features like eyes and nose. The last layers control colour and fine texture.

random numbers
      |
   [ mapping network ]  -> a "styles" vector
      |
      +---> layer 1 (coarse: pose, shape)
      +---> layer 2 (middle: eyes, nose)
      +---> layer 3 (fine: colour, skin texture)
                                |
                                v
                            the picture

Because each layer gets its own style, you can mix. Take the coarse styles from one face and the fine styles from another. You get the first person's pose with the second person's colouring.

Latent space, without the jargon

Latent space is the space of the random numbers that go in. Every point in it produces one picture.

Here is the useful part. Nearby points give similar pictures. Walk in a straight line from one point to another and the picture changes smoothly.

So directions in this space mean something. There is a direction that makes faces older. There is a direction that adds a smile. Finding those directions is how face editing apps work.

The catch nobody mentions first

Those directions are not clean. In real photos, smiling faces and open mouths go together. So the "smile" direction also opens mouths.

Worse, the training photos carry human bias. If most of the smiling photos were young people, the smile direction will also make faces younger.

This is not a bug in the code. It is the data showing through. That is worth sitting with.

The truncation dial

There is one more control worth knowing. You can pull every random point toward the average point.

Pull hard and every face looks safe, symmetric and a bit boring. Pull less and you get more variety, along with more broken results.

Every image tool has this dial under some name. It is a straight trade between quality and variety.

Where you have seen this

  • Photo apps that age a face, add a smile or change hairstyle.
  • "This person does not exist" style websites.
  • Avatar generators that let you slide between looks.

Remember this

  • StyleGAN injects style at every layer, so coarse and fine detail become separate controls.
  • Latent space is a map, and straight-line directions in it correspond to real changes.
  • Those directions come tangled, because the training photos were tangled.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy==1.26.4 scikit-learn==1.7.2 torch==2.5.1

The real StyleGAN checkpoints are large and need a GPU to sample comfortably. Every mechanism on this page can be studied without them. Use a synthetic latent space where you know the ground truth. That is the only way to check whether an editing method really works. With real faces you cannot see the true attribute directions.

Finding and using a latent direction

This is InterFaceGAN (Shen et al., 2020) end to end. Label an attribute, fit a linear boundary, then edit along its normal. Finally, remove the entanglement by projection.

latent_directions.py
import numpy as np
from sklearn.linear_model import LogisticRegression
rng = np.random.default_rng(0)
D = 16                                    # a toy stand-in for StyleGAN's 512-d W space

# Two attributes are baked into this fake W space, and they are correlated,
# exactly as "smiling" and "age" are correlated in the FFHQ face dataset.
W = rng.normal(size=(4000, D))
n_smile = np.zeros(D); n_smile[0] = 1.0                 # ground-truth smile direction
n_age   = np.zeros(D); n_age[0] = 0.6; n_age[1] = 0.8   # age shares component 0 with smile
smiling = (W @ n_smile > 0).astype(int)
old     = (W @ n_age   > 0).astype(int)
print(f"correlation between the two attributes in the data: {np.corrcoef(smiling, old)[0,1]:+.3f}")

# InterFaceGAN: fit a linear boundary to a labelled attribute, edit along its normal.
clf = LogisticRegression(max_iter=2000).fit(W, smiling)
n = clf.coef_[0] / np.linalg.norm(clf.coef_[0])
print(f"recovered smile normal, cosine to the true one: {n @ n_smile:.3f}")

def edit(w, direction, alpha):
    return w + alpha * direction

sample = W[:1000]
for alpha in (0.0, 1.0, 2.0, 3.0):
    e = edit(sample, n, alpha)
    print(f"alpha {alpha:3.1f}  smiling {np.mean(e @ n_smile > 0):6.1%}   "
          f"also became old {np.mean(e @ n_age > 0):6.1%}")

# Conditional manipulation: remove the component of the smile direction that
# points along the age direction, so editing one leaves the other alone.
a = n_age / np.linalg.norm(n_age)
n_cond = n - (n @ a) * a
n_cond /= np.linalg.norm(n_cond)
print("\nafter projecting the age component out of the smile direction:")
for alpha in (0.0, 1.0, 2.0, 3.0):
    e = edit(sample, n_cond, alpha)
    print(f"alpha {alpha:3.1f}  smiling {np.mean(e @ n_smile > 0):6.1%}   "
          f"also became old {np.mean(e @ n_age > 0):6.1%}")

# The truncation trick: pull every latent toward the average one.
w_avg = W.mean(0)
print("\npsi   spread of W   mean distance from the average face")
for psi in (1.0, 0.7, 0.5, 0.0):
    t = w_avg + psi * (W - w_avg)
    print(f"{psi:4.1f} {t.std():13.3f} {np.linalg.norm(t - w_avg, axis=1).mean():29.3f}")
Output
correlation between the two attributes in the data: +0.410
recovered smile normal, cosine to the true one: 1.000
alpha 0.0  smiling  49.0%   also became old  48.9%
alpha 1.0  smiling  84.2%   also became old  72.2%
alpha 2.0  smiling  97.0%   also became old  89.0%
alpha 3.0  smiling  99.8%   also became old  96.8%

after projecting the age component out of the smile direction:
alpha 0.0  smiling  49.0%   also became old  48.9%
alpha 1.0  smiling  79.0%   also became old  48.9%
alpha 2.0  smiling  94.0%   also became old  48.9%
alpha 3.0  smiling  98.7%   also became old  48.9%

psi   spread of W   mean distance from the average face
 1.0         1.000                         3.936
 0.7         0.700                         2.755
 0.5         0.500                         1.968
 0.0         0.020                         0.000

Reading the output

The classifier recovered the true direction perfectly (cosine 1.000). That is expected here, because the attribute really is a linear function of the latent. On real StyleGAN latents it will not be perfect, and the cosine is not measurable. That is why a synthetic check like this one is worth running first.

Naive editing drags a second attribute along. At alpha 3.0 the smile is at 99.8% and the unrelated attribute has been pushed from 48.9% to 96.8%. Nobody asked for that. A user sliding "smile" sees the face age, and calls your product broken.

The projection fixes it exactly. After removing the age component, the second attribute sits at 48.9% at every step. That is unchanged from where it started, while smiling still climbs to 98.7%. Two lines of linear algebra.

Truncation is a linear interpolation toward the mean. At psi = 0.5 the spread halves. At psi = 0 every latent is the average and you get one face. StyleGAN applies this in W, not in Z, and that matters. W is the space where attributes are closer to linearly separable.

Modulated convolution, the StyleGAN2 mechanism

StyleGAN1 injected style with AdaIN, normalising activations then rescaling them per channel. That produced the "water droplet" artefacts visible in StyleGAN1 samples. StyleGAN2 moved the same operation onto the weights.

modulated_conv.py
import torch
torch.manual_seed(0)

C_in, C_out, k = 8, 8, 3
w = torch.randn(C_out, C_in, k, k) / (C_in * k * k) ** 0.5   # unit-variance init
x = torch.randn(1, C_in, 16, 16)
style = torch.tensor([4.0, 0.25] + [1.0] * 6)                # one loud channel, one quiet

def conv(weight, inp):
    return torch.nn.functional.conv2d(inp, weight, padding=1)

print(f"input std                     {x.std():.3f}")
print(f"plain conv output std         {conv(w, x).std():.3f}")

# StyleGAN2 modulation: scale each INPUT channel of the weight by its style value.
w_mod = w * style.view(1, C_in, 1, 1)
print(f"after modulation only         {conv(w_mod, x).std():.3f}   <- style changed the volume")

# Demodulation: divide each output filter by the L2 norm of its modulated weights,
# which restores unit output variance without touching the activations.
sigma = w_mod.pow(2).sum(dim=(1, 2, 3), keepdim=True).add(1e-8).sqrt()
w_demod = w_mod / sigma
print(f"after modulate + demodulate   {conv(w_demod, x).std():.3f}   <- volume restored")

print("\nper-output-filter std, first four filters")
for name, weight in (("modulated  ", w_mod), ("demodulated", w_demod)):
    o = conv(weight, x)[0]
    print(f"  {name} {[round(float(o[c].std()), 3) for c in range(4)]}")
Output
input std                     1.000
plain conv output std         0.983
after modulation only         1.548   <- style changed the volume
after modulate + demodulate   0.990   <- volume restored

per-output-filter std, first four filters
  modulated   [1.673, 1.505, 0.929, 1.304]
  demodulated [0.991, 1.025, 0.977, 0.996]

Modulation alone lets one loud style channel inflate the whole activation. Demodulation divides it back out analytically, using the weights rather than the observed activations. That is the key difference from AdaIN. It never looks at the actual feature statistics, so the network cannot hide a signal in them.

Common mistakes

Editing in Z instead of W. Z is the raw Gaussian input; W is the output of the mapping network. Directions in Z are far more entangled. Every serious editing method operates in W or W+.

Assuming W+ and W are interchangeable. W+ gives each layer its own style vector, which makes inversion of a real photo much more accurate. It also leaves the space the generator was trained on, so edits generalise worse. Pick deliberately.

Turning truncation up to hide a weak model. Low psi makes samples look better and your FID worse, because diversity collapses. Report the psi you used or the number means nothing.

Reusing a direction found on one checkpoint. Latent directions are properties of a specific trained generator. Fine-tune the model and every saved direction is stale.

Try it yourself

Make the two attributes fully independent by setting n_age = [0, 1, 0, ...]. Rerun and confirm that naive editing now leaves the second attribute alone. That tells you the entanglement lives in the data, not in the method.

What to learn next

Researcher — Mathematics and papers.

Architecture

Karras et al. (2019), A Style-Based Generator Architecture for GANs (StyleGAN), restructure the generator into three parts.

Mapping network. An eight-layer MLP $f: \mathcal{Z} \to \mathcal{W}$, both $\mathbb{R}^{512}$. Its purpose is disentanglement. $\mathcal{Z}$ is constrained to a fixed Gaussian, so the generator must warp it to match the data. A learned $\mathcal{W}$ need not be Gaussian and can be less curved.

Synthesis network. Starts from a learned constant $4 \times 4 \times 512$ tensor. Style vectors enter at every resolution through learned affine maps $\mathbf{y} = A(\mathbf{w})$.

Noise inputs. Per-pixel Gaussian noise added after each convolution, scaled by learned per-channel weights. It supplies stochastic detail (hair strands, freckles) that the latent should not have to encode.

The StyleGAN1 style operation is AdaIN:

$$ \mathrm{AdaIN}(\mathbf{x}i, \mathbf{y}) = \mathbf{y}{s,i} \frac{\mathbf{x}_i - \mu(\mathbf{x}_i)}{\sigma(\mathbf{x}i)} + \mathbf{y}{b,i} $$

$\mathbf{x}i$ is feature map $i$, with $\mu$ and $\sigma$ its spatial mean and standard deviation. $\mathbf{y}{s,i}$ and $\mathbf{y}_{b,i}$ are the style scale and bias for that channel.

Weight demodulation

Karras et al. (2020) published Analyzing and Improving the Image Quality of StyleGAN, known as StyleGAN2. They trace the droplet artefacts to AdaIN destroying information carried in feature magnitudes. Their replacement modulates the weights:

$$ w'_{ijk} = s_i \cdot w_{ijk}, \qquad w''{ijk} = \frac{w'{ijk}}{\sqrt{\sum_{i,k} {w'_{ijk}}^2 + \epsilon}} $$

$i$ indexes input channels, $j$ output channels, and $k$ spatial positions. $s_i$ is the style scale for input channel $i$. The normalisation assumes unit-variance i.i.d. inputs, which makes it a statistical assumption rather than an exact operation on the actual activations. That is what removes the artefact. The network can no longer smuggle a signal into feature magnitudes for AdaIN to amplify.

The same paper adds path length regularisation. A fixed-size step in $\mathcal{W}$ should produce a fixed-size change in the image:

$$ \mathbb{E}{\mathbf{w}, \mathbf{y} \sim \mathcal{N}(0,I)}\left(|\mathbf{J}{\mathbf{w}}^{\top} \mathbf{y}|_2 - a\right)^2 $$

$\mathbf{J}_{\mathbf{w}} = \partial g(\mathbf{w}) / \partial \mathbf{w}$ is the generator Jacobian and $a$ a running average of the observed lengths. This makes the latent space better conditioned and measurably easier to invert.

Aliasing and StyleGAN3

Karras et al. (2021), Alias-Free Generative Adversarial Networks (StyleGAN3), identify "texture sticking". Fine detail in StyleGAN2 adheres to pixel coordinates rather than to the object. A translating face therefore drags a static texture. The cause is aliasing from pointwise nonlinearities and naive upsampling, which violate the sampling theorem.

Their fix treats every feature map as a continuous signal. It band-limits before each nonlinearity, and uses ideal low-pass filters for resampling. The result is exact translation and rotation equivariance, at higher compute cost. StyleGAN3 is the correct choice for video and animation, and StyleGAN2 usually remains better on still-image FID.

Latent space geometry

Perceptual path length measures how much the image changes for a small latent step:

$$ \mathrm{PPL} = \mathbb{E}\left[\frac{1}{\epsilon^2} \, d\big(G(\mathrm{lerp}(z_1, z_2; t)),\, G(\mathrm{lerp}(z_1, z_2; t + \epsilon))\big)\right] $$

$d$ is a perceptual distance (LPIPS), $t \sim U(0,1)$ and $\epsilon = 10^{-4}$. Lower PPL indicates a smoother, better-behaved latent space. It is consistently lower in $\mathcal{W}$ than in $\mathcal{Z}$, which is the quantitative justification for the mapping network.

Linear separability measures disentanglement. Fit a linear SVM per attribute in each space. Then compute the conditional entropy of the true label, given the classifier's side of the hyperplane. Again $\mathcal{W}$ wins.

Editing and inversion

InterFaceGAN (Shen et al., 2020) treats each binary attribute as a hyperplane in $\mathcal{W}$. Its unit normal is $\mathbf{n}$. Editing is $\mathbf{w}' = \mathbf{w} + \alpha \mathbf{n}$. For two entangled attributes with normals $\mathbf{n}_1, \mathbf{n}_2$, conditional manipulation projects:

$$ \mathbf{n}_1^{\perp} = \frac{\mathbf{n}_1 - (\mathbf{n}_1^{\top} \mathbf{n}_2)\mathbf{n}_2}{|\mathbf{n}_1 - (\mathbf{n}_1^{\top} \mathbf{n}_2)\mathbf{n}_2|} $$

GANSpace (Härkönen et al., 2020) finds directions unsupervised, by PCA on sampled $\mathbf{w}$ vectors. SeFa (Shen and Zhou, 2021) factorises the first affine layer's weights directly, needing no sampling at all. StyleCLIP (Patashnik et al., 2021) drives the search with a CLIP text encoder. The attribute is named rather than labelled.

Inversion, mapping a real photograph to a latent, is the practical bottleneck. Optimising directly in $\mathcal{W}$ reconstructs poorly. Optimising in $\mathcal{W}+$ reconstructs well, but lands off the learned manifold, so subsequent edits degrade. $\mathcal{W}+$ is a separate $\mathbf{w}$ per layer, $18 \times 512$ for a $1024^2$ generator. Tov et al. (2021), Designing an Encoder for StyleGAN Image Manipulation, formalise this as a distortion-editability trade-off. They train an encoder that targets a chosen point on it.

Honest limits

The attribute directions inherit every correlation in the training set. On FFHQ, "smile" correlates with age and gender presentation, and edits along a naive smile direction shift both. Projection removes only the correlations you thought to name.

That makes latent editing a fairness problem rather than a purely technical one. Karras et al. (2020), Training Generative Adversarial Networks with Limited Data, is also relevant here. Adaptive discriminator augmentation makes StyleGAN trainable on a few thousand images. That puts these systems within reach of small, unaudited datasets.

Papers

What to learn next