Deep Learning

GANs

A GAN trains two networks against each other, one inventing fakes and one catching them, so the invented data gets better without anyone ever writing down what good looks like.

On this page 8
  1. The short answer
  2. The forged signature
  3. Why it exists
  4. How it works
  5. Where you have already seen this
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A GAN is two networks trained against each other. One invents fake data, the other tries to catch it, and the fight makes the fakes convincing.

GAN stands for generative adversarial network — generative meaning it makes new things, adversarial meaning two parts competing.

The forged signature

A child needs a parent's signature on a school note and does not have one. The first attempt is a scribble, and the class teacher spots it in a second.

The child watches what gave it away. Too shaky. Wrong slant. Next week's attempt is closer. The teacher, now suspicious, starts checking the loop on the last letter.

Both of them get better, and neither of them was ever handed a rulebook. The child learns from getting caught. The teacher learns from what slipped through.

A GAN is that, running for a few hundred thousand rounds. The forger is one network. The teacher is another.

Why it exists

Judging is easy to teach a computer. Making is not.

Showing a network a million photos and asking "is this a real photograph" is ordinary classification. You have examples of both answers, and you can score every guess.

Now ask a network to produce a photograph of a face. What score do you give it? You cannot write down a number for "how face-like is this". Nobody can. Marking it against one particular real photo is worse: a perfectly good face that is not that face would score terribly.

This was the wall. Generating things needs a judge, and no one could write the judge down.

The idea behind a GAN is to stop trying. Train the judge from data at the same time as the maker, and let the maker's marking scheme improve alongside it.

How it works

   random numbers
        │
        ▼
   [ GENERATOR ]  ──► an invented sample ──┐
   the forger                              │
                                           ▼
                                    [ DISCRIMINATOR ] ──► "real" or "fake"
                                       the teacher              │
   real training data ─────────────────────▲                    │
                                                                │
        ▲                                                       │
        └───────── both learn from this verdict ────────────────┘

   the generator is rewarded when the verdict is wrong
   the discriminator is rewarded when the verdict is right

Read the last two lines again. The two networks want opposite things from the same number. That is the adversarial part, and it is what makes a GAN different from every other model on this site.

The generator never sees a single real example. It only ever sees the teacher's verdict on its own work. Everything it knows about real data arrives second-hand.

Where you have already seen this

  • "This person does not exist" websites showing faces of people who were never born.
  • Photo enhancement that turns a blurry old family photo into a sharp one.
  • Voice generation in navigation apps and audiobooks, where a GAN turns a sound pattern into an audio waveform.
  • Video game and video upscaling, filling in detail that was never recorded.
  • Deepfakes. The same technology, used to put words in someone's mouth. This is a real harm, and it exists because the technique works.

What is honestly hard here

GANs are the most difficult model on this site to train, and pretending otherwise would waste your time.

There is no score that tells you it is working. In every other model, the loss falls and you know you are making progress. In a GAN, both losses hover around the same value whether the model is excellent or useless. Two networks pushing against each other produce a stalemate reading either way. People judge GAN training by looking at the output with their own eyes.

It can collapse to one answer. Suppose the generator finds a single output the teacher accepts. Producing that same thing forever is then a winning strategy. It stops making faces and starts making one face. This is called mode collapse. You will watch it happen in the Developer section, with a model that ignores half of its training data.

It can fall over entirely. If the teacher gets too good too fast, it rejects everything with total confidence. The forger then receives no useful hint about what to change, and training stops dead.

Diffusion models have taken over image generation. Every current image tool you have heard of is a diffusion model, not a GAN. They are far easier to train and they cover the whole range of the data instead of collapsing.

So why learn GANs at all? Three honest reasons.

  • They are fast at making things. A GAN produces an image in one pass. A diffusion model takes many steps. Where speed matters — real-time video, a phone, a game — that gap decides it.
  • They still win specific jobs. Sharpening a photo, and turning sound patterns into audio, are both still done with GANs in shipping products.
  • The adversarial idea is used inside other models. Several parts of modern image systems are trained with a small GAN-style critic attached, to stop outputs looking soft and smeared.

Remember this

  • A GAN is a maker and a judge, trained together, each trying to beat the other.
  • It exists because nobody could write down a score for "does this look real", so the score is learned.
  • It is powerful and genuinely unstable, and diffusion models have replaced it for most image work.

What to learn next

  • Autoencoders — the other route to generating data, and why its output is smoother.
  • Diffusion models — the approach that replaced GANs for image generation.
  • Loss functions — why a GAN's loss tells you so little.

Developer — Code and libraries.

Setup

bash
pip install torch

Both examples run on CPU in about ten seconds. Everything is one-dimensional, so you can read the result as numbers instead of squinting at pictures.

A GAN you can read as numbers

The real data is a bell curve with mean 7 and spread 1.5. The generator is never told either number. It sees only the discriminator's verdict on its own output.

tiny_gan.py
import torch
import torch.nn as nn

torch.manual_seed(0)

TRUE_MEAN, TRUE_SD = 7.0, 1.5      # the distribution the generator is never told


def real_batch(n):
    return torch.randn(n, 1) * TRUE_SD + TRUE_MEAN


# The forger: 4 random numbers in, one invented sample out.
G = nn.Sequential(nn.Linear(4, 32), nn.ReLU(), nn.Linear(32, 32), nn.ReLU(), nn.Linear(32, 1))
# The checker: one sample in, one score out. High means "I think this is real".
D = nn.Sequential(nn.Linear(1, 32), nn.ReLU(), nn.Linear(32, 32), nn.ReLU(), nn.Linear(32, 1))

optG = torch.optim.Adam(G.parameters(), lr=0.002)
optD = torch.optim.Adam(D.parameters(), lr=0.002)
bce = nn.BCEWithLogitsLoss()
N = 256

print("step    generated mean   generated sd   D(real)   D(fake)")
for step in range(1, 3001):
    real = real_batch(N)
    fake = G(torch.randn(N, 4)).detach()      # detach: this step trains D only
    optD.zero_grad()
    lossD = bce(D(real), torch.ones(N, 1)) + bce(D(fake), torch.zeros(N, 1))
    lossD.backward()
    optD.step()

    optG.zero_grad()
    made = G(torch.randn(N, 4))               # no detach: the gradient must reach G
    lossG = bce(D(made), torch.ones(N, 1))    # G wants D to call its output real
    lossG.backward()
    optG.step()

    if step in (1, 200, 600, 1500, 3000):
        with torch.no_grad():
            s = G(torch.randn(4000, 4))
            d_real = torch.sigmoid(D(real_batch(4000))).mean().item()
            d_fake = torch.sigmoid(D(s)).mean().item()
            print(f"{step:5d}   {s.mean().item():13.2f}   {s.std().item():12.2f}"
                  f"   {d_real:7.2f}   {d_fake:7.2f}")

print(f"\nthe real distribution:   mean {TRUE_MEAN:.2f}   sd {TRUE_SD:.2f}")
Output
step    generated mean   generated sd   D(real)   D(fake)
    1            0.10           0.06      0.46      0.51
  200            4.82           1.23      0.59      0.56
  600            7.11           1.60      0.50      0.50
 1500            7.42           1.37      0.48      0.48
 3000            7.03           1.27      0.50      0.50

the real distribution:   mean 7.00   sd 1.50

Four things in this table are worth slowing down for.

At step 1 the generator produces 0.10 ± 0.06. An untrained network is close to a constant. It is nowhere near the real data and does not know that the real data exists.

By step 600 it produces 7.11 ± 1.60, against a truth of 7.00 ± 1.50. It never saw one real number. It only ever saw a verdict, and the verdicts alone were enough to locate a distribution in space.

The two right-hand columns both converge to 0.50. This is the equilibrium the method is aiming for. D(real) = 0.50 means the discriminator, shown a genuine sample, is reduced to a coin flip. When the fake data matches the real data, no discriminator can do better than guessing — and that is exactly the theoretical fixed point, which the researcher block derives.

More training did not help. At step 600 the spread was 1.60. By step 3000 it was 1.27, further from the true 1.50 than it had been 2400 steps earlier. The two discriminator columns read the same at both points.

That last observation is the practical problem with GANs in one line. The numbers on your screen do not tell you which of those two models is better. For a real GAN you save checkpoints along the way and choose by looking.

Watching mode collapse happen

Now give the generator a target with two separate peaks: half the real data near -4, half near +4.

mode_collapse.py
import torch
import torch.nn as nn

torch.manual_seed(3)


def real_batch(n):
    # Half the real data sits near -4, half near +4. Two peaks, equally common.
    side = torch.randint(0, 2, (n, 1)).float() * 8 - 4
    return side + torch.randn(n, 1) * 0.4


G = nn.Sequential(nn.Linear(2, 32), nn.ReLU(), nn.Linear(32, 32), nn.ReLU(), nn.Linear(32, 1))
D = nn.Sequential(nn.Linear(1, 32), nn.ReLU(), nn.Linear(32, 32), nn.ReLU(), nn.Linear(32, 1))
optG = torch.optim.Adam(G.parameters(), lr=0.002)
optD = torch.optim.Adam(D.parameters(), lr=0.002)
bce = nn.BCEWithLogitsLoss()
N = 256

print("step   share of samples near -4   near +4")
for step in range(1, 4001):
    real = real_batch(N)
    optD.zero_grad()
    (bce(D(real), torch.ones(N, 1)) +
     bce(D(G(torch.randn(N, 2)).detach()), torch.zeros(N, 1))).backward()
    optD.step()
    optG.zero_grad()
    bce(D(G(torch.randn(N, 2))), torch.ones(N, 1)).backward()
    optG.step()
    if step in (50, 150, 400, 1000, 4000):
        with torch.no_grad():
            s = G(torch.randn(4000, 2))
            left = (s < 0).float().mean().item()
            print(f"{step:5d}   {left:23.2f}   {1 - left:7.2f}")
print("\nthe real data is 0.50 / 0.50")
Output
step   share of samples near -4   near +4
   50                      0.00      1.00
  150                      0.00      1.00
  400                      0.00      1.00
 1000                      0.00      1.00
 4000                      0.00      1.00

the real data is 0.50 / 0.50

Every sample the generator ever produces is on the right. Half of the training data might as well not exist.

Nothing crashed and no error was raised. Inside its chosen peak the generator's output is excellent, and the discriminator cannot fault it. From the generator's point of view this is a winning strategy, because its only instruction is "make the discriminator say real". Nothing in the objective rewards variety.

Change torch.manual_seed(3) to 0 and it collapses onto the left peak instead. Try 5 and it covers both, at roughly 0.46 / 0.54. Four of five seeds we tried collapsed, and which side won was luck.

That is the honest picture of GAN training. If your generated images ever start looking suspiciously similar to each other, this is what is happening, and the fix is a change of objective — Wasserstein loss, or minibatch features — rather than more training.

Common mistakes

Forgetting .detach() on the fake batch in the discriminator step. Without it, lossD.backward() sends gradients into the generator as well, and the generator gets trained to help the discriminator catch it. The model quietly refuses to improve and no error appears.

Calling optG.step() after optD.step() on stale outputs. The generator's loss must be computed from a fresh forward pass through the updated discriminator, as in the code above. Reusing the fake tensor gives the generator a gradient through a discriminator that no longer exists.

Reading the losses as progress. Shown above. Both losses sit near log 2 ≈ 0.69 at equilibrium regardless of quality. Look at samples.

Using the saturating generator loss. The original paper's log(1 − D(G(z))) has almost no gradient early in training, when D(G(z)) is near zero and the generator most needs signal. Maximise log D(G(z)) instead — which is what bce(D(made), ones) above computes. Every practical implementation does this.

Letting one side win. If the discriminator reaches near-perfect accuracy, the generator's gradient vanishes. Common repairs: a lower learning rate for D, label smoothing (train D towards 0.9 rather than 1.0), or spectral normalisation on D's layers.

Using batch normalisation in the discriminator with a gradient penalty. They conflict, because the penalty is defined per sample and batch norm mixes samples. Use layer or instance normalisation there.

Reaching for a GAN for image generation in a new project. Use a diffusion model unless you specifically need single-pass speed.

Try it yourself

In tiny_gan.py, give the discriminator a head start: run its update block five times per generator update, inside the loop.

Write down your prediction first. On this one-dimensional problem the result comes out at roughly mean 7.04, sd 1.51 against a truth of 7.00, 1.50 — better than the baseline's 1.27. A stronger judge helped here.

Now go the other way and update D only once every five generator steps. This one falls apart: the generator drifts to a mean near 14, about double the truth. A judge that has not kept up cannot tell the generator it has wandered off, so it wanders further.

Two things to take from that pair.

The widely repeated warning that a too-strong discriminator kills training is real, but it is a statement about high-dimensional data, where a discriminator can separate real from fake perfectly and then has nothing informative left to say. On an easy low-dimensional problem the opposite risk dominates entirely.

So the balance point is not a constant you can look up. It depends on your data, and finding it by hand for each new dataset is a large part of what training a GAN actually involves.

What to learn next

  • Autoencoders — the other route to generating data, and why its output is smoother.
  • Diffusion models — the approach that replaced GANs for image generation.
  • Loss functions — why a GAN's loss tells you so little.

Researcher — Mathematics and papers.

The minimax objective

Goodfellow et al. (2014) define a two-player zero-sum game between a generator G: Z → X and a discriminator D: X → (0,1):

min_G  max_D  V(D, G)  =  E_{x ~ p_data} [ log D(x) ]  +  E_{z ~ p_z} [ log(1 − D(G(z))) ]
  • p_data — the true data distribution over X
  • p_z — a fixed prior over the latent space, typically N(0, I)
  • p_g — the distribution induced on X by pushing p_z through G
  • D(x) — the probability the discriminator assigns to x being real

The optimal discriminator, and what the generator is really minimising

For fixed G, the inner maximisation has a closed-form solution obtained pointwise:

D*_G(x)  =  p_data(x) / ( p_data(x) + p_g(x) )

Substituting back gives the generator's effective objective:

C(G)  =  2 · JSD( p_data ‖ p_g )  −  log 4

where JSD is the Jensen-Shannon divergence. It is non-negative and zero only when p_data = p_g, so the global minimum is C(G) = −log 4 ≈ −1.386, attained uniquely at p_g = p_data, where D*(x) = 1/2 everywhere.

The Developer section's D(real) → 0.50 and D(fake) → 0.50 is this fixed point observed numerically. It is also why the losses carry no quality information: at equilibrium the discriminator loss is 2 log 2 ≈ 1.386 whether the game converged to a good G or is oscillating around a bad one.

The convergence proof in the paper holds in function space, assuming D is optimised to convergence at each step. Neither assumption is met in practice, and the gap between the two is where GAN training pathology lives.

Why the objective fails in practice

Vanishing gradients. Early in training D(G(z)) ≈ 0, and ∇_G log(1 − D(G(z))) is then near zero. The paper's own remedy is the non-saturating loss: maximise log D(G(z)) instead. Same fixed point, very different gradient magnitude when the generator is losing.

Disjoint supports. Arjovsky & Bottou (2017) show the deeper problem. If p_data is supported on a low-dimensional manifold in R^n — which is the standing assumption for natural images — and p_g is supported on another such manifold, then with probability one the two are disjoint or intersect in a measure-zero set. A perfect discriminator then exists, JSD is constant at log 2, and its gradient with respect to G is zero almost everywhere.

So the theory that justifies the objective breaks precisely under the conditions the method is applied in. This result reframed the whole field.

Wasserstein GANs

Arjovsky, Chintala & Bottou (2017) replace JSD with the Earth Mover distance, which stays finite and differentiable for disjoint supports. Under the Kantorovich-Rubinstein duality:

W(p_data, p_g)  =  sup_{‖f‖_L ≤ 1}  E_{x~p_data}[ f(x) ]  −  E_{x~p_g}[ f(x) ]

f is a 1-Lipschitz critic producing an unbounded real score rather than a probability. Enforcing the Lipschitz constraint is the entire engineering problem:

  • Weight clipping (original WGAN). Works, but biases the critic towards simple functions and is sensitive to the clip value.
  • Gradient penalty, WGAN-GP (Gulrajani et al., 2017): add λ E[ (‖∇_x̂ f(x̂)‖₂ − 1)² ] at points x̂ interpolated between real and fake samples. The standard choice.
  • Spectral normalisation (Miyato et al., 2018): divide each weight matrix by its largest singular value, estimated by one power iteration per step. Cheap, stable, and now the most widely used constraint.

The claimed benefit — that the critic loss correlates with sample quality — holds partially. It tracks progress within a single run and does not compare across architectures.

Mode collapse

The Developer section shows a generator abandoning half its target. Formally, p_g concentrates on a subset of supp(p_data), and no term in the objective penalises this. The reverse-KL-like behaviour of the practical non-saturating loss is mode-seeking rather than mode-covering, which is the same reason a variational autoencoder blurs while a GAN sharpens: the two sit on opposite sides of the same trade.

Interventions, roughly in order of adoption:

  • Minibatch discrimination (Salimans et al., 2016) gives D features computed across the batch, so a batch of near-identical samples is detectable. Minibatch standard deviation, its cheap descendant, is standard in the StyleGAN line.
  • Unrolled GANs (Metz et al., 2017) differentiate the generator update through k steps of the discriminator's future updates, so G cannot exploit a D that is about to adapt.
  • Two time-scale update rule (Heusel et al., 2017) uses different learning rates for G and D, with a proof of convergence to a local Nash equilibrium under conditions.
  • PacGAN (Lin et al., 2018) feeds D several samples jointly, making a collapsed generator detectable by construction.

Evaluation

Inception Score (Salimans et al., 2016) — exp( E_x KL( p(y|x) ‖ p(y) ) ). Rewards confident, diverse ImageNet class predictions. Insensitive to intra-class collapse and to matching the real data at all, since it never looks at real images. Barratt & Sharma (2018) document its failure modes; it should not be used alone.

Fréchet Inception Distance (Heusel et al., 2017) — fit Gaussians to Inception pool3 features of real and generated sets and compute

FID = ‖μ_r − μ_g‖²  +  Tr( Σ_r + Σ_g − 2 (Σ_r Σ_g)^{1/2} )

The standard metric, and it is biased upward at small sample sizes, sensitive to the resizing and JPEG pipeline, and reliant on a Gaussian assumption that Inception features do not satisfy. Kynkäänniemi et al. (2023) show FID's ranking can disagree with human judgement on modern generators. KID (Bińkowski et al., 2018) uses an unbiased MMD estimator and behaves better on small samples.

Precision and recall for generative models (Sajjadi et al., 2018; Kynkäänniemi et al., 2019) separate fidelity from coverage, which a single scalar cannot. This is the right instrument for measuring mode collapse, and it is under-used.

Architecture lineage

  • DCGAN (Radford, Metz & Chintala, 2015) — the convolutional recipe that made GANs trainable: strided convolutions instead of pooling, batch norm, no fully-connected hidden layers.
  • Progressive growing (Karras et al., 2018) — start at 4×4 and add layers, reaching 1024×1024 faces.
  • StyleGAN / StyleGAN2 / StyleGAN3 (Karras et al., 2019, 2020, 2021) — a mapping network to an intermediate latent space W, per-layer style modulation, and in v3 a redesign to make generation equivariant to translation and rotation, removing texture "sticking" in video.
  • BigGAN (Brock et al., 2019) — class-conditional at scale, with the truncation trick trading diversity for fidelity at sampling time.
  • Conditional GANs (Mirza & Osindero, 2014), pix2pix (Isola et al., 2017) for paired translation, and CycleGAN (Zhu et al., 2017) for unpaired translation via a cycle-consistency loss.

Where GANs stand now

Diffusion models overtook GANs on image synthesis (Dhariwal & Nichol, 2021) and the gap widened with latent diffusion. The reasons are structural: a diffusion model optimises a stable likelihood-like objective with no adversarial equilibrium to balance, and mode coverage follows from the training objective rather than needing to be defended.

Three roles remain, all genuine:

One-step generation. A GAN samples in a single forward pass. This is why diffusion distillation — turning a many-step sampler into a one-step generator — reaches for adversarial objectives: Adversarial Diffusion Distillation (Sauer et al., 2023) is a GAN loss wrapped around a distilled diffusion student. GigaGAN (Kang et al., 2023) scales a GAN to text-to-image at competitive quality with far faster sampling.

Perceptual super-resolution. SRGAN (Ledig et al., 2017), ESRGAN and Real-ESRGAN remain the practical choice for upscaling, where the adversarial term supplies plausible high-frequency detail that a reconstruction loss averages away.

Neural vocoders. HiFi-GAN (Kong, Kim & Bae, 2020) converts mel-spectrograms to waveforms in real time and is deployed widely in text-to-speech.

As a loss term inside other models. The autoencoder in VQGAN (Esser, Rombach & Ommer, 2021) and in latent diffusion is trained with a patch-based adversarial loss alongside reconstruction. Without it the decoder produces blurred output. So the adversarial objective survives inside the very systems that displaced GANs as standalone generators.

Papers

What to learn next

  • Autoencoders — the other route to generating data, and why its output is smoother.
  • Diffusion models — the approach that replaced GANs for image generation.
  • Loss functions — why a GAN's loss tells you so little.

What to learn next

These follow on from what you just read.

  • Deep Learning

    Autoencoders

    An autoencoder squeezes its input through a narrow middle and rebuilds it, and the squeezed middle turns out to be a more useful description of the data than the input was.

  • Deep Learning

    Transformers

    A transformer reads every word at once and lets each word decide which other words matter to it, which is the architecture behind almost every modern AI model.

  • Deep Learning

    PyTorch basics

    PyTorch is the library most neural networks are written in, and its central trick is keeping a record of every calculation so it can trace an error back to every weight that caused it.