Diffusion model
In one sentence A diffusion model generates images by learning to remove noise step by step, starting from pure static and sculpting it into a picture.
Updated
A diffusion model generates images (or audio, or video) by starting from pure random noise and removing the noise in many small steps until a coherent picture remains.
The training idea sounds backwards and is the whole trick. Take a real photo. Add a little static; add more; keep going until nothing but noise remains — that direction is easy and mechanical. Now train a network to run the film in reverse: at every noise level, predict what noise was added, and strip a little of it away. A restorer who has cleaned ten thousand soot-darkened paintings learns, stroke by stroke, what soot looks like versus what painting looks like.
Generation then starts from pure noise — no painting underneath at all — and applies the learned cleaning repeatedly. Twenty to fifty steps later, a novel image has been sculpted out of static, the way one might see a face in the clouds and progressively make it real. Text guidance ("a chai stall in the rain, watercolour") steers every denoising step toward images matching the caption, using text embeddings learned from millions of captioned pictures.
This recipe powers Stable Diffusion, Midjourney, DALL·E and the image modes of frontier assistants, having displaced GANs as the dominant image generator — easier to train, better coverage of variety. The efficiency trick worth knowing: latent diffusion runs the whole process in a compressed latent-space (via an autoencoder) rather than on raw pixels, which is what made image generation fit on consumer GPUs. The main cost stays visible in products: many steps per image means seconds of wait, and step-reduction research is ongoing.
Where to go next
- Full lesson: How LLMs work
- Related terms: gan, latent-space, multimodal, autoencoder