CNN (convolutional neural network)
In one sentence A CNN learns small visual patterns and slides them across an image, so it can find the same pattern anywhere in the picture.
Updated
A convolutional neural network, or CNN, is a network that learns small visual patterns and slides each one across the whole image to find where it occurs.
Imagine a small stencil cut out in the shape of an edge. You slide it over a photograph, and wherever the photo matches the stencil, a light comes on. Now imagine hundreds of stencils, and imagine that nobody drew them — the network worked out the useful shapes by itself during training. That is a convolution layer.
The reason this design beat everything else at images for a decade is that it takes a fact about pictures seriously: a cat's ear looks like a cat's ear whether it sits in the top-left corner or the bottom-right. One set of learned weights is reused at every position, which cuts the number of parameters enormously and means the network does not have to relearn the same shape in every location.
Layers build on each other
Photo → conv: edges and colours
→ conv: corners, curves, textures
→ conv: eyes, ears, fur patches
→ conv: whole faces
→ "cat", 94% confidentEarly layers find tiny generic features; later layers combine them into recognisable parts. That layering is why a CNN trained on millions of general photos transfers so well — you keep the early layers, replace the last one, and train on your own few thousand images. CNNs still power most practical vision work: defect detection on production lines, medical scans, document scanning, on-device photo search. Vision transformers match or beat them at very large scale, while CNNs remain the stronger choice when your dataset is modest.
Where to go next
- Full lesson: Convolutional neural networks (CNN)
- Related terms: tensor, parameter, activation-function, transformer