Deep Learning

Convolutional neural networks (CNN)

A CNN looks at a small patch of an image at a time and slides the same pattern detector across the whole picture, which makes vision models small enough to train.

Read these first

On this page 7
  1. Why this had to be invented
  2. How it works
  3. The other piece: shrinking
  4. Where you have already seen this
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A CNN is a neural network for images. It studies one small patch at a time, sliding the same detector across the whole picture.

Think about a power cut at night. You hold a torch and sweep it across the wall, hunting for the switchboard. You never see the whole wall at once — only a small bright circle that you move around.

Your eyes do not change as you sweep. The same pair of eyes checks every patch. If you can spot a switchboard in the top corner, you can spot one in the bottom corner too.

A CNN works this way. One small detector, slid across the entire image, checking every patch with the same set of eyes.

Why this had to be invented

Try feeding a photo to an ordinary neural network. A modest phone photo has around fifty thousand pixels, and each colour channel triples that.

An ordinary network wires every pixel to every neuron — one small unit that adds up its inputs and passes on a single number. Do that here and you get around 150 million knobs. That is in the first layer alone, meaning the very first row of those units. It is too many to train, too many to store, and far more than the data can support.

There was a second, worse problem. An ordinary network learns each pixel position separately. Teach it to spot a cat in the centre of a photo. It has learned nothing about cats in the top-left corner, and every position starts from scratch.

Both problems disappear with the torch idea. One small detector, reused everywhere, is a tiny number of knobs — and it works at every position automatically.

How it works

The small detector is called a filter: a little grid of numbers that lights up when it finds its pattern.

   image (a big grid of pixels)
   ┌───────────────┐
   │ ▒▒▒░░░░       │        filter (a small grid)
   │ ▒▒▒░░░░       │        ┌───┐
   │ ▒▒▒░░░░       │   +    │ ▓ │  slide it over every position
   │ ▒▒▒░░░░       │        └───┘
   └───────────────┘
            │
            ↓
   a map showing WHERE that pattern was found

The first layer's filters find dull, simple things: an edge here, a change of colour there. Nothing exciting on its own.

The second layer looks at the first layer's maps and combines edges into corners and curves. The third combines those into shapes like an eye or a wheel. A later layer combines those into "cat" or "car".

Nobody designs this ladder. The network works out what each filter should look for, from examples alone.

The other piece: shrinking

Between layers, a CNN throws away detail on purpose. It takes each small square of a map and keeps only the strongest response. This is called pooling — shrinking a map by keeping the highest value from each patch.

   4 5 │ 1 0            keep the biggest
   2 1 │ 3 2    →       from each block    →    5 3
   ────┼────                                    8 6
   8 3 │ 6 1
   0 2 │ 4 5

This matters more than it looks. Knowing an edge is "roughly here" is usually enough. It makes every later step cheaper, and less fussy about exact position.

Where you have already seen this

  • Google Photos finding every picture of your dog.
  • Face unlock on your phone.
  • Number plate readers at toll booths on the highway.
  • Medical scan tools that flag a suspicious area for a doctor to review.

What is honestly hard here

CNNs are startlingly easy to fool. Change a handful of pixels in ways your eye cannot detect, and a confident "panda" becomes a confident "gibbon".

This is a real, active research problem called adversarial vulnerability, and it is not solved. It is one reason these systems should not be trusted alone for anything with serious consequences.

Remember this

  • A CNN slides one small filter across the whole image, reusing the same knobs everywhere.
  • That reuse cuts the number of knobs enormously, and finds a pattern anywhere it appears.
  • Early layers find edges; later layers combine them into objects.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

The PyTorch section further down needs pip install torch, but the core idea needs nothing beyond NumPy.

Convolution, written out in full

convolution.py
import numpy as np

# A 6x6 greyscale "image": the left half is bright, the right half is dark.
# There is one vertical edge in it, between column 2 and column 3.
image = np.array([
    [1., 1., 1., 0., 0., 0.],
    [1., 1., 1., 0., 0., 0.],
    [1., 1., 1., 0., 0., 0.],
    [1., 1., 1., 0., 0., 0.],
    [1., 1., 1., 0., 0., 0.],
    [1., 1., 1., 0., 0., 0.],
])

vertical = np.array([[1., 0., -1.],      # bright on the left, dark on the right
                     [1., 0., -1.],
                     [1., 0., -1.]])

horizontal = np.array([[1., 1., 1.],     # bright on top, dark on the bottom
                       [0., 0., 0.],
                       [-1., -1., -1.]])


def show(img, title):
    print(title)
    for r in img:
        print("   " + " ".join("#" if v > 0.5 else "." for v in r))


def convolve(img, kernel):
    kh, kw = kernel.shape
    out_h = img.shape[0] - kh + 1        # no padding, so the result is slightly smaller
    out_w = img.shape[1] - kw + 1
    out = np.zeros((out_h, out_w))
    for i in range(out_h):
        for j in range(out_w):
            patch = img[i:i + kh, j:j + kw]
            out[i, j] = np.sum(patch * kernel)   # one number per window position
    return out


def maxpool2(x):
    out = np.zeros((x.shape[0] // 2, x.shape[1] // 2))
    for i in range(out.shape[0]):
        for j in range(out.shape[1]):
            out[i, j] = x[2 * i:2 * i + 2, 2 * j:2 * j + 2].max()
    return out


show(image, "the image (# = bright, . = dark):")

print("\nvertical-edge filter slid over it:")
v_map = convolve(image, vertical)
print(v_map)

print("\nhorizontal-edge filter slid over the same image:")
print(convolve(image, horizontal))

print("\nafter 2x2 max pooling:")
print(maxpool2(v_map))

# Move the edge one column right and the response moves with it.
moved = np.roll(image, 1, axis=1)
show(moved, "\nsame edge, shifted one column right:")
print("vertical filter response:")
print(convolve(moved, vertical))

print("\n--- why not use a plain dense layer on images? ---")
pixels = 224 * 224 * 3
dense = pixels * 1000 + 1000                 # every pixel wired to every one of 1000 units
conv = 3 * 3 * 3 * 64 + 64                   # 64 filters, each 3x3 across 3 colour channels
print(f"dense layer on a 224x224 colour photo: {dense:,} parameters")
print(f"conv layer with 64 3x3 filters:        {conv:,} parameters")
Output
the image (# = bright, . = dark):
   # # # . . .
   # # # . . .
   # # # . . .
   # # # . . .
   # # # . . .
   # # # . . .

vertical-edge filter slid over it:
[[0. 3. 3. 0.]
 [0. 3. 3. 0.]
 [0. 3. 3. 0.]
 [0. 3. 3. 0.]]

horizontal-edge filter slid over the same image:
[[0. 0. 0. 0.]
 [0. 0. 0. 0.]
 [0. 0. 0. 0.]
 [0. 0. 0. 0.]]

after 2x2 max pooling:
[[3. 3.]
 [3. 3.]]

same edge, shifted one column right:
   . # # # . .
   . # # # . .
   . # # # . .
   . # # # . .
   . # # # . .
   . # # # . .
vertical filter response:
[[-3.  0.  3.  3.]
 [-3.  0.  3.  3.]
 [-3.  0.  3.  3.]
 [-3.  0.  3.  3.]]

--- why not use a plain dense layer on images? ---
dense layer on a 224x224 colour photo: 150,529,000 parameters
conv layer with 64 3x3 filters:        1,792 parameters

Reading the output carefully

The vertical filter fired at the edge and nowhere else. The map is 0 3 3 0 on every row. Zero where the patch is uniformly bright or uniformly dark, and 3 where the window straddles the boundary. A filter is a pattern detector, and this is what "detecting" looks like numerically.

The horizontal filter returned all zeros. Same image, different filter, nothing found — correctly, because there is no horizontal edge in it. Different filters see different things, which is why a real layer has 64 or 256 of them rather than one.

Shifting the image shifted the response. The pair of 3s moved one column right, tracking the edge. This property is called translation equivariance: move the input, and the output moves the same way. It is the payoff for reusing one filter everywhere.

The -3 values deserve an explanation rather than being glossed over. np.roll wraps the rightmost column around to the left, which creates a second edge at the left border — dark then bright, the opposite polarity. Our filter is [1, 0, -1], so bright-to-dark gives +3 and dark-to-bright gives -3. The sign carries the direction of the edge, and that is real information, not an artefact.

Parameter counts. 150,529,000 against 1,792. That is a factor of about 84,000, and it is the whole reason CNNs exist.

Shapes in a real network

cnn_shapes.py
import torch
import torch.nn as nn

net = nn.Sequential(
    nn.Conv2d(1, 8, kernel_size=3, padding=1),    # 1 grey channel in, 8 feature maps out
    nn.ReLU(),
    nn.MaxPool2d(2),                              # halves height and width
    nn.Conv2d(8, 16, kernel_size=3, padding=1),
    nn.ReLU(),
    nn.MaxPool2d(2),
    nn.Flatten(),
    nn.Linear(16 * 7 * 7, 10),                    # 10 digit classes
)

x = torch.zeros(1, 1, 28, 28)                     # one fake 28x28 image
print(f"{'input':12s} -> {tuple(x.shape)}")
for layer in net:
    x = layer(x)
    print(f"{layer.__class__.__name__:12s} -> {tuple(x.shape)}")

total = sum(p.numel() for p in net.parameters())
print("\ntrainable parameters:", f"{total:,}")
Output
input        -> (1, 1, 28, 28)
Conv2d       -> (1, 8, 28, 28)
ReLU         -> (1, 8, 28, 28)
MaxPool2d    -> (1, 8, 14, 14)
Conv2d       -> (1, 16, 14, 14)
ReLU         -> (1, 16, 14, 14)
MaxPool2d    -> (1, 16, 7, 7)
Flatten      -> (1, 784)
Linear       -> (1, 10)

trainable parameters: 9,098

Follow the two numbers that move in opposite directions. Spatial size shrinks 28 → 14 → 7 while channel count grows 1 → 8 → 16. Every CNN does this: trade where something is for what it is. By the final layers the network barely knows position but knows a great deal about content.

The tensor layout is (batch, channels, height, width), called NCHW. TensorFlow defaults to NHWC instead, with channels last — a frequent source of confusion when porting code.

Common mistakes

Getting Linear input size wrong. 16 * 7 * 7 = 784 has to match whatever Flatten produces. Change a pooling layer or a padding value and this breaks with RuntimeError: mat1 and mat2 shapes cannot be multiplied. Print shapes layer by layer, exactly as above, instead of computing it in your head. Or use nn.LazyLinear, which infers it on the first forward pass.

Forgetting padding=1 with a 3x3 kernel. Without padding, every conv layer shrinks the image by two pixels. Stack twenty of them and a 28x28 image runs out of pixels. padding=1 with kernel_size=3 keeps the size unchanged, which is why that pairing is nearly universal.

Forgetting the batch dimension. A single image is (1, 28, 28), but Conv2d wants (N, C, H, W). Use x.unsqueeze(0) to add it.

Normalising incorrectly, or not at all. Feed raw 0-255 pixel values and training is slow and unstable. Scale to 0-1, then standardise with the dataset's channel means and standard deviations.

Leaving the model in training mode when evaluating. Call model.eval() before inference and wrap it in torch.no_grad(). Otherwise dropout stays active and batch-norm keeps updating its statistics, and your accuracy numbers are wrong.

Try it yourself

Change vertical to a diagonal detector: [[1,0,0],[0,0,0],[0,0,-1]]. Predict the response on the vertical-edge image before running it. Then build an image with a horizontal edge and confirm the horizontal filter finally fires.

What to learn next

Researcher — Mathematics and papers.

The operation

What deep learning calls convolution is technically cross-correlation — the kernel is not flipped. For input $X$, kernel $K$ of size $k_h \times k_w$, and output $Y$:

$$ Y[i, j] = \sum_{m=0}^{k_h - 1} \sum_{n=0}^{k_w - 1} X[i + m,\; j + n] \cdot K[m, n] $$

Since $K$ is learned, the flip is absorbed into the learned parameters and the distinction has no practical consequence.

With $C_{\text{in}}$ input channels and $C_{\text{out}}$ filters:

$$ Y[c_{\text{out}}, i, j] = b[c_{\text{out}}] + \sum_{c=0}^{C_{\text{in}}-1} \sum_{m} \sum_{n} X[c,\, i \cdot s + m,\; j \cdot s + n] \cdot K[c_{\text{out}}, c, m, n] $$

Where $s$ is the stride. Each filter spans the full channel depth, so a "3x3 filter" on RGB input is really a $3 \times 3 \times 3$ tensor.

Output size

$$ H_{\text{out}} = \left\lfloor \frac{H_{\text{in}} + 2P - D(k - 1) - 1}{S} \right\rfloor + 1 $$

Where $P$ is padding, $S$ is stride, $k$ is kernel size, and $D$ is dilation. With $D = 1$ this reduces to the familiar $\lfloor (H + 2P - k)/S \rfloor + 1$.

Same padding requires $P = \lfloor (k-1)/2 \rfloor$ at $S = 1$, which is why odd kernel sizes dominate — an even kernel cannot pad symmetrically.

Cost

For one conv layer:

  • Parameters: $C_{\text{out}} \cdot C_{\text{in}} \cdot k_h \cdot k_w + C_{\text{out}}$ — independent of input resolution.
  • FLOPs: $\approx 2 \cdot H_{\text{out}} W_{\text{out}} \cdot C_{\text{out}} C_{\text{in}} k_h k_w$ — strongly dependent on resolution.

That split explains a common surprise: doubling image resolution leaves the parameter count unchanged while quadrupling compute.

Depthwise separable convolution (Chollet, 2017, Xception; Howard et al., 2017, MobileNets) factorises the operation into a per-channel spatial convolution followed by a $1 \times 1$ pointwise convolution, cutting cost by a factor of:

$$ \frac{1}{C_{\text{out}}} + \frac{1}{k^2} $$

For $k=3$ and $C_{\text{out}} = 256$, that is roughly a $9\times$ reduction — the basis of every mobile vision architecture.

Receptive field

The receptive field is the region of input influencing one output unit. Stacking layers with kernel size $k_l$ and stride $s_l$:

$$ r_{l} = r_{l-1} + (k_l - 1) \prod_{i=1}^{l-1} s_i $$

Two stacked $3\times3$ convolutions give a $5\times5$ receptive field using $2 \cdot 9 C^2 = 18C^2$ parameters, against $25C^2$ for a single $5\times5$ — fewer parameters, plus an extra non-linearity. This argument, from Simonyan and Zisserman (2015), VGG, is why $3\times3$ became the standard kernel.

Luo et al. (2016), Understanding the effective receptive field, show the effective receptive field is Gaussian-weighted and substantially smaller than the theoretical one, growing as $O(\sqrt{L})$ rather than $O(L)$.

Equivariance, and its limits

Convolution is equivariant to translation: $f(T_{\delta} x) = T_{\delta} f(x)$ for a shift $T_\delta$. Adding global pooling yields approximate invariance.

Three honest caveats:

  • Strided convolution and pooling break exact equivariance by violating the Nyquist criterion. Zhang (2019), Making convolutional networks shift-invariant again, shows a one-pixel shift can change predictions, and fixes it with anti-aliased downsampling.
  • Convolution is not equivariant to rotation or scale. Those must be learned from augmented data, or built in explicitly (Cohen and Welling, 2016, Group equivariant CNNs).
  • Zero padding leaks absolute position information, so networks are less translation-invariant than the theory suggests (Islam et al., 2020).

Architectural lineage

YearModelContribution
1980Neocognitron (Fukushima)Hierarchical local features, no backpropagation
1998LeNet-5 (LeCun et al.)Convolution trained end to end by backpropagation
2012AlexNet (Krizhevsky et al.)ReLU, dropout, GPUs; halved ImageNet error
2014VGG (Simonyan and Zisserman)Depth via stacked 3x3 kernels
2015ResNet (He et al.)Residual connections; trained 152 layers
2017MobileNet / XceptionDepthwise separable convolution
2019EfficientNet (Tan and Le)Compound scaling of depth, width and resolution
2022ConvNeXt (Liu et al.)CNN modernised with transformer design choices; matches ViT

Residual connections are the single most important entry. Before ResNet, deeper networks performed worse on the training set — an optimisation failure, not overfitting. The identity shortcut $y = \mathcal{F}(x) + x$ gives gradients a path with derivative exactly $1$, and depth became trainable.

CNNs versus vision transformers

ViTs (Dosovitskiy et al., 2021) overtook CNNs on large-scale benchmarks, but the comparison is frequently misreported. ViTs lack the locality and translation-equivariance priors that CNNs build in, so they need far more data or heavy augmentation to match them. Liu et al. (2022), A ConvNet for the 2020s, show a modernised CNN matches Swin Transformer at equal compute — evidence that much of the reported gap came from training recipes rather than architecture.

CNNs remain the better choice for small datasets, high-resolution dense prediction, and edge deployment.

What to learn next