What is computer vision?
Computer vision is teaching a computer to pull meaning out of pictures and video, when all it actually receives is a grid of brightness values.
- 17 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Computer vision is teaching a computer to pull meaning out of pictures and video.
Stand at an airport belt and watch the bags come round. Twenty of them are black, roughly the same size, from the same three brands. You spot yours in under a second.
Now try to write down how you did it. Not "it looked like mine" — the actual rule, in words, that would work for a stranger. You cannot. Nobody can.
That gap is the whole field. Seeing is effortless for you and brutally hard to describe. So it had to be learned from examples, not written down as rules.
Why it exists
A computer does not receive a picture. It receives a grid of dots, and a set of numbers describing the colour of each dot. Nothing else.
A photo, zoomed in until you can see the dots:
+--+--+--+--+--+--+
|::|::|::|##|##|##| Each little square is one "pixel",
+--+--+--+--+--+--+ which is one dot of the picture.
|::|::|::|##|##|##|
+--+--+--+--+--+--+ Each pixel stores how much red,
|::|::|::|##|##|##| green and blue it holds.
+--+--+--+--+--+--+A phone photo has about twelve million of those dots. Somewhere in that pile of numbers is a dog. Nothing in the numbers says so.
For thirty years, people tried to write the rules by hand. Find edges. Find corners. Measure the shape. It worked in a lab and fell over outdoors.
Move the lamp and every brightness value changes. Turn the dog sideways and the shape rules break. Photograph a black dog at night and half the numbers vanish. A rule that survived all of that was never written.
The answer, as with most of AI, was to stop writing rules. Show the computer a million labelled photos and let it work out the pattern itself.
How it works
camera -> grid of dots -> model finds patterns -> "a dog"The model does not look at the whole picture at once. It slides a small window across the image, over and over, checking for one thing at a time.
The early windows look for very plain things. A line here. A dark-to-bright edge there. Then later windows look at the answers from the earlier ones.
Layer 1 -> edges and blobs
|
v
Layer 2 -> corners, curves, simple textures
|
v
Layer 3 -> eyes, wheels, leaf shapes
|
v
Layer 4 -> faces, dogs, carsNobody told the model to build up in that order. It falls out of training. This layering is why it handles a dog in a new pose. Or a new light. Or a new corner of the frame.
Where you have already seen it
- FASTag at a toll gate. A camera reads your number plate as you drive through, in daylight, rain and headlight glare.
- Face unlock on your phone. It works with a beard, with glasses, in a dark room.
- Google Lens. Point at a plant, a signboard or a dish and get a name back.
- Google Photos. Search "cake" and it finds every birthday photo you never tagged.
- Crop disease apps. A farmer photographs a leaf and gets a likely disease and a treatment.
- Medical screening. Chest X-ray systems flag likely tuberculosis cases so a radiologist reviews them first. The machine sorts the queue. A doctor still decides.
Where it genuinely fails
This is worth knowing before you build anything.
It breaks outside what it was shown. A model trained on clean daytime road photos will not handle fog. It does not tell you it is out of its depth. It answers with the same confidence as always.
It can be fooled. Change a small number of pixels in ways your eye cannot see, and a confident "panda" becomes a confident "gibbon". This is a real, reproducible property of these models, not a party trick.
It is not equally good for everyone. Buolamwini and Gebru measured commercial face-analysis products in 2018. Error rates for lighter-skinned men were under one in a hundred. For darker-skinned women they were as high as one in three. The training photos were not representative, so the systems were not either.
That last one matters more than any accuracy number. A face system can decide who gets through a gate. Whose face it was trained on then becomes a question about people, not code.
Remember this
- A computer receives a grid of numbers, never a picture. All the meaning has to be worked out.
- Hand-written rules failed. Learning from large sets of labelled images is what worked.
- Models build up from edges to shapes to objects, one layer at a time, without being told to.
What to learn next
- How images are stored in a computer — the grid of numbers, in detail.
- Image classification — the first real task to build.
- Convolutional neural networks — the sliding-window model itself.
Developer — Code and libraries.
Two short programs, both pure NumPy. The first shows what a pixel really is. The second builds a convolution by hand — the operation that every vision model is made of.
No image files, no downloads, no GPU. If convolution has felt like a black box, twenty lines of loops will fix that.
Setup
pip install numpyA pixel is three numbers
import numpy as np
pixel = np.array([210, 105, 30]) # one pixel, as red / green / blue
print("red:", pixel[0], " green:", pixel[1], " blue:", pixel[2])
print("brightness:", round(float(pixel.mean()), 1))red: 210 green: 105 blue: 30 brightness: 115.0
Each channel runs from 0 to 255, so one pixel fits in three bytes. A 4000 by 3000 photo is therefore 36 million numbers before any compression.
Two things bite people here.
The range 0 to 255 comes from uint8 storage, and uint8 arithmetic wraps around. Add 10 to a pixel holding 250 and you get 4, not 260. Convert to float before doing arithmetic on images.
And the channel order is not universal. OpenCV reads images as blue-green-red, while Pillow, matplotlib and most model code use red-green-blue. Load with one and display with the other and your photo comes out blue. This is the single most common first-day bug in computer vision.
Convolution, written out by hand
A convolution slides a small grid of numbers, called a kernel or filter, across the image. At each position it multiplies overlapping numbers and adds them up.
That one operation, repeated with learned kernels, is the whole of a convolutional network.
import numpy as np
# A tiny 6x6 grayscale image: left half black, right half white.
img = np.zeros((6, 6))
img[:, 3:] = 255
# A vertical-edge detector. It asks: is the left side darker than the right?
kernel = np.array([[-1, 0, 1],
[-1, 0, 1],
[-1, 0, 1]], dtype=float)
kh, kw = kernel.shape
out = np.zeros((img.shape[0] - kh + 1, img.shape[1] - kw + 1))
for y in range(out.shape[0]):
for x in range(out.shape[1]):
patch = img[y:y + kh, x:x + kw] # the 3x3 window under the filter
out[y, x] = (patch * kernel).sum()
print("the image the computer stores:")
print(img.astype(int))
print()
print("what the edge detector found:")
print(out.astype(int))the image the computer stores: [[ 0 0 0 255 255 255] [ 0 0 0 255 255 255] [ 0 0 0 255 255 255] [ 0 0 0 255 255 255] [ 0 0 0 255 255 255] [ 0 0 0 255 255 255]] what the edge detector found: [[ 0 765 765 0] [ 0 765 765 0] [ 0 765 765 0] [ 0 765 765 0]]
Line by line
The output lights up only at the edge. Columns 1 and 2 sit where black meets white and score 765. Columns 0 and 3 sit over flat regions and score 0.
That is what an edge detector does, and it is worth stating plainly: a flat area produces zero. The filter responds to change, not to brightness. A pure white wall and a pure black wall both give 0.
Why 765. Each row of the window contributes 255 - 0, and there are three rows, giving 3 * 255. This is also a warning. Convolution output routinely exceeds 255 and can go negative, so it is not an image any more. Normalise before displaying, or everything clips to white.
Why the output is 4 by 4 and not 6 by 6. A 3-wide window cannot be centred on the outermost column, so you lose one pixel from each side. The general rule is out = in - kernel + 1. Stack twenty such layers and your image has shrunk by 40 pixels, which is why real networks pad the borders.
The kernel here was written by hand. In a CNN it is learned. This exact filter is the Prewitt operator from the 1970s. A trained network discovers filters very much like it in its first layer, without being told to. That convergence is one of the more satisfying results in the field.
Common mistakes
Channel order. OpenCV gives you BGR. Nearly everything else expects RGB. Convert explicitly with cv2.cvtColor(img, cv2.COLOR_BGR2RGB) rather than hoping.
Integer overflow. uint8 wraps silently. np.uint8(250) + np.uint8(10) is 4. Cast to float, do the work, clip to the 0 to 255 range, then cast back.
Forgetting to divide by 255. Most pretrained models expect inputs scaled to a 0-to-1 or normalised range. Feed raw 0-to-255 values and accuracy collapses with no error message at all.
Different preprocessing at training and inference. Resize, crop and normalise identically in both places. A model trained on centre-cropped 224-pixel squares and served stretched full frames will underperform for reasons that look mysterious for days.
Reading accuracy without looking at the images it got wrong. Always print or view a handful of failures. Ten minutes of that reveals more than a week of tuning.
Try it yourself
Rotate the kernel to detect horizontal edges instead:
kernel = np.array([[-1, -1, -1],
[ 0, 0, 0],
[ 1, 1, 1]], dtype=float)Run it against the same image:
what the edge detector found: [[0 0 0 0] [0 0 0 0] [0 0 0 0] [0 0 0 0]]
Every value is 0. A left-to-right split contains no horizontal edge, so this filter finds nothing at all.
Now change the image to img[3:, :] = 255, so the split runs top to bottom, and run both kernels again. The horizontal filter lights up with 765 across the middle two rows. The vertical filter now returns all zeros.
Each filter is blind to one orientation. That is the reason a real layer holds 64 or more kernels rather than one. Each is blind to most of the world, so you need many of them looking for different things.
What to learn next
- How images are stored in a computer — colour spaces, formats and precision.
- OpenCV — the standard library for this work.
- Convolutional neural networks — stacking learned kernels into a model.
Researcher — Mathematics and papers.
Image formation
An image is a sampled, quantised measurement of irradiance on a sensor plane:
I(x, y, c) = quantise( integral over t, over lambda of E(x, y, lambda, t) * S_c(lambda) dlambda dt )I(x, y, c)is the stored value at pixel(x, y)in channelc.Eis spectral irradiance at wavelengthlambdaand timet.S_cis the spectral sensitivity of the sensor's colour filter for channelc.quantisemaps to 8 bits in consumer formats, 10 to 14 bits in raw capture.
Three consequences follow, and each one produces a class of failure in deployed systems. The mapping from scene to pixels is many-to-one, so it is not invertible. It is scene-dependent through illumination. And it is lossy at both the integration and quantisation stages.
Nearly every robustness problem in vision has the same root. The model learns statistics of E and S_c, the capture conditions, while the task depends only on scene content.
Discrete convolution
(I * k)(x, y) = sum over i, j of I(x + i, y + j) * k(i, j)kis the kernel of sizek_h x k_w.iandjrange over the kernel extent.
Deep learning frameworks implement cross-correlation, without the kernel flip of the signal-processing definition. Since k is learned, the distinction is immaterial for training and matters only when porting classical filters.
Output spatial size with padding p, stride s, dilation d:
out = floor( (in + 2p - d*(k - 1) - 1) / s ) + 1Parameters for a conv layer with C_in input channels, C_out output channels, kernel k x k: C_out * (C_in * k^2 + 1), the +1 being bias.
FLOPs for output of spatial size H_out x W_out: approximately 2 * H_out * W_out * C_out * C_in * k^2, counting multiply and add separately.
The two structural properties that make this the right prior for images:
Translation equivariance. conv(shift(I)) = shift(conv(I)). A feature detected in one location is detected identically elsewhere, so training signal from one position transfers to all positions. Note that this is equivariance, not invariance — pooling and global aggregation supply approximate invariance later.
Parameter sharing. A 3 x 3 kernel on a 224-pixel input uses 9 weights per channel pair rather than one weight per input position. This is a strong, hand-specified prior about spatial locality, and it is precisely the prior a Vision Transformer discards.
Receptive field. Stacking L layers of kernel size k and stride 1 gives an effective receptive field of L * (k - 1) + 1. Two 3 x 3 layers cover 5 x 5 with 18 weights rather than 25, with an extra nonlinearity in between — the VGG argument.
The architectural sequence
| Year | Model | Key idea |
|---|---|---|
| 1998 | LeNet-5 (LeCun et al.) | Convolution and pooling, trained end to end by backpropagation |
| 2012 | AlexNet (Krizhevsky et al.) | ReLU, dropout, GPU training; ImageNet top-5 error 15.3% against 26.2% |
| 2014 | VGG (Simonyan & Zisserman) | Depth from stacked 3 x 3 kernels |
| 2015 | ResNet (He et al.) | Residual connections make 100-plus layers trainable |
| 2019 | EfficientNet (Tan & Le) | Compound scaling of depth, width and resolution together |
| 2021 | ViT (Dosovitskiy et al.) | Patch embeddings plus transformer; no convolution at all |
| 2021 | CLIP (Radford et al.) | Image-text contrastive pre-training, zero-shot transfer |
| 2022 | ConvNeXt (Liu et al.) | Modernised ConvNet matching ViT at equal compute |
| 2023 | SAM (Kirillov et al.) | Promptable segmentation, class-agnostic |
Two entries deserve comment. ResNet's contribution was optimisation, not representation. The residual path gives gradients a route that does not pass through every weight matrix. That removes the degradation which made deep plain networks worse than shallow ones.
And ConvNeXt matters as a negative result. It demonstrated that most of ViT's advantage came from training recipe and data scale rather than from architecture. Any claim that attention is intrinsically better for vision has to contend with that paper.
Data scale and inductive bias
ViT underperforms ResNets on ImageNet-1k alone and overtakes them when pre-trained on JFT-300M. The convolutional prior is worth roughly an order of magnitude of data. With enough data, the model learns the prior; without it, the prior must be supplied by the architecture.
This trade-off recurs throughout the field. Hold onto it when choosing an architecture for a domain with a few thousand labelled images, which describes most real projects.
Evaluation
| Task | Metric | Definition |
|---|---|---|
| Classification | Top-1 / Top-5 | Correct label is the argmax / among the 5 highest |
| Detection | mAP@[.5:.95] | Mean average precision over IoU thresholds 0.50 to 0.95 |
| Segmentation | mIoU | Mean over classes of intersection over union |
IoU = area(prediction AND ground truth) / area(prediction OR ground truth).
Calibration is measured separately and is routinely neglected. Guo et al. (2017), On Calibration of Modern Neural Networks (arXiv:1706.04599), showed modern networks are systematically overconfident. Expected calibration error partitions predictions into confidence bins and measures the gap between mean confidence and observed accuracy. If a downstream decision uses a confidence threshold, report ECE alongside accuracy or the threshold means nothing.
Robustness, and why benchmark numbers mislead
Four results should temper any reported accuracy figure.
Texture bias. Geirhos et al. (2019) showed ImageNet-trained CNNs classify primarily by texture, not shape. A cat silhouette with elephant skin texture is classified as elephant. Human observers do the opposite.
Natural distribution shift. Recht et al. (2019) built ImageNetV2 following the original collection protocol as closely as possible. Accuracy dropped 11 to 14 points across every model tested. The ranking of models was preserved, so the benchmark still compares models fairly. What it does not do is estimate real-world accuracy.
Corruptions and adversarial examples. ImageNet-C (Hendrycks & Dietterich, 2019) applies blur, noise, weather and compression artefacts, and accuracy falls sharply. Szegedy et al. (2014) showed imperceptible perturbations flip predictions, and adversarial training remains the only defence with a durable track record — at a real cost in clean accuracy.
Label noise. Beyer et al. (2020), Are we done with ImageNet? (arXiv:2006.07159), found that a substantial share of ImageNet validation "errors" are label problems, including images with multiple valid objects and a single ground-truth label. Reported top-1 above roughly 85 percent is partly measuring agreement with the labelling protocol.
Key references
- LeCun, Y., Bottou, L., Bengio, Y. & Haffner, P. (1998). Gradient-Based Learning Applied to Document Recognition. Proc. IEEE 86(11).
- Lowe, D. (2004). Distinctive Image Features from Scale-Invariant Keypoints. IJCV 60(2). SIFT.
- Dalal, N. & Triggs, B. (2005). Histograms of Oriented Gradients for Human Detection. CVPR.
- Deng, J. et al. (2009). ImageNet: A Large-Scale Hierarchical Image Database. CVPR.
- Krizhevsky, A., Sutskever, I. & Hinton, G. (2012). ImageNet Classification with Deep Convolutional Neural Networks. NeurIPS 25.
- He, K., Zhang, X., Ren, S. & Sun, J. (2016). Deep Residual Learning for Image Recognition. arXiv:1512.03385
- Dosovitskiy, A. et al. (2021). An Image is Worth 16x16 Words. arXiv:2010.11929
- Radford, A. et al. (2021). Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.00020
- Liu, Z. et al. (2022). A ConvNet for the 2020s. arXiv:2201.03545
- Kirillov, A. et al. (2023). Segment Anything. arXiv:2304.02643
- Oquab, M. et al. (2023). DINOv2: Learning Robust Visual Features without Supervision. arXiv:2304.07193
- Buolamwini, J. & Gebru, T. (2018). Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. PMLR 81.
Current state and open problems
The field has moved to large pre-trained backbones plus light task-specific heads. DINOv2 produces features strong enough for dense prediction without labels; CLIP gives an open-vocabulary interface; SAM gives class-agnostic segmentation from prompts. Training a vision model from scratch is now the unusual choice.
Four problems remain genuinely open.
Out-of-distribution behaviour. Models remain confident when they are wrong. Detection of distribution shift at inference is unsolved, and it is the difference between a demo and a deployable system.
3-D and physical reasoning. Recognition is strong. Reasoning about occlusion, contact, support and stability from images remains weak.
Video understanding. Most video models are still image models applied per frame with light temporal pooling. Long-horizon temporal reasoning is far behind image recognition.
Evaluation and fairness. The Gender Shades methodology — disaggregated error rates across intersectional subgroups — should be routine for any deployed vision system and still is not. Aggregate accuracy conceals exactly the disparities that matter when a system is deciding about people.
What to learn next
- Convolutional neural networks — architecture and training dynamics in depth.
- Image classification — losses, metrics and transfer learning.
- Vision transformers — attention applied to patches.