Image classification
Image classification is giving a whole picture one label from a fixed list of choices, which is the first real task most vision projects start with.
- 17 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Image classification is looking at a whole picture and giving it one label from a fixed list.
Walk through a mango mandi early in the morning. Workers pick up each mango, turn it once, and drop it into a crate. Grade one here, grade two there, reject in the third.
One fruit, one crate, decided in under a second. Watch closely and you notice something. There is no crate for guavas. If a guava came down the line, it would still land in a mango crate. Those are the only crates there are.
Hold on to that. It is the most important limitation of image classification, and it catches people out constantly.
Why it exists
Almost every useful vision question starts as a sorting question.
Is this X-ray normal or does it need a doctor's eyes today? Is this leaf healthy or infected? Is this photo a receipt or a selfie? Is this weld good or cracked?
Each one is a pile of pictures and a small set of crates. Get that working and you have solved a real problem for real people, with the simplest tool in computer vision.
It is also the doorway to everything else. Detection, segmentation and face recognition all reuse the machinery you build here.
How it works
The model never picks one answer directly. It gives a score to every label on the list, and the highest score wins.
+-------------------------+
| score for every label |
photo -> model ->| |
| dog ####### <- highest, so this is the answer
| cat #
| horse #
| cow .
+-------------------------+The scores always add up to a whole, like slicing one roti between the labels. So a big slice for "dog" leaves small slices for everything else.
That slicing has an odd side effect. Show it a photo of a guava, and the roti still gets sliced between dog, cat, horse and cow. Something must win. The model has no way to say "this is not on my list".
Classification is not the only vision task
People mix these up constantly, so here they are side by side.
CLASSIFICATION "there is a dog in this photo" one label, whole image
DETECTION "there is a dog, here, in this box" label plus location
SEGMENTATION "these exact dots are the dog" a label for every dotStart with classification. It needs the least labelling work by a wide margin, and it answers more questions than people expect.
Where you have already seen it
- Google Photos albums. Every photo gets sorted into people, places and things.
- Your spam folder for images. Screenshots and memes get separated from documents.
- Plant and bird apps. Photograph it, get a species name and a confidence.
- Bank apps. Photograph a cheque and it checks the picture is a cheque before reading it.
- Factory quality checks. A camera above the belt sorts good parts from cracked ones.
Be careful with the accuracy number
You will train something and it will report 99 out of 100 correct. That feels like finishing. It usually is not.
If the crates are lopsided, accuracy lies. Say 99 of every 100 X-rays are normal. A model that answers "normal" every single time, with no thought at all, scores 99 out of 100. It also misses every sick patient. The number looks excellent and the system is worthless.
A test that resembles the real world is the only test that counts. Photos taken in your office, with your lighting, on your phone, are not the photos your users will send.
Always look at the pictures the model got wrong. Ten minutes of looking beats a week of tuning. This is the habit that separates people who ship working systems from people who ship good numbers.
Remember this
- Classification gives one label per picture, chosen from a fixed list decided before training.
- It cannot say "none of these". Anything you show it lands in some crate.
- Accuracy alone is misleading whenever your classes are uneven. Look at what it got wrong.
What to learn next
- Object detection — finding where things are, not only what is there.
- Convolutional neural networks — the model doing the work.
- Model evaluation — metrics that survive uneven classes.
Developer — Code and libraries.
You can train a working image classifier in fifteen lines, on a CPU, in about two seconds. Do that before touching a neural network.
scikit-learn ships a small handwritten-digit dataset inside the library, so there is nothing to download. That makes it perfect for learning the shape of the problem.
Setup
pip install scikit-learnNo GPU, no dataset download, no network access at runtime.
First, look at your data
Never train on data you have not looked at. Here is what the model actually receives.
from sklearn.datasets import load_digits
digits = load_digits() # ships inside scikit-learn, nothing to download
image = digits.images[0] # one 8x8 grid of ink values, 0 to 16
for row in image:
# heavier ink prints as a heavier character, so the shape becomes visible
print("".join("#" if v > 8 else "+" if v > 2 else "." for v in row))
print("label:", digits.target[0])
print("dataset:", digits.images.shape)..+##... ..####+. .+#..#+. .+#..++. .++..#+. .+#..#+. ..#+##.. ..+##... label: 0 dataset: (1797, 8, 8)
A zero, drawn with 64 numbers. That is the entire input. No strokes, no order of writing, no pen pressure — 64 brightness values in a grid.
Train the classifier
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC
X, y = load_digits(return_X_y=True) # X is 1797 rows of 64 numbers each
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=0)
clf = SVC(gamma=0.001) # gamma sets how local each decision is
clf.fit(X_train, y_train)
predicted = clf.predict(X_test)
correct = (predicted == y_test).sum()
print("trained on:", X_train.shape[0], "images")
print("tested on: ", X_test.shape[0], "images")
print(f"correct: {correct} out of {len(y_test)}")
print(f"accuracy: {correct / len(y_test):.3f}")
for true, guess in zip(y_test[predicted != y_test], predicted[predicted != y_test]):
print(f" mistake: a {true} was read as a {guess}")trained on: 1347 images tested on: 450 images correct: 448 out of 450 accuracy: 0.996 mistake: a 9 was read as a 5 mistake: a 5 was read as a 9
Your counts should match exactly, because random_state=0 fixes the split and SVC has no randomness in fitting. If your scikit-learn version differs by a lot, the two mistakes could shift by one either way.
Line by line
random_state=0 fixes the shuffle before splitting. Without it, every run gives a different score and you cannot tell a real improvement from noise. Fix the seed, and report results across several seeds when it matters.
Fit on train, score on test. train_test_split holds back 450 images the model never sees during fitting. Score on data the model trained on and you measure memorisation, which is always near-perfect and always meaningless.
gamma=0.001 controls how far a single training example's influence reaches in an RBF kernel. Large gamma means each example only affects its immediate neighbourhood, which memorises. Small gamma means broad, smooth decision boundaries. This value is the standard one for this dataset and was found by search, not derived.
Printing the mistakes is the most valuable line in the file. Both errors are 5 against 9, in both directions. At 8 by 8 resolution, a 5 and a 9 differ by a handful of pixels in the upper curve. The model is not failing randomly — it is failing on the genuinely ambiguous pair, which is what a healthy model looks like.
Do not read too much into 0.996
This dataset is unusually kind, and it is worth naming why.
Every digit is centred, upright, cropped and drawn in the same style. There is no background, no lighting variation, no camera angle. Real photographs have every one of those problems.
The dataset is also perfectly balanced — around 180 examples of each of the ten digits — so accuracy is a fair measure here. It will not be fair on your data.
Treat this as learning the shape of a classification pipeline, not as evidence that image classification is a solved problem. It is not.
When to move up to a neural network
Straight lines through pixel values stop working the moment position matters. That is nearly always, for real photos.
The step up is a convolutional network. The practical route is transfer learning. Take a model already trained on millions of images, keep everything except the last layer, and train a new last layer on your own data.
That works with a few hundred examples per class, on a CPU if you are patient. Training from scratch needs tens of thousands per class. Almost nobody should train from scratch.
Common mistakes
Scaling mismatch. These digits are 0 to 16. Photographs are 0 to 255. Pretrained models expect their own specific normalisation. Feed the wrong range and accuracy collapses with no error raised. Apply exactly the same preprocessing at training and at inference.
Leaking data between the splits. Near-duplicate images — burst shots, frames from one video, the same object photographed twice — must land on the same side of the split. Otherwise the test set contains what the model already learned, and the score is fiction. Split by source, patient or session, never by row.
Reporting accuracy on uneven classes. With 95 percent normal cases, answering "normal" always scores 95 percent. Report per-class recall and a confusion matrix. sklearn.metrics.classification_report gives both in one line.
Forgetting the closed-set problem. Your model assigns every input to one of its known classes. Feed it a class it never saw and it answers confidently and wrongly. If unknown inputs are possible in production, add a rejection class or a confidence threshold, and validate that threshold on real unknowns.
Tuning against the test set. Try twenty settings, keep the best test score, and that score is now optimistic. Use a separate validation split for tuning and touch the test set once.
Try it yourself
Replace SVC(gamma=0.001) with SVC(), which uses the default gamma="scale", and re-run. Then look at the mistakes it prints, not only the accuracy.
After that, get the per-class report instead of a single number:
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC
from sklearn.metrics import classification_report
X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=0)
predicted = SVC(gamma=0.001).fit(X_train, y_train).predict(X_test)
print(classification_report(y_test, predicted)) precision recall f1-score support
0 1.00 1.00 1.00 37
1 1.00 1.00 1.00 43
2 1.00 1.00 1.00 44
3 1.00 1.00 1.00 45
4 1.00 1.00 1.00 38
5 0.98 0.98 0.98 48
6 1.00 1.00 1.00 52
7 1.00 1.00 1.00 48
8 1.00 1.00 1.00 48
9 0.98 0.98 0.98 47
accuracy 1.00 450
macro avg 1.00 1.00 1.00 450
weighted avg 1.00 1.00 1.00 450Read down the recall column. Classes 5 and 9 sit at 0.98 and every other digit is perfect.
Notice also that the accuracy row prints 1.00 while two images were in fact wrong. It is rounding 0.996 up. That is exactly why the raw count of mistakes belongs next to any percentage you report.
The support column is the number of test images per class, and it is the first thing to check on your own data. Even support means accuracy is fair. Lopsided support means it is not.
What to learn next
- Convolutional neural networks — the architecture for real photographs.
- Object detection — adding location to the label.
- Model evaluation — precision, recall and confusion matrices in depth.
Researcher — Mathematics and papers.
Formulation
Given input x in R^(H x W x C) and label y in {1, ..., K}, learn f_theta mapping images to a distribution over K classes.
H,W,Care height, width and channel count.Kis the number of classes, fixed before training.thetais the parameter vector.
The network emits logits z = f_theta(x) in R^K, converted to probabilities by softmax:
p_i = exp(z_i) / sum over j = 1..K of exp(z_j)z_iis the logit for classi.p_iis the predicted probability of classi, with allp_isumming to 1.
Note that softmax is shift-invariant: adding a constant to every logit leaves p unchanged. Implementations subtract max(z) before exponentiating for numerical stability, which is why raw logits are not comparable across examples.
Cross-entropy loss
L = -(1/N) * sum over n = 1..N of sum over i = 1..K of y_{n,i} * log(p_{n,i})Nis the batch size.y_{n,i}is 1 if examplenhas classi, else 0 (one-hot).- For hard labels this collapses to
L = -(1/N) * sum over n of log(p_{n, y_n}).
The gradient with respect to the logits is exactly dL/dz_i = p_i - y_i. This clean form is the reason softmax and cross-entropy are paired: the softmax Jacobian and the log cancel, leaving a difference of probabilities.
Label smoothing (Szegedy et al., 2016) replaces the one-hot target with y_smooth = (1 - eps) * y + eps / K, where eps is typically 0.1. It bounds the logit gap the model drives toward, improving calibration and usually top-1 by a few tenths of a point. It also worsens the quality of learned features for downstream retrieval, so it is not free.
Class imbalance
For imbalance ratio r between the largest and smallest class, three standard treatments:
| Method | Mechanism | Failure mode |
|---|---|---|
| Re-weighting | Loss term scaled by 1 / n_c | Unstable at extreme r; noisy minority gradients dominate |
| Re-sampling | Oversample minority classes | Overfits the duplicated minority examples |
| Focal loss | (1 - p_t)^gamma factor | An extra hyperparameter; less effective for classification than detection |
Focal loss (Lin et al., 2017) down-weights well-classified examples:
FL(p_t) = -alpha_t * (1 - p_t)^gamma * log(p_t)p_tis the predicted probability of the true class.gamma >= 0sets the down-weighting strength;gamma = 2is standard.alpha_tis an optional per-class weight.
Cui et al. (2019) offer class-balanced weighting by effective number of samples, (1 - beta^n_c) / (1 - beta), which behaves better than 1/n_c at high imbalance.
Augmentation
Augmentation encodes invariances the architecture does not supply. It is the highest-return intervention on small datasets, and the specific transforms must match the domain.
- Geometric: random resized crop, horizontal flip, rotation. Horizontal flip is wrong for text and for digit recognition — a flipped 2 is not a 2.
- Photometric: colour jitter, grayscale, blur. Wrong wherever colour is the label, as in medical staining or fruit ripeness grading.
- Mixing: mixup (Zhang et al., 2018) interpolates image and label pairs; CutMix (Yun et al., 2019) pastes patches and mixes labels by area. Both improve calibration measurably.
- Learned policies: AutoAugment and RandAugment (Cubuk et al., 2019, 2020). RandAugment reduces the search to two hyperparameters and is the usual default.
Transfer learning
Pre-train on a large corpus, then adapt. Three regimes, chosen by dataset size:
| Target data | Approach | Typical setting |
|---|---|---|
| Under ~1k per class | Linear probe on frozen features | Backbone entirely frozen |
| 1k to 10k per class | Fine-tune, discriminative LRs | Later layers at higher LR |
| Over ~10k per class | Full fine-tune | Uniform LR, longer schedule |
Kornblith et al. (2019) found ImageNet accuracy correlates strongly with transfer accuracy across architectures — but the correlation weakens sharply when the target domain is far from natural images. Medical imaging is the standard counterexample: Raghu et al. (2019), Transfusion, showed ImageNet pre-training gives limited benefit for medical images. Much of the observed gain came from better weight scaling rather than from transferred features.
Parameter-efficient methods — LoRA, adapters, BitFit — now cover most fine-tuning at a fraction of the memory. See LoRA.
Open-set recognition
Softmax over K classes has no mechanism to abstain. p sums to 1 by construction, so an input from an unseen class is forced into the existing simplex.
Standard approaches:
- Maximum softmax probability thresholding (Hendrycks & Gimpel, 2017) — the baseline, and surprisingly hard to beat.
- ODIN — temperature scaling plus input perturbation.
- Mahalanobis distance in feature space, fitted per class.
- Energy-based scores,
-logsumexp(z), which use the unnormalised logits and outperform MSP consistently.
None of these is reliable enough to deploy unsupervised. Yang et al. (2022) survey the field, and the honest summary is that OOD detection performance on far-OOD data looks good and on near-OOD data remains poor. Near-OOD is what production actually encounters.
Calibration
Guo et al. (2017) established that modern networks are overconfident, and that the effect worsens with depth and with training beyond convergence.
Expected calibration error partitions predictions into M confidence bins:
ECE = sum over m = 1..M of (|B_m| / n) * | acc(B_m) - conf(B_m) |B_mis the set of predictions in binm.acc(B_m)is observed accuracy in that bin;conf(B_m)is mean predicted confidence.nis the total number of predictions.
Temperature scaling divides logits by a single scalar T, fitted on a validation set. It removes most of the miscalibration at zero cost to accuracy, since it does not change the argmax. Fit T, do not guess it.
Key references
- Krizhevsky, A., Sutskever, I. & Hinton, G. (2012). ImageNet Classification with Deep CNNs. NeurIPS 25.
- He, K. et al. (2016). Deep Residual Learning for Image Recognition. arXiv:1512.03385
- Szegedy, C. et al. (2016). Rethinking the Inception Architecture for Computer Vision. arXiv:1512.00567 — label smoothing.
- Guo, C. et al. (2017). On Calibration of Modern Neural Networks. arXiv:1706.04599
- Lin, T.-Y. et al. (2017). Focal Loss for Dense Object Detection. arXiv:1708.02002
- Zhang, H. et al. (2018). mixup: Beyond Empirical Risk Minimization. arXiv:1710.09412
- Kornblith, S., Shlens, J. & Le, Q. (2019). Do Better ImageNet Models Transfer Better? arXiv:1805.08974
- Raghu, M. et al. (2019). Transfusion: Understanding Transfer Learning for Medical Imaging. arXiv:1902.07208
- Recht, B. et al. (2019). Do ImageNet Classifiers Generalize to ImageNet? arXiv:1902.10811
- Beyer, L. et al. (2020). Are we done with ImageNet? arXiv:2006.07159
- Dosovitskiy, A. et al. (2021). An Image is Worth 16x16 Words. arXiv:2010.11929
- Radford, A. et al. (2021). Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.00020
Current state and open problems
Supervised ImageNet-1k classification is saturated as a research target. Top-1 above 90 percent is routine, and Beyer et al. showed the remaining headroom is substantially label noise rather than model error. Reporting further gains on that benchmark carries little information.
The practical frontier has moved to three places.
Zero-shot and open-vocabulary classification. CLIP-style models classify against a text-defined label set supplied at inference, removing the fixed-K constraint that defines classical classification. Accuracy trails a fine-tuned specialist on any specific task, and the flexibility usually wins in practice.
Label efficiency. Self-supervised backbones — DINOv2, MAE — produce features where a linear probe on a few hundred labels per class approaches full fine-tuning. This matters far more for real projects than another point on ImageNet.
Robustness and reliability. Distribution shift, calibration and open-set rejection are all unsolved and all directly determine whether a classifier can be deployed where a wrong answer has a cost. A system that is 99 percent accurate in-distribution and silently confident out of it is not ready. No benchmark in common use will tell you that.
What to learn next
- Object detection — localisation, anchors and mAP.
- Vision transformers — patch-based classification without convolution.
- Model evaluation — the metric machinery in full.