Computer Vision

Image classification

Image classification is giving a whole picture one label from a fixed list of choices, which is the first real task most vision projects start with.

On this page 7
  1. Why it exists
  2. How it works
  3. Classification is not the only vision task
  4. Where you have already seen it
  5. Be careful with the accuracy number
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Image classification is looking at a whole picture and giving it one label from a fixed list.

Walk through a mango mandi early in the morning. Workers pick up each mango, turn it once, and drop it into a crate. Grade one here, grade two there, reject in the third.

One fruit, one crate, decided in under a second. Watch closely and you notice something. There is no crate for guavas. If a guava came down the line, it would still land in a mango crate. Those are the only crates there are.

Hold on to that. It is the most important limitation of image classification, and it catches people out constantly.

Why it exists

Almost every useful vision question starts as a sorting question.

Is this X-ray normal or does it need a doctor's eyes today? Is this leaf healthy or infected? Is this photo a receipt or a selfie? Is this weld good or cracked?

Each one is a pile of pictures and a small set of crates. Get that working and you have solved a real problem for real people, with the simplest tool in computer vision.

It is also the doorway to everything else. Detection, segmentation and face recognition all reuse the machinery you build here.

How it works

The model never picks one answer directly. It gives a score to every label on the list, and the highest score wins.

                       +-------------------------+
                       |  score for every label  |
   photo   ->  model ->|                         |
                       |   dog     ####### <- highest, so this is the answer
                       |   cat     #
                       |   horse   #
                       |   cow     .
                       +-------------------------+

The scores always add up to a whole, like slicing one roti between the labels. So a big slice for "dog" leaves small slices for everything else.

That slicing has an odd side effect. Show it a photo of a guava, and the roti still gets sliced between dog, cat, horse and cow. Something must win. The model has no way to say "this is not on my list".

Classification is not the only vision task

People mix these up constantly, so here they are side by side.

   CLASSIFICATION    "there is a dog in this photo"        one label, whole image
   DETECTION         "there is a dog, here, in this box"   label plus location
   SEGMENTATION      "these exact dots are the dog"        a label for every dot

Start with classification. It needs the least labelling work by a wide margin, and it answers more questions than people expect.

Where you have already seen it

  • Google Photos albums. Every photo gets sorted into people, places and things.
  • Your spam folder for images. Screenshots and memes get separated from documents.
  • Plant and bird apps. Photograph it, get a species name and a confidence.
  • Bank apps. Photograph a cheque and it checks the picture is a cheque before reading it.
  • Factory quality checks. A camera above the belt sorts good parts from cracked ones.

Be careful with the accuracy number

You will train something and it will report 99 out of 100 correct. That feels like finishing. It usually is not.

If the crates are lopsided, accuracy lies. Say 99 of every 100 X-rays are normal. A model that answers "normal" every single time, with no thought at all, scores 99 out of 100. It also misses every sick patient. The number looks excellent and the system is worthless.

A test that resembles the real world is the only test that counts. Photos taken in your office, with your lighting, on your phone, are not the photos your users will send.

Always look at the pictures the model got wrong. Ten minutes of looking beats a week of tuning. This is the habit that separates people who ship working systems from people who ship good numbers.

Remember this

  • Classification gives one label per picture, chosen from a fixed list decided before training.
  • It cannot say "none of these". Anything you show it lands in some crate.
  • Accuracy alone is misleading whenever your classes are uneven. Look at what it got wrong.

What to learn next

Developer — Code and libraries.

You can train a working image classifier in fifteen lines, on a CPU, in about two seconds. Do that before touching a neural network.

scikit-learn ships a small handwritten-digit dataset inside the library, so there is nothing to download. That makes it perfect for learning the shape of the problem.

Setup

bash
pip install scikit-learn

No GPU, no dataset download, no network access at runtime.

First, look at your data

Never train on data you have not looked at. Here is what the model actually receives.

look.py
from sklearn.datasets import load_digits

digits = load_digits()          # ships inside scikit-learn, nothing to download
image = digits.images[0]        # one 8x8 grid of ink values, 0 to 16

for row in image:
    # heavier ink prints as a heavier character, so the shape becomes visible
    print("".join("#" if v > 8 else "+" if v > 2 else "." for v in row))

print("label:", digits.target[0])
print("dataset:", digits.images.shape)
Output
..+##...
..####+.
.+#..#+.
.+#..++.
.++..#+.
.+#..#+.
..#+##..
..+##...
label: 0
dataset: (1797, 8, 8)

A zero, drawn with 64 numbers. That is the entire input. No strokes, no order of writing, no pen pressure — 64 brightness values in a grid.

Train the classifier

classify.py
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC

X, y = load_digits(return_X_y=True)      # X is 1797 rows of 64 numbers each
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=0)

clf = SVC(gamma=0.001)                   # gamma sets how local each decision is
clf.fit(X_train, y_train)

predicted = clf.predict(X_test)
correct = (predicted == y_test).sum()

print("trained on:", X_train.shape[0], "images")
print("tested on: ", X_test.shape[0], "images")
print(f"correct:    {correct} out of {len(y_test)}")
print(f"accuracy:   {correct / len(y_test):.3f}")

for true, guess in zip(y_test[predicted != y_test], predicted[predicted != y_test]):
    print(f"  mistake: a {true} was read as a {guess}")
Output
trained on: 1347 images
tested on:  450 images
correct:    448 out of 450
accuracy:   0.996
  mistake: a 9 was read as a 5
  mistake: a 5 was read as a 9

Your counts should match exactly, because random_state=0 fixes the split and SVC has no randomness in fitting. If your scikit-learn version differs by a lot, the two mistakes could shift by one either way.

Line by line

random_state=0 fixes the shuffle before splitting. Without it, every run gives a different score and you cannot tell a real improvement from noise. Fix the seed, and report results across several seeds when it matters.

Fit on train, score on test. train_test_split holds back 450 images the model never sees during fitting. Score on data the model trained on and you measure memorisation, which is always near-perfect and always meaningless.

gamma=0.001 controls how far a single training example's influence reaches in an RBF kernel. Large gamma means each example only affects its immediate neighbourhood, which memorises. Small gamma means broad, smooth decision boundaries. This value is the standard one for this dataset and was found by search, not derived.

Printing the mistakes is the most valuable line in the file. Both errors are 5 against 9, in both directions. At 8 by 8 resolution, a 5 and a 9 differ by a handful of pixels in the upper curve. The model is not failing randomly — it is failing on the genuinely ambiguous pair, which is what a healthy model looks like.

Do not read too much into 0.996

This dataset is unusually kind, and it is worth naming why.

Every digit is centred, upright, cropped and drawn in the same style. There is no background, no lighting variation, no camera angle. Real photographs have every one of those problems.

The dataset is also perfectly balanced — around 180 examples of each of the ten digits — so accuracy is a fair measure here. It will not be fair on your data.

Treat this as learning the shape of a classification pipeline, not as evidence that image classification is a solved problem. It is not.

When to move up to a neural network

Straight lines through pixel values stop working the moment position matters. That is nearly always, for real photos.

The step up is a convolutional network. The practical route is transfer learning. Take a model already trained on millions of images, keep everything except the last layer, and train a new last layer on your own data.

That works with a few hundred examples per class, on a CPU if you are patient. Training from scratch needs tens of thousands per class. Almost nobody should train from scratch.

Common mistakes

Scaling mismatch. These digits are 0 to 16. Photographs are 0 to 255. Pretrained models expect their own specific normalisation. Feed the wrong range and accuracy collapses with no error raised. Apply exactly the same preprocessing at training and at inference.

Leaking data between the splits. Near-duplicate images — burst shots, frames from one video, the same object photographed twice — must land on the same side of the split. Otherwise the test set contains what the model already learned, and the score is fiction. Split by source, patient or session, never by row.

Reporting accuracy on uneven classes. With 95 percent normal cases, answering "normal" always scores 95 percent. Report per-class recall and a confusion matrix. sklearn.metrics.classification_report gives both in one line.

Forgetting the closed-set problem. Your model assigns every input to one of its known classes. Feed it a class it never saw and it answers confidently and wrongly. If unknown inputs are possible in production, add a rejection class or a confidence threshold, and validate that threshold on real unknowns.

Tuning against the test set. Try twenty settings, keep the best test score, and that score is now optimistic. Use a separate validation split for tuning and touch the test set once.

Try it yourself

Replace SVC(gamma=0.001) with SVC(), which uses the default gamma="scale", and re-run. Then look at the mistakes it prints, not only the accuracy.

After that, get the per-class report instead of a single number:

report.py
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC
from sklearn.metrics import classification_report

X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=0)

predicted = SVC(gamma=0.001).fit(X_train, y_train).predict(X_test)
print(classification_report(y_test, predicted))
Output
              precision    recall  f1-score   support

           0       1.00      1.00      1.00        37
           1       1.00      1.00      1.00        43
           2       1.00      1.00      1.00        44
           3       1.00      1.00      1.00        45
           4       1.00      1.00      1.00        38
           5       0.98      0.98      0.98        48
           6       1.00      1.00      1.00        52
           7       1.00      1.00      1.00        48
           8       1.00      1.00      1.00        48
           9       0.98      0.98      0.98        47

    accuracy                           1.00       450
   macro avg       1.00      1.00      1.00       450
weighted avg       1.00      1.00      1.00       450

Read down the recall column. Classes 5 and 9 sit at 0.98 and every other digit is perfect.

Notice also that the accuracy row prints 1.00 while two images were in fact wrong. It is rounding 0.996 up. That is exactly why the raw count of mistakes belongs next to any percentage you report.

The support column is the number of test images per class, and it is the first thing to check on your own data. Even support means accuracy is fair. Lopsided support means it is not.

What to learn next

Researcher — Mathematics and papers.

Formulation

Given input x in R^(H x W x C) and label y in {1, ..., K}, learn f_theta mapping images to a distribution over K classes.

  • H, W, C are height, width and channel count.
  • K is the number of classes, fixed before training.
  • theta is the parameter vector.

The network emits logits z = f_theta(x) in R^K, converted to probabilities by softmax:

text
p_i = exp(z_i) / sum over j = 1..K of exp(z_j)
  • z_i is the logit for class i.
  • p_i is the predicted probability of class i, with all p_i summing to 1.

Note that softmax is shift-invariant: adding a constant to every logit leaves p unchanged. Implementations subtract max(z) before exponentiating for numerical stability, which is why raw logits are not comparable across examples.

Cross-entropy loss

text
L = -(1/N) * sum over n = 1..N of sum over i = 1..K of  y_{n,i} * log(p_{n,i})
  • N is the batch size.
  • y_{n,i} is 1 if example n has class i, else 0 (one-hot).
  • For hard labels this collapses to L = -(1/N) * sum over n of log(p_{n, y_n}).

The gradient with respect to the logits is exactly dL/dz_i = p_i - y_i. This clean form is the reason softmax and cross-entropy are paired: the softmax Jacobian and the log cancel, leaving a difference of probabilities.

Label smoothing (Szegedy et al., 2016) replaces the one-hot target with y_smooth = (1 - eps) * y + eps / K, where eps is typically 0.1. It bounds the logit gap the model drives toward, improving calibration and usually top-1 by a few tenths of a point. It also worsens the quality of learned features for downstream retrieval, so it is not free.

Class imbalance

For imbalance ratio r between the largest and smallest class, three standard treatments:

MethodMechanismFailure mode
Re-weightingLoss term scaled by 1 / n_cUnstable at extreme r; noisy minority gradients dominate
Re-samplingOversample minority classesOverfits the duplicated minority examples
Focal loss(1 - p_t)^gamma factorAn extra hyperparameter; less effective for classification than detection

Focal loss (Lin et al., 2017) down-weights well-classified examples:

text
FL(p_t) = -alpha_t * (1 - p_t)^gamma * log(p_t)
  • p_t is the predicted probability of the true class.
  • gamma >= 0 sets the down-weighting strength; gamma = 2 is standard.
  • alpha_t is an optional per-class weight.

Cui et al. (2019) offer class-balanced weighting by effective number of samples, (1 - beta^n_c) / (1 - beta), which behaves better than 1/n_c at high imbalance.

Augmentation

Augmentation encodes invariances the architecture does not supply. It is the highest-return intervention on small datasets, and the specific transforms must match the domain.

  • Geometric: random resized crop, horizontal flip, rotation. Horizontal flip is wrong for text and for digit recognition — a flipped 2 is not a 2.
  • Photometric: colour jitter, grayscale, blur. Wrong wherever colour is the label, as in medical staining or fruit ripeness grading.
  • Mixing: mixup (Zhang et al., 2018) interpolates image and label pairs; CutMix (Yun et al., 2019) pastes patches and mixes labels by area. Both improve calibration measurably.
  • Learned policies: AutoAugment and RandAugment (Cubuk et al., 2019, 2020). RandAugment reduces the search to two hyperparameters and is the usual default.

Transfer learning

Pre-train on a large corpus, then adapt. Three regimes, chosen by dataset size:

Target dataApproachTypical setting
Under ~1k per classLinear probe on frozen featuresBackbone entirely frozen
1k to 10k per classFine-tune, discriminative LRsLater layers at higher LR
Over ~10k per classFull fine-tuneUniform LR, longer schedule

Kornblith et al. (2019) found ImageNet accuracy correlates strongly with transfer accuracy across architectures — but the correlation weakens sharply when the target domain is far from natural images. Medical imaging is the standard counterexample: Raghu et al. (2019), Transfusion, showed ImageNet pre-training gives limited benefit for medical images. Much of the observed gain came from better weight scaling rather than from transferred features.

Parameter-efficient methods — LoRA, adapters, BitFit — now cover most fine-tuning at a fraction of the memory. See LoRA.

Open-set recognition

Softmax over K classes has no mechanism to abstain. p sums to 1 by construction, so an input from an unseen class is forced into the existing simplex.

Standard approaches:

  • Maximum softmax probability thresholding (Hendrycks & Gimpel, 2017) — the baseline, and surprisingly hard to beat.
  • ODIN — temperature scaling plus input perturbation.
  • Mahalanobis distance in feature space, fitted per class.
  • Energy-based scores, -logsumexp(z), which use the unnormalised logits and outperform MSP consistently.

None of these is reliable enough to deploy unsupervised. Yang et al. (2022) survey the field, and the honest summary is that OOD detection performance on far-OOD data looks good and on near-OOD data remains poor. Near-OOD is what production actually encounters.

Calibration

Guo et al. (2017) established that modern networks are overconfident, and that the effect worsens with depth and with training beyond convergence.

Expected calibration error partitions predictions into M confidence bins:

text
ECE = sum over m = 1..M of (|B_m| / n) * | acc(B_m) - conf(B_m) |
  • B_m is the set of predictions in bin m.
  • acc(B_m) is observed accuracy in that bin; conf(B_m) is mean predicted confidence.
  • n is the total number of predictions.

Temperature scaling divides logits by a single scalar T, fitted on a validation set. It removes most of the miscalibration at zero cost to accuracy, since it does not change the argmax. Fit T, do not guess it.

Key references

  • Krizhevsky, A., Sutskever, I. & Hinton, G. (2012). ImageNet Classification with Deep CNNs. NeurIPS 25.
  • He, K. et al. (2016). Deep Residual Learning for Image Recognition. arXiv:1512.03385
  • Szegedy, C. et al. (2016). Rethinking the Inception Architecture for Computer Vision. arXiv:1512.00567 — label smoothing.
  • Guo, C. et al. (2017). On Calibration of Modern Neural Networks. arXiv:1706.04599
  • Lin, T.-Y. et al. (2017). Focal Loss for Dense Object Detection. arXiv:1708.02002
  • Zhang, H. et al. (2018). mixup: Beyond Empirical Risk Minimization. arXiv:1710.09412
  • Kornblith, S., Shlens, J. & Le, Q. (2019). Do Better ImageNet Models Transfer Better? arXiv:1805.08974
  • Raghu, M. et al. (2019). Transfusion: Understanding Transfer Learning for Medical Imaging. arXiv:1902.07208
  • Recht, B. et al. (2019). Do ImageNet Classifiers Generalize to ImageNet? arXiv:1902.10811
  • Beyer, L. et al. (2020). Are we done with ImageNet? arXiv:2006.07159
  • Dosovitskiy, A. et al. (2021). An Image is Worth 16x16 Words. arXiv:2010.11929
  • Radford, A. et al. (2021). Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.00020

Current state and open problems

Supervised ImageNet-1k classification is saturated as a research target. Top-1 above 90 percent is routine, and Beyer et al. showed the remaining headroom is substantially label noise rather than model error. Reporting further gains on that benchmark carries little information.

The practical frontier has moved to three places.

Zero-shot and open-vocabulary classification. CLIP-style models classify against a text-defined label set supplied at inference, removing the fixed-K constraint that defines classical classification. Accuracy trails a fine-tuned specialist on any specific task, and the flexibility usually wins in practice.

Label efficiency. Self-supervised backbones — DINOv2, MAE — produce features where a linear probe on a few hundred labels per class approaches full fine-tuning. This matters far more for real projects than another point on ImageNet.

Robustness and reliability. Distribution shift, calibration and open-set rejection are all unsolved and all directly determine whether a classifier can be deployed where a wrong answer has a cost. A system that is 99 percent accurate in-distribution and silently confident out of it is not ready. No benchmark in common use will tell you that.

What to learn next