Build a handwritten digit recogniser
Train a small neural network to read handwritten digits from 1,797 images that ship inside scikit-learn, with no download and no GPU.
- 19 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A digit recogniser looks at a picture of a handwritten number and tells you which number it is.
The analogy you have already lived
Somebody scribbles a phone number on a scrap of paper and hands it to you. You squint at one character. Is that a one or a seven? A four or a nine?
You decide anyway. You compare the shape against the thousands of handwritten digits you have seen since you were five years old, and you pick the closest match.
You are not following a rule about what a four looks like. You are matching against memory. That is exactly what this project builds.
Why this project exists
Every day, machines must read numbers a human wrote by hand. Bank cheques. Postal PIN codes on envelopes. Exam roll numbers on answer sheets. Meter readings.
Writing rules for this is hopeless. Your four has a closed top; mine has an open one. A doctor's four looks like nothing on earth. There is no rulebook that covers every human hand.
So this became one of the first real successes of the field. In the 1990s, systems reading handwritten digits were sorting post and clearing cheques at scale, long before anyone said "AI" in an advertisement.
It is the standard first project in computer vision for a good reason. The problem is easy to state, the data is small, and a working model is genuinely useful.
What you are actually going to build
The pictures are tiny — eight squares across and eight squares down. Each square holds one shade of grey, from blank white to solid black.
You get 1,797 of these pictures. Every one comes with the correct answer attached.
You show the model most of them, keeping some hidden. Then you test it on the hidden ones, which it has never seen. That hidden set is the only honest way to know whether it learned anything.
How it works
handwritten digit, photographed
|
v
[ chopped into 64 tiny squares ]
each square = one shade of grey
|
v
[ the model ]
|
v
"this is a 4" (98% sure)The model never sees a "picture" the way you do. It sees sixty-four numbers in a fixed order. Every image, no matter whose handwriting, becomes sixty-four numbers.
You will print these images as text on your screen, using symbols for dark and light. That way you see exactly what the model sees.
Where you have already seen this
- Cheque clearing in banks, reading the handwritten amount.
- India Post sorting machines, reading the six-digit PIN code.
- Google Lens, reading a number off a bill through your camera.
- Exam answer sheets where a machine reads your roll number.
The honest part
Your model will get about ninety-seven out of every hundred right. That sounds excellent. Think about what it means for a bank.
Three cheques in every hundred read wrong. On a million cheques a day, that is thirty thousand errors. Nobody ships that.
Real systems handle this by refusing to answer. When the model is unsure, it hands the image to a human instead of guessing. You will see this in action, because the mistakes your model makes are mostly ones where it was hesitant.
The second honest thing: these images are small and clean and centred. A photo from your phone is none of those. Cropping and straightening real handwriting is often harder than the recognition itself.
Remember this
- An image is nothing but numbers — one number per tiny square of grey.
- The model learns by comparing against thousands of labelled examples, not by following shape rules.
- Ninety-seven percent sounds great and is not good enough for money or post — knowing when to say "I am not sure" matters more.
What to learn next
- Convolutional neural networks — the architecture built for images, and the next version of this project.
- Image classification — the same task on photographs instead of 8x8 grids.
- Model evaluation — confusion matrices, precision and recall in depth.
Developer — Code and libraries.
The problem, stated precisely
Given a grid of pixel values, output which of ten digits it shows. This is multi-class classification: one input, ten mutually exclusive labels.
Our plan:
- Load 1,797 labelled 8x8 grayscale images that ship inside scikit-learn.
- Print a few as text, so you can see the actual input.
- Train a small neural network, measure it, and look at every mistake it makes.
- Race five different algorithms and see which wins.
Setup
pip install scikit-learn numpyNothing is downloaded. The dataset is a file inside the scikit-learn package you already installed. The whole script trains in a few seconds on CPU.
The dataset
load_digits() gives you 1,797 images of handwritten digits, each 8 by 8 pixels, each pixel an integer from 0 to 16. It comes from the UCI Optical Recognition of Handwritten Digits collection, contributed by Alpaydin and Kaynak in 1998.
Those originals were 32x32 black-and-white bitmaps. Each was cut into non-overlapping 4x4 blocks and the dark pixels in each block counted, giving an 8x8 grid of counts from 0 to 16. That is why the values stop at 16, which puzzles people the first time.
It is a smaller cousin of MNIST, the famous 28x28 dataset of 70,000 digits. MNIST is an 11 MB download and needs a minute of training. This one is instant, and everything you learn transfers directly.
The full build
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import MinMaxScaler
from sklearn.neural_network import MLPClassifier
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix
digits = load_digits() # ships inside scikit-learn, nothing is downloaded
X, y = digits.data, digits.target
print("images:", digits.images.shape, " flattened:", X.shape)
print("labels present:", np.unique(y))
print("pixel range:", X.min(), "to", X.max())
def show(image, label):
"""Print an 8x8 image as text, darkest pixels as '@'."""
ramp = " .:-=+*#%@"
for row in image:
print("".join(ramp[int(p) * (len(ramp) - 1) // 16] for p in row))
print("label:", label, "\n")
show(digits.images[0], y[0])
show(digits.images[8], y[8])
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=0, stratify=y
)
# pixels are 0-16; neural networks train far better on 0-1 inputs
scaler = MinMaxScaler()
X_train_s = scaler.fit_transform(X_train)
X_test_s = scaler.transform(X_test)
model = MLPClassifier(hidden_layer_sizes=(64,), max_iter=600, random_state=0)
model.fit(X_train_s, y_train)
pred = model.predict(X_test_s)
print("training examples:", X_train.shape[0], " test examples:", X_test.shape[0])
print("test accuracy: %.4f" % accuracy_score(y_test, pred))
print()
print(classification_report(y_test, pred, digits=3))images: (1797, 8, 8) flattened: (1797, 64)
labels present: [0 1 2 3 4 5 6 7 8 9]
pixel range: 0.0 to 16.0
:#+
#%+%:
.%. *=
:* ==
:= +=
:* *-
.#:+*
-#+
label: 0
+#=
*##*
++ %:
.@*#.
:@@.
.@=+#.
% .@=
*@%*
label: 8
training examples: 1347 test examples: 450
test accuracy: 0.9733
precision recall f1-score support
0 1.000 1.000 1.000 45
1 0.918 0.978 0.947 46
2 1.000 0.977 0.989 44
3 0.939 1.000 0.968 46
4 1.000 0.978 0.989 45
5 0.957 0.978 0.968 46
6 1.000 0.978 0.989 45
7 1.000 0.978 0.989 45
8 0.952 0.930 0.941 43
9 0.977 0.933 0.955 45
accuracy 0.973 450
macro avg 0.974 0.973 0.973 450
weighted avg 0.974 0.973 0.973 450Look at those two text pictures. A zero with a hole in the middle, and an eight with two loops. Eight by eight is barely enough resolution for a human, and the model works from exactly the same sixty-four numbers.
Now look at every mistake
Accuracy is one number covering 450 decisions. The interesting information is in the twelve it got wrong.
print("confusion matrix (row = true digit, column = predicted digit):")
print(confusion_matrix(y_test, pred))
print()
wrong = np.where(pred != y_test)[0]
print("mistakes:", len(wrong), "out of", len(y_test))
print()
for i in wrong[:3]:
conf = model.predict_proba(X_test_s[i].reshape(1, -1)).max()
show(X_test[i].reshape(8, 8), f"true {y_test[i]}, guessed {pred[i]} ({conf:.0%} sure)")confusion matrix (row = true digit, column = predicted digit):
[[45 0 0 0 0 0 0 0 0 0]
[ 0 45 0 0 0 0 0 0 1 0]
[ 0 1 43 0 0 0 0 0 0 0]
[ 0 0 0 46 0 0 0 0 0 0]
[ 0 0 0 0 44 0 0 0 1 0]
[ 0 0 0 1 0 45 0 0 0 0]
[ 0 1 0 0 0 0 44 0 0 0]
[ 0 0 0 0 0 0 0 44 0 1]
[ 0 2 0 0 0 1 0 0 40 0]
[ 0 0 0 2 0 1 0 0 0 42]]
mistakes: 12 out of 450
.:+
.%=#.
=- +-
.*=%*
-.#.
-+
+- @
.*@@.
label: true 9, guessed 3 (55% sure)
*-
:@..#.
+% *%
-@%@:
-%*
@:
+*
#+
label: true 4, guessed 8 (77% sure)
#+==-
:@@@@%.
+@:
#*
-%
=%
-+*
#@:
label: true 5, guessed 3 (52% sure)Read the confusion matrix — it is the most useful table in classification
Every row is a true digit; every column is what the model said. The diagonal is correct answers. Anything off the diagonal is a mistake, and its position tells you exactly which mistake.
The digit 8 was called 1 twice, and 9 was called 3 twice. These are not random errors. In eight-by-eight resolution, a narrow 8 loses its loops and becomes a vertical stroke. A 9 with a flat tail genuinely does resemble a 3.
Print the three failures as text and you can see it yourself. That first one, labelled 9, is barely a 9 to human eyes either.
Two of the three mistakes came with low confidence: 55% and 52%. The model was hesitating, and it told you so. That is directly usable:
probs = model.predict_proba(X_test_s)
confident = probs.max(axis=1) > 0.90
print("kept:", confident.sum(), "of", len(y_test))
print("accuracy on kept:", accuracy_score(y_test[confident], pred[confident]))No output block for that snippet, because the exact split depends on the threshold you pick — run it and see. The pattern matters more than the number: trading coverage for accuracy by refusing uncertain cases is how real cheque readers work. They send the hesitant ones to a human.
The middle mistake is the troubling one. A true 4 called an 8 at 77% confidence. Confident and wrong. Neural network probabilities are not honest probabilities out of the box, which is a real limitation covered in the researcher block.
Race five algorithms
A neural network is not automatically the right tool. Find out.
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import MinMaxScaler
from sklearn.neural_network import MLPClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.svm import SVC
from sklearn.ensemble import RandomForestClassifier
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score
digits = load_digits()
X, y = digits.data, digits.target
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=0, stratify=y)
sc = MinMaxScaler().fit(X_train)
Xtr, Xte = sc.transform(X_train), sc.transform(X_test)
models = {
"always most common (baseline)": DummyClassifier(strategy="most_frequent"),
"logistic regression": LogisticRegression(max_iter=2000),
"random forest": RandomForestClassifier(n_estimators=200, random_state=0),
"neural net (64 hidden)": MLPClassifier(hidden_layer_sizes=(64,), max_iter=600, random_state=0),
"support vector machine": SVC(gamma="scale"),
}
for name, m in models.items():
m.fit(Xtr, y_train)
print(f"{name:32} {accuracy_score(y_test, m.predict(Xte)):.4f}")always most common (baseline) 0.1022 logistic regression 0.9689 random forest 0.9867 neural net (64 hidden) 0.9733 support vector machine 0.9889
The neural network came fourth. A support vector machine, an algorithm from 1995 with no hidden layers at all, beat it. So did a random forest.
This is worth sitting with. On 1,347 training images of 64 pixels each, a small dataset with clean, centred, low-resolution inputs, classical methods are excellent. Neural networks pull decisively ahead when data grows into the tens of thousands and inputs get messier. Choose by measuring, not by fashion.
Also note the baseline: 0.1022. Always guessing the most common digit gets you about ten percent, because there are ten roughly equal classes. Every accuracy number needs that reference point beside it. Always fit a DummyClassifier first.
Common mistakes
Skipping the scaler. Feeding raw 0-16 values to MLPClassifier still works here, but converges slower and less reliably. On real image data with a 0-255 range, unscaled inputs cause saturated activations and a stalled loss. Scale before any neural network, always.
Calling fit_transform on the test set. scaler.fit_transform(X_test) computes new minimum and maximum values from the test data. That is data leakage: information from the graded set flows into preprocessing. Fit on train, transform both.
Reshaping wrong. X is flat, shape (n, 64). digits.images is (n, 8, 8). Mixing them gives ValueError: cannot reshape array of size 64 into shape (8,8,3) or, worse, a silently transposed image. Print .shape whenever an image changes form.
Ignoring max_iter warnings. ConvergenceWarning: Stochastic Optimizer: Maximum iterations reached means training stopped early, not that it finished. Raise max_iter or the model is underfitted.
Trusting one split. A single random_state=0 split gives one sample of the accuracy. Use cross_val_score(model, X, y, cv=5) before you believe any difference under about one percent.
How to make this genuinely good
- Move to MNIST.
fetch_openml("mnist_784", version=1, as_frame=False)— 70,000 images at 28x28, about 11 MB, still CPU-trainable in minutes. Everything above works unchanged. - Augment the training data. Shift each image by one pixel in four directions and add the copies to the training set. This roughly quadruples your data for free and reliably improves accuracy, because it teaches the model that position should not change the answer.
- Use a convolutional network. A CNN shares its filters across the whole image, so it does not have to learn "an edge here" and "an edge two pixels left" separately. This is the single biggest architectural win on images.
- Calibrate the probabilities. Wrap the model in
CalibratedClassifierCVso that "77% sure" means it is right about 77% of the time. - Add a reject option. Return
Nonebelow your confidence threshold and route those to a human. Measure coverage and accuracy-on-covered as two separate numbers. - Test on your own handwriting. Write digits on paper, photograph them, crop to squares, convert to grayscale, resize to 8x8, invert if needed so ink is high-valued.
That last step will humble the model, and it should. It is the honest test.
Try it yourself
Change hidden_layer_sizes=(64,) to (8,) and re-run. Then try (256, 128).
Note what happens to accuracy in each case. The biggest network does not automatically win.
Then look at which digits the tiny network confuses. A model with too little capacity does not fail evenly. It collapses the pairs that look most alike first.
What to learn next
- Convolutional neural networks — the architecture built for images, and the next version of this project.
- Image classification — the same task on photographs instead of 8x8 grids.
- Model evaluation — confusion matrices, precision and recall in depth.
Researcher — Mathematics and papers.
Formulas here are written in plain text, since the site renders no maths typesetting library.
The classification head
For K classes, the network outputs a vector of logits z in R^K, converted to a distribution by softmax:
softmax(z)_k = exp(z_k - max_j z_j) / sum_i exp(z_i - max_j z_j)The subtraction of max_j z_j is mathematically a no-op and numerically essential: without it, exp overflows for logits above roughly 709 in float64. Every production implementation does this.
Training minimises categorical cross-entropy over N examples:
L = -(1/N) * sum_n sum_k y_nk * log( p_nk )y_nkis 1 if examplenhas true classk, else 0 (one-hot).p_nkis the softmax probability the model assigns examplento classk.
With one-hot targets this collapses to -(1/N) * sum_n log p_n,c(n), the negative log-likelihood of the correct class. The gradient with respect to the logits is exactly p - y, which is why softmax and cross-entropy are always implemented as a fused operation.
Why an MLP is the wrong prior for images
The model above flattens an 8x8 grid into a 64-vector. That discards the spatial arrangement entirely: permute the 64 pixels with any fixed permutation, retrain, and you get the same accuracy. The network never knew the pixels were adjacent.
Convolutional networks (LeCun et al., 1989; LeCun et al., 1998) impose two structural priors:
- Translation equivariance. A convolution commutes with translation, so a feature detector learned at one location applies at every location. An MLP must learn each location separately.
- Locality and weight sharing. A
k x kkernel overC_ininput channels andC_outoutput channels hask^2 * C_in * C_out + C_outparameters regardless of image size. A dense layer over aH x Wimage scales withH * W.
Concretely, for 28x28 MNIST: a dense layer to 128 units costs 784 * 128 + 128 = 100,480 parameters. A 3x3 convolution to 32 channels costs 9 * 1 * 32 + 32 = 320 parameters and produces a 32-channel feature map. Two orders of magnitude fewer parameters, and better accuracy.
Note that convolution gives equivariance, not invariance. Pooling and augmentation supply approximate invariance. This distinction is the subject of Cohen and Welling (2016) on group-equivariant networks.
Dataset provenance and its limits
load_digits is the Optical Recognition of Handwritten Digits set (Alpaydin and Kaynak, 1998), 1,797 samples in scikit-learn's copy. Preprocessing: 32x32 binary bitmaps, partitioned into non-overlapping 4x4 blocks, with the count of on-pixels per block forming an 8x8 integer matrix in [0, 16]. This block-count step is itself a hand-designed dimensionality reduction, and part of why simple classifiers do so well.
MNIST (LeCun, Cortes and Burges, 1998) is derived from NIST Special Database 3 and 1, size-normalised to a 20x20 box and centred by centre of mass in a 28x28 field. State of the art sits around 99.8% test accuracy, which means the benchmark carries roughly 20 informative examples. It is saturated and should not be used to compare modern methods.
Recht et al. (2018, 2019) constructed fresh test sets for CIFAR-10 and ImageNet following the original collection protocols, and observed accuracy drops of 3 to 15 points across a wide range of models, with the drop increasing for higher-scoring models. The ranking of models was largely preserved, so this is distribution shift from protocol details rather than pure adaptive overfitting — but it bounds how much any absolute number on a reused test set should be trusted.
Calibration
The 77%-confident error in the developer block is not an anomaly. Guo et al. (2017), On Calibration of Modern Neural Networks, show that deeper and wider networks are systematically overconfident, and that this worsened as architectures grew, even as accuracy improved.
Expected calibration error partitions predictions into M confidence bins and measures the gap between confidence and accuracy:
ECE = sum_{m=1..M} ( |B_m| / n ) * | acc(B_m) - conf(B_m) |B_m is the set of predictions whose confidence falls in bin m, acc(B_m) their empirical accuracy, conf(B_m) their mean predicted confidence.
Temperature scaling — dividing logits by a single scalar T fitted on a validation set by minimising NLL — repairs most of the gap at essentially zero cost and without changing the argmax, so accuracy is unaffected. It should be the default final step for any classifier whose probabilities are consumed downstream.
For a reject option, the decision-theoretic result is Chow's rule (Chow, 1970): under a known cost d for rejection and 0/1 loss for errors, the optimal policy rejects when max_k p(k|x) < 1 - d. This is optimal only for calibrated posteriors, which closes the loop back to why calibration matters.
Cost
For an MLP with layer widths [d_0, ..., d_L] and batch size B:
- Parameters:
sum_l ( d_l * d_{l-1} + d_l ). - Forward FLOPs: about
2 * B * sum_l d_l * d_{l-1}. - Backward: roughly twice the forward cost.
For the (64, 64, 10) network above: 64*64 + 64 = 4,160 plus 64*10 + 10 = 650, giving 4,810 parameters. That is small enough to run on a microcontroller, which is precisely why digit recognition was deployable in the 1990s on hardware weaker than a modern washing machine.
For a convolution producing H_out x W_out x C_out from C_in channels with a k x k kernel, FLOPs are about 2 * H_out * W_out * C_out * C_in * k^2 per image. Compute scales with spatial size while parameters do not, which is the opposite of the dense case and drives entirely different engineering trade-offs.
What production handwriting recognition looks like
Single-digit classification is a solved subproblem, not the product. A cheque or address reader is a pipeline:
- Detection and layout analysis — locate text regions, correct skew and perspective.
- Binarisation and normalisation — adaptive thresholding (Sauvola or Niblack) under uneven lighting.
- Sequence recognition — not per-character classification. Segmentation-free CTC (Graves et al., 2006) or attention-based encoder-decoders read a whole line, because character boundaries in cursive handwriting are ill-defined. This "Sayre's paradox" — you cannot segment without recognising, nor recognise without segmenting — is why per-character pipelines were abandoned.
- Language modelling — a lexicon or n-gram prior over valid PIN codes, amounts or names, which corrects a large fraction of visual errors.
- Rejection and human routing, with cost-sensitive thresholds set per field.
The recognition network is often the least troublesome component. Steps 1, 2 and 4 usually determine whether the system works.
Papers
- LeCun, Boser, Denker et al., Backpropagation Applied to Handwritten Zip Code Recognition, Neural Computation, 1989.
- LeCun, Bottou, Bengio and Haffner, Gradient-Based Learning Applied to Document Recognition, Proc. IEEE, 1998 — introduces LeNet-5 and MNIST.
- Alpaydin and Kaynak, Optical Recognition of Handwritten Digits, UCI Machine Learning Repository, 1998 — the dataset used here.
- Chow, On Optimum Recognition Error and Reject Tradeoff, IEEE Trans. Information Theory, 1970.
- Graves, Fernández, Gomez and Schmidhuber, Connectionist Temporal Classification, ICML 2006.
- Cohen and Welling, Group Equivariant Convolutional Networks, ICML 2016 — arxiv.org/abs/1602.07576
- Guo, Pleiss, Sun and Weinberger, On Calibration of Modern Neural Networks, ICML 2017 — arxiv.org/abs/1706.04599
- Recht, Roelofs, Schmidt and Shankar, Do ImageNet Classifiers Generalize to ImageNet?, ICML 2019 — arxiv.org/abs/1902.10811
What to learn next
- Convolutional neural networks — the architecture built for images, and the next version of this project.
- Image classification — the same task on photographs instead of 8x8 grids.
- Model evaluation — confusion matrices, precision and recall in depth.