Handwriting recognition
Handwriting is hard because every writer has a different alphabet, letters connect, and the same person writes differently twice — and the fix is mostly data, not architecture.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Handwriting recognition reads writing done by hand, where every writer draws letters differently.
The analogy you have already lived
Think of your school notebooks being marked by a teacher. She read forty different handwritings a night, including the boy whose a looked like an o.
She managed because she had seen thousands of pages. A new teacher, on her first week, struggled with the same notebooks.
That gap is exactly the gap between a model trained on printed text and one trained on handwriting. Nothing about the eyes changed. Only the practice.
Why it is harder than printed text
Printed text is a solved problem because it is repetitive. Handwriting breaks every property that made it solvable.
Every writer has a private alphabet. There is no single k shape, there are millions.
Letters connect. In cursive, one continuous stroke covers a whole word. There is nowhere to cut.
The same person is inconsistent. Write your name five times. The five are not identical.
Baselines drift. Words sag, rise, and slant differently down a page.
Spacing is unreliable. Some writers leave no gap between words, and big gaps inside them.
What actually fixes it
Two things, in this order.
Read the whole line, never cut into letters. This is the segmentation-free idea from CRNN text recognition, and it is not optional here. Cutting cursive is impossible, so a model that must cut cannot work.
Train on handwriting, with heavy variation added. This is the part people underestimate. A perfectly good architecture trained on clean print will fail on handwriting, and the failure looks like a model problem. It is a data problem.
The code below shows exactly that, measured. One network, two training sets. Trained on neat print it reads neat print perfectly and messy writing almost never. The identical network trained on messy writing reads both.
How the shapes are handled
printed: R E N T letters separated, upright, same size
|
handwritten: RENT slanted, touching, wobbling on the baseline
|
the reader never asks "where does R end?"
it sweeps left to right and reports what it seesWhere you have already seen it
- Post offices sorting letters by reading handwritten PIN codes.
- Bank cheque readers reading the amount written in words.
- Note apps that convert your scribbles into typed text.
- Exam and survey forms filled in by hand and processed in bulk.
The honest part
Handwriting recognition is still genuinely unsolved for free-form writing.
Constrained handwriting works well. A box per character, a known alphabet, a known field type — those systems are reliable and are deployed everywhere. Free-flowing cursive on a lined page is much harder, and the best systems still make errors a human would not.
Anyone selling you a general handwriting reader with 99% accuracy is describing a narrow test set.
Remember this
- Handwriting is hard because there is no fixed alphabet and letters connect.
- The reader must never try to cut a word into letters.
- Most of the accuracy comes from training data variety, not from a better network.
What to learn next
- Document layout analysis — finding the lines before reading them.
- Data augmentation — the lever this lesson measured.
- CTC loss — why the failed reads came back empty rather than wrong.
Developer — Code and libraries.
Setup
pip install torch==2.5.1 opencv-python==4.10.0.84 numpy==1.26.4Real handwriting datasets need a download and a licence click. The point being made here does not, and it is worth making cleanly: the same architecture succeeds or fails purely on what it was shown.
We generate two synthetic corpora. One is neat print. One is slanted, touching, wobbling and variable-stroke — the properties that make handwriting hard, without the download. Then we train the identical network on each and cross-test.
This takes about ninety seconds on a laptop CPU.
Same network, two training sets
import cv2
import numpy as np
import torch
import torch.nn as nn
FONT = {
"A": ["..#..", ".#.#.", "#...#", "#...#", "#####", "#...#", "#...#"],
"C": [".###.", "#...#", "#....", "#....", "#....", "#...#", ".###."],
"E": ["#####", "#....", "#....", "####.", "#....", "#....", "#####"],
"L": ["#....", "#....", "#....", "#....", "#....", "#....", "#####"],
"N": ["#...#", "##..#", "#.#.#", "#.#.#", "#..##", "#...#", "#...#"],
"R": ["####.", "#...#", "#...#", "####.", "#.#..", "#..#.", "#...#"],
"T": ["#####", "..#..", "..#..", "..#..", "..#..", "..#..", "..#.."],
"1": ["..#..", ".##..", "..#..", "..#..", "..#..", "..#..", ".###."],
"9": [".###.", "#...#", "#...#", ".####", "....#", "#...#", ".###."],
}
ALPHABET = "".join(sorted(FONT))
GLYPH = {c: np.array([[x == "#" for x in r] for r in FONT[c]], np.float32) for c in FONT}
IMG_H, IMG_W, SCALE = 24, 96, 2
def draw(word, rng, messy):
"""messy=0 is neat print. Higher values slant, overlap and wobble the letters."""
canvas = np.zeros((IMG_H, IMG_W), np.float32)
x = 4 + int(rng.integers(0, 4))
for ch in word:
g = np.kron(GLYPH[ch], np.ones((SCALE, SCALE), np.float32))
h, w = g.shape
if messy:
thick = int(rng.integers(0, 2)) # pen pressure varies
if thick:
g = cv2.dilate(g, np.ones((2, 2), np.float32))
top = (IMG_H - h) // 2 + int(rng.integers(-messy, messy + 1))
top = max(0, min(IMG_H - h, top))
if x + w > IMG_W:
break
canvas[top:top + h, x:x + w] = np.maximum(canvas[top:top + h, x:x + w], g)
gap = 2 - int(rng.integers(0, messy + 1)) # letters start to touch
x += w + max(gap, -2)
if messy:
slant = float(rng.uniform(-0.35, 0.35)) * (messy / 3.0) # a slanted hand
M = np.array([[1, slant, -slant * IMG_H / 2], [0, 1, 0]], np.float32)
canvas = cv2.warpAffine(canvas, M, (IMG_W, IMG_H))
return np.clip(canvas, 0, 1)
def batch(n, rng, messy, lo=3, hi=6):
words = ["".join(rng.choice(list(ALPHABET), rng.integers(lo, hi + 1))) for _ in range(n)]
imgs = torch.tensor(np.stack([draw(w, rng, messy) for w in words]))[:, None]
targets = torch.tensor([ALPHABET.index(c) + 1 for w in words for c in w])
lengths = torch.tensor([len(w) for w in words])
return words, imgs, targets, lengths
class CRNN(nn.Module):
def __init__(self, n_classes):
super().__init__()
self.cnn = nn.Sequential(
nn.Conv2d(1, 16, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
nn.Conv2d(16, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d((2, 1)),
nn.Conv2d(32, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d((2, 1)),
)
self.rnn = nn.LSTM(32 * 3, 64, bidirectional=True)
self.head = nn.Linear(128, n_classes)
def forward(self, x):
f = self.cnn(x).permute(3, 0, 1, 2).flatten(2)
return self.head(self.rnn(f)[0])
def decode(logits):
out = []
for row in logits.argmax(2).T.tolist():
chars, prev = [], -1
for k in row:
if k != prev and k != 0:
chars.append(ALPHABET[k - 1])
prev = k
out.append("".join(chars))
return out
def train(messy, steps=900):
torch.manual_seed(0)
rng = np.random.default_rng(1)
model = CRNN(len(ALPHABET) + 1)
loss_fn = nn.CTCLoss(blank=0, zero_infinity=True)
opt = torch.optim.Adam(model.parameters(), lr=3e-3)
for step in range(steps):
_, imgs, targets, lengths = batch(16, rng, messy)
logits = model(imgs)
lens = torch.full((16,), logits.shape[0], dtype=torch.long)
loss = loss_fn(logits.log_softmax(2), targets, lens, lengths)
opt.zero_grad()
loss.backward()
opt.step()
model.eval()
return model, loss.item()
def evaluate(model, messy, n=200):
rng = np.random.default_rng(99)
words, imgs, _, _ = batch(n, rng, messy)
with torch.no_grad():
read = decode(model(imgs))
exact = sum(a == b for a, b in zip(words, read)) / n
return exact, list(zip(words, read))[:4]
print("the same word, printed neatly and written messily:")
for label, messy in [("print", 0), ("messy", 3)]:
img = draw("RENT", np.random.default_rng(7), messy)
print(f" {label}:")
for row in img[3:22:2]:
print(" " + "".join("#" if v > 0.5 else "." for v in row[:62]))
print("\ntraining a reader on NEAT PRINT only ...")
printed_model, loss_p = train(0)
print(f" final training loss {loss_p:.3f}")
for messy in [0, 1, 2, 3]:
acc, sample = evaluate(printed_model, messy)
print(f" tested at messiness {messy}: exact-word accuracy {acc:.3f} e.g. {sample[0]}")
print("\ntraining the SAME network on messy writing instead ...")
messy_model, loss_m = train(3)
print(f" final training loss {loss_m:.3f}")
for messy in [0, 3]:
acc, sample = evaluate(messy_model, messy)
print(f" tested at messiness {messy}: exact-word accuracy {acc:.3f} e.g. {sample[0]}")the same word, printed neatly and written messily:
print:
..............................................................
.......########....##########..##......##..##########.........
.......##......##..##..........####....##......##.............
.......##......##..##..........##..##..##......##.............
.......########....########....##..##..##......##.............
.......##..##......##..........##....####......##.............
.......##....##....##..........##......##......##.............
.......##......##..##########..##......##......##.............
..............................................................
..............................................................
messy:
.......................##......##.............................
.......................####....##.............................
......####################..##..##............................
......###.....############..##..##.##########.................
.......###.....####......##....####.....##....................
.......####################......##.....##....................
........###.###..###########......##.....##...................
........###...######.....................##...................
.........###.....###########..............##..................
..........................................##..................
training a reader on NEAT PRINT only ...
final training loss 0.040
tested at messiness 0: exact-word accuracy 1.000 e.g. ('ENL9ET', 'ENL9ET')
tested at messiness 1: exact-word accuracy 0.085 e.g. ('ENL9ET', '1N9T')
tested at messiness 2: exact-word accuracy 0.015 e.g. ('ENL9ET', 'AL9RT')
tested at messiness 3: exact-word accuracy 0.010 e.g. ('ENL9ET', 'T')
training the SAME network on messy writing instead ...
final training loss 0.093
tested at messiness 0: exact-word accuracy 0.960 e.g. ('ENL9ET', 'ENL9ET')
tested at messiness 3: exact-word accuracy 0.960 e.g. ('ENL9ET', 'ENL9ET')Reading the output carefully
Look at the ASCII pictures first. In the messy version the R and E have merged into one blob, the N sits high, the T sits low, and the whole word leans. Nothing about it is unreadable to you, and everything about it is different from the print sample.
1.000 down to 0.085 at messiness 1. One notch of wobble and slant costs the print-trained reader over ninety per cent of its exact-word accuracy. It had never been shown a letter that touched its neighbour.
0.010 at messiness 3, and the sample read is 'T'. The model is not producing wrong letters so much as producing almost nothing. Under CTC, a model that recognises nothing outputs blank, and blank collapses to the empty string. Deletion, not substitution, is the signature of a domain mismatch in a CTC reader.
0.960 on both, from the identical network. This is the whole lesson. No layer changed. No hyperparameter changed. The training distribution changed.
The messy-trained model also reads clean print at 0.960. Training on the harder distribution costs almost nothing on the easier one. That asymmetry is why augmentation is nearly always worth adding, and why "train on clean, deploy on messy" is nearly always a mistake.
The augmentations that matter for real handwriting
Ordered roughly by how much they buy, on real data:
| Augmentation | What it simulates | Typical range |
|---|---|---|
| Shear / slant | Writing angle | ±0.3 shear factor |
| Elastic distortion | Wobbly strokes | small local displacements |
| Stroke thickness (dilate/erode) | Pen and pressure | ±1 pixel |
| Baseline shift and rotation | Unlined paper | ±3 degrees |
| Contrast and blur | Scan quality | mild |
| Random word spacing | Personal spacing habits | ±50% |
Elastic distortion (Simard et al., ICDAR 2003) is the most valuable single one for handwriting and is missing from the demo above only to keep the code short. torchvision.transforms.v2.ElasticTransform implements it.
Common mistakes
Evaluating on the same writers you trained on. Handwriting datasets must be split by writer. A random split leaks a writer's style into both sides and inflates accuracy substantially. This is data leakage in its most specific form.
Reporting character error rate only. For a form field, a single wrong character makes the field wrong. Report exact-match accuracy per field alongside CER, as the code above does.
Normalising line height by stretching. Resize by height while preserving aspect ratio, then pad. Squashing a long line to a fixed width destroys stroke geometry that the model needs.
Skipping slant correction. For classical pipelines, estimating and removing the writing slant before recognition is a large, cheap win, and it is the geometric equivalent of the augmentation above.
Assuming a printed-text checkpoint will transfer. It will not, and the run above shows the size of the gap. Fine-tune on handwriting, or start from a handwriting checkpoint.
Try it yourself
Train at messiness 1 and test across 0 to 3. You will see partial transfer — better than the print-trained model at every level, and worse than the messy-trained one at level 3. That curve is the practical shape of "how much augmentation is enough".
What to learn next
- Document layout analysis — finding the lines before reading them.
- Data augmentation — the lever this lesson measured.
- CTC loss — why the failed reads came back empty rather than wrong.
Researcher — Mathematics and papers.
Offline against online
Two distinct problems that share a name.
Online handwriting recognition receives the pen trajectory: an ordered sequence of $(x, y, t)$ points, plus pen-up and pen-down events. Stroke order and direction are known, which removes most of the ambiguity. Accuracy on online data is far higher, and this is what a stylus-based note app solves.
Offline recognition receives only an image. All temporal information is gone, and a crossing of two strokes is indistinguishable from a single stroke. Everything in this lesson is offline.
Papers reporting spectacular handwriting numbers should be checked for which of the two they solved.
The standard benchmarks
- IAM (Marti and Bunke, 2002) — 1,539 pages of English from 657 writers. The standard offline benchmark. Splits are defined by writer, and the Aachen split is the one most papers use; comparisons across different splits are not comparable.
- RIMES — French administrative correspondence.
- Bentham and READ / ICFHR collections — historical manuscripts, where the alphabet and the language model both shift.
- IFN/ENIT — Arabic handwritten town names, where cursive is the norm rather than the exception.
Reported IAM word error rates in the low single digits usually include a strong language model trained on the same corpus. The recogniser's own error rate is considerably higher, and the distinction matters when you have no in-domain text to build a language model from.
Architectures
MDLSTM (Graves and Schmidhuber, NIPS 2008) scanned the image with LSTMs in all four diagonal directions and dominated handwriting benchmarks for a decade. It is expensive and awkward to parallelise.
CNN plus BiLSTM plus CTC — the CRNN family — replaced it, matching or beating MDLSTM at a fraction of the cost (Puigcerver, ICDAR 2017, is the clean demonstration). This remains a strong default.
Attention encoder-decoder and transformer models are stronger on hard scripts because the decoder supplies an implicit language model, and TrOCR's handwritten checkpoints sit here. They carry the hallucination risk described in TrOCR and transformer OCR, which is why cheque and form pipelines often still use CTC.
Hybrid CTC-attention training gets most of both: the CTC head enforces monotonic alignment while the attention head models output dependencies.
Where the accuracy actually comes from
Three levers, in descending order of effect on real projects.
Synthetic data at scale. Rendering millions of lines with many handwriting fonts, then fine-tuning on a few thousand real lines, is the standard recipe and generally beats any architecture change. The experiment in the developer block is a miniature of this.
A language model at decode time. Prefix beam search with a word-level or character-level language model, shallow-fused, is worth several absolute points of WER on IAM. It also introduces a bias toward in-vocabulary words, which is harmful for names, codes and amounts — precisely the fields most business documents care about. Disable it for those fields.
Writer adaptation. When many pages come from one writer, a few labelled lines from that writer, used for a short fine-tune or for a learned writer embedding, gives a large gain. Historical archive projects exploit this heavily.
The unsolved part
Free-form, unconstrained, multi-writer handwriting on unlined paper with mixed scripts remains open. Reported gains often come from constrained settings — known field types, boxed characters, a closed lexicon — and those constraints do more work than the model does. When evaluating a claim, look for the constraint before looking at the number.
What to learn next
- Document layout analysis — finding the lines before reading them.
- Data augmentation — the lever this lesson measured.
- CTC loss — why the failed reads came back empty rather than wrong.