How an OCR system is put together
OCR is not one model but a short assembly line — clean the image, find the text, read each piece, then put the page back together in reading order.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
OCR turns a picture of text back into text you can copy, search and edit.
OCR stands for optical character recognition: looking at letters and working out which ones they are.
The analogy you have already lived
Think of reading a handwritten shopping list left on the fridge. Your eyes do four things without you noticing. You hold the paper straight. You spot where the writing is. You read each word. You keep the lines in order.
An OCR system does those same four things, one after another, as separate steps.
Why it exists
A photo of a bill is a grid of coloured dots. Your phone stores millions of dots and knows nothing about what they spell.
You cannot search that photo for the word "total". You cannot add up the amounts. You cannot paste a line into a message.
Every one of those becomes possible the moment the dots become letters. That single conversion is what OCR sells.
The assembly line
photo of a page
|
[1] clean it up straighten, remove shadow, make ink black
|
[2] find the text draw a box around every word
|
[3] read each box box of pixels -> "TOTAL"
|
[4] rebuild the page put boxes back in reading order
|
text you can searchEach stage can fail on its own. That is good news, because it tells you where to look when the output is wrong.
If stage 2 misses a word, no amount of clever reading fixes it — the word was never sent. If stage 4 gets confused, you get every correct word in nonsense order.
The part people get wrong
Many people picture OCR as one giant model that swallows a page and produces text. Some newer systems do behave that way, and we cover them later in this section.
The four-stage line still matters, because it is how you debug. When a scanned invoice comes out as gibberish, you check the stages in order. Nine times out of ten the picture was bad, not the reader.
Where you have already seen it
- Google Lens copying text out of a signboard you photographed.
- Your bank app reading a cheque you photographed for deposit.
- Toll booths reading the number plate of your car.
- A PDF where you can select the text of a scanned old book.
- Aadhaar or PAN details auto-filled from a photo of the card.
The honest part
OCR looks solved because it works beautifully on clean, printed, straight, English pages. That is the easy corner of the problem.
Move to a crumpled receipt, a Devanagari handwritten form, or a photo taken at an angle. Accuracy falls hard. Most of the real work in a document project is fighting those cases, not choosing a model.
Remember this
- OCR is four stages: clean, find, read, reorder.
- A wrong answer usually comes from one stage, and you can find out which.
- Clean printed text is easy; handwriting, angles and bad light are not.
What to learn next
- Preprocessing scans for OCR — stage one, and the cheapest accuracy you will ever buy.
- Text detection models — how stage two finds words in a photograph.
- Document AI — what you do with the text once you have it.
Developer — Code and libraries.
Setup
pip install numpy==1.26.4The point of this lesson is that a working OCR pipeline is small enough to read in one sitting. So we build one, end to end, with no OCR library at all.
To keep the pixels identical on every machine, the font is defined in the file as a 5x7 bitmap rather than loaded from disk.
A whole OCR pipeline in one file
import numpy as np
# A 5x7 bitmap font. '#' is ink. Defining it here keeps the lesson free of
# font files, so the same pixels appear on every machine.
FONT = {
"A": ["..#..", ".#.#.", "#...#", "#...#", "#####", "#...#", "#...#"],
"C": [".###.", "#...#", "#....", "#....", "#....", "#...#", ".###."],
"E": ["#####", "#....", "#....", "####.", "#....", "#....", "#####"],
"I": ["#####", "..#..", "..#..", "..#..", "..#..", "..#..", "#####"],
"L": ["#....", "#....", "#....", "#....", "#....", "#....", "#####"],
"N": ["#...#", "##..#", "#.#.#", "#.#.#", "#..##", "#...#", "#...#"],
"1": ["..#..", ".##..", "..#..", "..#..", "..#..", "..#..", ".###."],
"4": ["...#.", "..##.", ".#.#.", "#..#.", "#####", "...#.", "...#."],
"9": [".###.", "#...#", "#...#", ".####", "....#", "#...#", ".###."],
}
GLYPHS = {ch: np.array([[c == "#" for c in row] for row in rows], dtype=float)
for ch, rows in FONT.items()}
def render(text, pad=3, gap=2):
"""Draw a line of text into a 0/1 image, with blank columns between letters."""
cells = [np.zeros((7, gap))]
for ch in text:
cells.append(np.zeros((7, 3)) if ch == " " else GLYPHS[ch])
cells.append(np.zeros((7, gap)))
return np.pad(np.hstack(cells), pad)
def find_boxes(image):
"""Step 1 and 2: crop to the text line, then cut it at blank columns."""
inked_rows = np.where(image.sum(axis=1) > 0)[0]
band = image[inked_rows[0]:inked_rows[-1] + 1]
inked_cols = np.append(band.sum(axis=0) > 0, False)
boxes, start = [], None
for x, has_ink in enumerate(inked_cols):
if has_ink and start is None:
start = x
elif not has_ink and start is not None:
boxes.append((start, x)) # one run of inked columns = one glyph
start = None
return band, boxes
def to_cell(patch, width=5):
"""Glyphs like '1' are narrower than their cell, so re-centre before matching."""
left = (width - patch.shape[1]) // 2
return np.pad(patch, ((0, 0), (left, width - patch.shape[1] - left)))
def recognise(patch):
"""Step 3: nearest template wins — fewest disagreeing pixels."""
cell = to_cell(patch)
scored = sorted((np.abs(g - cell).sum(), ch) for ch, g in GLYPHS.items())
return scored[0][1], scored
page = render("CLEAN LINE 1949")
print(f"page: {page.shape[0]} rows x {page.shape[1]} columns of pixels\n")
for row in page[3:10]:
print(" " + "".join("#" if v else "." for v in row))
band, boxes = find_boxes(page)
print(f"\nboxes found: {len(boxes)}")
print("first three (start column, end column):", boxes[:3])
gaps = [boxes[i + 1][0] - boxes[i][1] for i in range(len(boxes) - 1)]
space_at = (min(gaps) + max(gaps)) / 2 # word gaps are far wider than letter gaps
print("gaps between boxes:", gaps)
print(f"gap sizes run {min(gaps)} to {max(gaps)}, so anything over {space_at} is a space")
read = ""
for i, (x0, x1) in enumerate(boxes):
read += recognise(band[:, x0:x1])[0]
if i < len(gaps) and gaps[i] > space_at:
read += " "
print("\nrecognised:", repr(read))
_, scored = recognise(band[:, boxes[1][0]:boxes[1][1]])
print("what box 2 was choosing between:", [(c, int(s)) for s, c in scored[:3]])page: 13 rows x 109 columns of pixels
......###...#......#####....#....#...#.......#......#####..#...#..#####.........#.....###......#....###......
.....#...#..#......#.......#.#...##..#.......#........#....##..#..#............##....#...#....##...#...#.....
.....#......#......#......#...#..#.#.#.......#........#....#.#.#..#.............#....#...#...#.#...#...#.....
.....#......#......####...#...#..#.#.#.......#........#....#.#.#..####..........#.....####..#..#....####.....
.....#......#......#......#####..#..##.......#........#....#..##..#.............#........#..#####......#.....
.....#...#..#......#......#...#..#...#.......#........#....#...#..#.............#....#...#.....#...#...#.....
......###...#####..#####..#...#..#...#.......#####..#####..#...#..#####........###....###......#....###......
boxes found: 13
first three (start column, end column): [(5, 10), (12, 17), (19, 24)]
gaps between boxes: [2, 2, 2, 2, 7, 2, 2, 2, 8, 3, 2, 2]
gap sizes run 2 to 8, so anything over 5.0 is a space
recognised: 'CLEAN LINE 1949'
what box 2 was choosing between: [('L', 0), ('E', 7), ('C', 8)]Reading the output carefully
The four stages are visible in the four print statements. Crop to the line, cut into boxes, match each box, join with spaces. Every real OCR engine has these same joints, with far stronger parts bolted in.
gaps between boxes: [2, 2, 2, 2, 7, 2, 2, 2, 8, 3, 2, 2] is the whole word-splitting problem in one list. Letters sit 2 or 3 columns apart, words 7 or 8. Nothing in the image says "space" — the space is inferred from a gap being unusually wide, and that inference is a common source of TOTALAMOUNT style errors on real receipts.
[('L', 0), ('E', 7), ('C', 8)] shows the recogniser's confidence, honestly. A cost of 0 means a perfect pixel match, and the runner-up is 7 pixels away. A large margin like that is a confident read. When the top two costs are close, the character is genuinely ambiguous — that is where O/0 and l/1/I errors are born.
The 3 in the gap list is a near miss. It sits between 1 and 9, because the 1 glyph does not fill its cell. One column more and our rule would have split 1949 into two words.
Why real engines are bigger than this
Our version assumes a great deal, and each assumption is a whole lesson later in this section.
| Our assumption | Reality | Where it is handled |
|---|---|---|
| Perfectly black-and-white pixels | Grey, shadowed, JPEG-blurred scans | Preprocessing scans |
| One straight line of text | Curved, angled, multi-column pages | Text detection models |
| Letters never touch | Printed letters merge constantly | CRNN text recognition |
| We know where each letter starts | Almost never true in cursive | CTC loss |
| One font, one size | Thousands of fonts | Trained models, not templates |
The jump from templates to neural networks happens at row three. Once letters touch, cutting between them stops working, and the segment-then-classify design collapses. The fix was to stop segmenting and read a whole word strip at once.
Common mistakes
Feeding a colour photo straight to an OCR engine. Convert to greyscale and binarise first. Most engines do a rough job internally, and doing it deliberately is usually better.
Blaming the recogniser for a detection failure. Before tuning the model, draw the detected boxes on the image and look. Missing text is a stage 2 bug, and no model change fixes it.
Assuming reading order comes free. Boxes come out of a detector in whatever order the network produced. Two-column pages, tables and sidebars all need explicit sorting, which is document layout analysis.
Measuring accuracy by eye. "It looks fine" hides a 4% character error rate. Score with character error rate on a held-out set of your own documents, not the vendor's demo page.
Try it yourself
Add a glyph for O and 0 to FONT, then render "1O0". Print the full candidate list for each box. You will see the top two costs sitting within a pixel or two of each other, which is exactly why every real OCR system has post-processing rules for digits inside amounts.
What to learn next
- Preprocessing scans for OCR — stage one, and the cheapest accuracy you will ever buy.
- Text detection models — how stage two finds words in a photograph.
- Document AI — what you do with the text once you have it.
Researcher — Mathematics and papers.
The classical decomposition
The pipeline formalised by the OCR literature is a chain of conditional models:
$$ P(\text{text} \mid \text{image}) \approx \prod_{\text{stage}} P(\text{stage}i \mid \text{stage}{i-1}) $$
The chain is trained stage-wise and run greedily, so errors compound. If detection recall is $r$ and per-word recognition accuracy is $a$, end-to-end word accuracy is bounded above by $r \cdot a$. A detector at $0.95$ recall paired with a recogniser at $0.97$ gives at most $0.92$ — which is why detection recall is usually the first thing to fix.
The three architectural eras
Segment-then-classify (1970s–2000s). Characters are isolated by connected components or projection profiles, then classified individually. Tesseract's pre-4.0 engine is the best-known survivor. It fails on touching, broken or cursive glyphs, and the failure is unrecoverable because segmentation is committed to before classification.
Segmentation-free sequence models (2015–2021). A detector emits word or line boxes; a recogniser maps a whole strip to a character sequence without cutting it. Shi, Bai and Yao (2017), An end-to-end trainable neural network for image-based sequence recognition (CRNN), plus Graves et al. (2006), Connectionist temporal classification, are the two papers that made this era work. Detection moved to segmentation-style networks: EAST (Zhou et al., 2017), CRAFT (Baek et al., 2019), and DB (Liao et al., AAAI 2020), extended as DBNet++ in TPAMI 2022.
End-to-end vision-language models (2022–now). A single encoder-decoder consumes the page and emits structured text, skipping explicit detection. Donut (Kim et al., ECCV 2022) removed OCR from document understanding entirely; Nougat (Blecher et al., 2023) targeted scientific PDFs; general VLMs and OCR-specialised derivatives now handle full-page markdown extraction. The 2026 open-source landscape splits cleanly along this line: pipeline engines such as Tesseract, PaddleOCR and docTR on one side, VLM readers on the other.
What each era trades
| Property | Pipeline | End-to-end VLM |
|---|---|---|
| Word-level bounding boxes | Native | Usually absent or approximate |
| Failure mode | Localised, debuggable | Global, silent |
| Hallucination | Structurally impossible | Documented and real |
| Cost per page | Milliseconds on CPU | Seconds on GPU |
| Reading order, tables | Needs explicit modelling | Learned, often better |
The hallucination row is decisive for regulated work. A CTC-based recogniser cannot emit a character that no pixel column supported. An autoregressive decoder can, and does, invent plausible account numbers when the crop is illegible. Any pipeline touching money or identity should either stay on the left column or carry a verification pass.
Evaluation
Character error rate is Levenshtein distance normalised by reference length:
$$ \mathrm{CER} = \frac{S + D + I}{N} $$
Where $S$, $D$ and $I$ are substitutions, deletions and insertions in the optimal alignment, and $N$ is the number of characters in the reference. Word error rate is the same quantity computed over word tokens.
Two cautions that trip up published comparisons. CER is unbounded above, since insertions can exceed $N$ — a hallucinating model can score above 1.0. And CER is undefined for reading order: a system that emits every word correctly but interleaves two columns can post an excellent per-line CER and be useless. Structure-aware metrics (TEDS for tables, GriTS for table structure) exist precisely because of that gap.
Open datasets worth knowing
- ICDAR Robust Reading challenges (2013, 2015, 2019 ArT, 2019 MLT) — scene text detection and recognition.
- IAM and RIMES — offline handwriting, still the standard handwriting benchmarks.
- PubLayNet (Zhong et al., ICDAR 2019) and DocLayNet (Pfitzmann et al., KDD 2022) — layout. DocLayNet is human-annotated across six document domains and is the more honest benchmark; PubLayNet is auto-annotated from PubMed and is heavily biased toward two-column scientific articles.
- FUNSD, CORD, SROIE, XFUND — form and receipt understanding, with key-value labels.
- PubTables-1M (Smock et al., CVPR 2022) — table detection and structure.
What to learn next
- Preprocessing scans for OCR — stage one, and the cheapest accuracy you will ever buy.
- Text detection models — how stage two finds words in a photograph.
- Document AI — what you do with the text once you have it.