How it works
How OCR works
OCR cleans up the photo, finds where the text lines are, reads each line as a sequence rather than letter by letter, and uses language knowledge to fix what the pixels left unclear.
- 4 min read
- Updated
On this page 7
Point your camera at a shop sign and your phone translates it. Deposit a cheque by photographing it. Search a word inside a scanned PDF. All of this is OCR — optical character recognition, turning pictures of text into actual text a computer can edit and search. Here is the pipeline from photo to characters.
The pipeline at a glance
photo of a page / sign / receipt
|
v
[1. preprocess] straighten, clean, boost contrast
|
v
[2. detect] find the boxes where text lives
|
v
[3. recognise] read each box as a character sequence
|
v
[4. correct] language knowledge fixes 0/O, 1/l
|
v
[5. structure] reading order, fields, tablesStage 1 — clean up the image
Real photos are tilted, shadowed and crumpled. Preprocessing straightens and scrubs: rotate so lines run horizontal (deskewing), boost contrast, remove noise, and often reduce the image toward plain dark-ink-on-light-paper.
The old rule holds hard here: garbage in, garbage out. A blurry photo does not contain the information, and no later stage can invent it. Most "OCR is bad" complaints are actually "photo was bad".
Stage 2 — find the text
Before reading anything, the system must find where text is at all. A text detection model — a neural network trained on images with text regions marked — outputs boxes around every word or line. Three lines up top, a paragraph, a phone number at an angle across a poster.
Detection and reading are deliberately separate stages. The detector answers only "where is text?", so the next model can concentrate on the far harder question — "what does it say?" — one small strip at a time.
Stage 3 — read each strip as a sequence
Each detected strip is read by a recognition model — typically a network that sees the strip's visual features (a CNN, a network built for images) feeding a sequence model that outputs characters in order.
The word sequence is the important one. Early OCR cut strips into single characters and classified each in isolation, which collapsed the moment letters touched or fonts got creative. Modern OCR reads the whole strip in one pass and emits the character sequence directly — nothing has to decide in advance where one letter ends and the next begins. The model learns the segmentation implicitly, from data.
This is also what makes handwriting readable at all: in joined-up writing there are no clean letter boundaries to find, and you read words, not letters. So does the model.
Stage 4 — let language fix what pixels cannot
Some shapes are genuinely ambiguous. 0 and O. 1, l and I. rn pretending to be m. Pixels alone cannot settle these — but language can. "L0ND0N" is almost certainly "LONDON"; a price of "1OO" is "100".
So recognition output passes through a correction layer with a feel for likely character sequences — from simple dictionaries up to neural language models. Domain rules help too: a cheque amount must be digits, a number-plate slot follows a fixed pattern. Note the honest cost: this stage fixes common words and occasionally breaks rare ones, "correcting" an unusual surname into a dictionary word.
Stage 5 — put the page back together
Characters are not the finish line. The boxes must be ordered into reading flow, which is tricky with columns, tables and stamps. For forms and receipts they must also be mapped to meaning: this string is the date, that one is the total. This layer, document AI, is where OCR meets language models, and it is why you can photograph a bill and have the amount land in the right field of an app.
Which lessons teach each stage
- The field this lives in: What is computer vision?
- Stages 2–3, networks that read images: CNNs and Image classification
- Stage 3, models that handle sequences: RNNs
- How any of these networks learned: How neural networks learn