How it works

How OCR works

OCR cleans up the photo, finds where the text lines are, reads each line as a sequence rather than letter by letter, and uses language knowledge to fix what the pixels left unclear.

On this page 7
  1. The pipeline at a glance
  2. Stage 1 — clean up the image
  3. Stage 2 — find the text
  4. Stage 3 — read each strip as a sequence
  5. Stage 4 — let language fix what pixels cannot
  6. Stage 5 — put the page back together
  7. Which lessons teach each stage

Point your camera at a shop sign and your phone translates it. Deposit a cheque by photographing it. Search a word inside a scanned PDF. All of this is OCR — optical character recognition, turning pictures of text into actual text a computer can edit and search. Here is the pipeline from photo to characters.

The pipeline at a glance

 photo of a page / sign / receipt
        |
        v
 [1. preprocess]   straighten, clean, boost contrast
        |
        v
 [2. detect]       find the boxes where text lives
        |
        v
 [3. recognise]    read each box as a character sequence
        |
        v
 [4. correct]      language knowledge fixes 0/O, 1/l
        |
        v
 [5. structure]    reading order, fields, tables

Stage 1 — clean up the image

Real photos are tilted, shadowed and crumpled. Preprocessing straightens and scrubs: rotate so lines run horizontal (deskewing), boost contrast, remove noise, and often reduce the image toward plain dark-ink-on-light-paper.

The old rule holds hard here: garbage in, garbage out. A blurry photo does not contain the information, and no later stage can invent it. Most "OCR is bad" complaints are actually "photo was bad".

Stage 2 — find the text

Before reading anything, the system must find where text is at all. A text detection model — a neural network trained on images with text regions marked — outputs boxes around every word or line. Three lines up top, a paragraph, a phone number at an angle across a poster.

Detection and reading are deliberately separate stages. The detector answers only "where is text?", so the next model can concentrate on the far harder question — "what does it say?" — one small strip at a time.

Stage 3 — read each strip as a sequence

Each detected strip is read by a recognition model — typically a network that sees the strip's visual features (a CNN, a network built for images) feeding a sequence model that outputs characters in order.

The word sequence is the important one. Early OCR cut strips into single characters and classified each in isolation, which collapsed the moment letters touched or fonts got creative. Modern OCR reads the whole strip in one pass and emits the character sequence directly — nothing has to decide in advance where one letter ends and the next begins. The model learns the segmentation implicitly, from data.

This is also what makes handwriting readable at all: in joined-up writing there are no clean letter boundaries to find, and you read words, not letters. So does the model.

Stage 4 — let language fix what pixels cannot

Some shapes are genuinely ambiguous. 0 and O. 1, l and I. rn pretending to be m. Pixels alone cannot settle these — but language can. "L0ND0N" is almost certainly "LONDON"; a price of "1OO" is "100".

So recognition output passes through a correction layer with a feel for likely character sequences — from simple dictionaries up to neural language models. Domain rules help too: a cheque amount must be digits, a number-plate slot follows a fixed pattern. Note the honest cost: this stage fixes common words and occasionally breaks rare ones, "correcting" an unusual surname into a dictionary word.

Stage 5 — put the page back together

Characters are not the finish line. The boxes must be ordered into reading flow, which is tricky with columns, tables and stamps. For forms and receipts they must also be mapped to meaning: this string is the date, that one is the total. This layer, document AI, is where OCR meets language models, and it is why you can photograph a bill and have the amount land in the right field of an app.

Which lessons teach each stage