Messy Real-World Text

Fixing OCR errors

OCR turns scanned images into text, and it makes specific, predictable mistakes — like reading "1" as "l" — that correction rules can target directly.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

OCR (Optical Character Recognition) turns a scanned image of text into actual text a computer can read. It makes very specific, repeatable mistakes while doing it.

Think about reading a faded photocopy. Most words come through fine. Some letters genuinely look alike on the page. A capital "I," a lowercase "l," and the digit "1" can look nearly identical in some fonts. An old, much-copied page makes it worse.

OCR software has the same problem, at scale, on every scanned page it processes. The errors are not random typos — they are shape confusions, and they repeat in predictable ways.

Why it exists

Scanned invoices, old books, government forms and photographed documents all need to become searchable, editable text. OCR does that conversion. But the software is reading pixel shapes, not meaning, so it occasionally reads the wrong shape.

Because the confusions are visual, not random, they follow patterns. "0" and "O." "1," "l" and "I." "rn" and "m." "5" and "S." Correction tools built for OCR specifically target these known confusions, rather than treating every error as an unpredictable typo.

How it works

  scanned page  -->  OCR engine  -->  "1nvoice"  ("1" read where "I" belongs)
                                            |
                                            v
                          known confusion table + dictionary check
                                            |
                                            v
                                       "invoice"

Where you have already seen it

  • Searchable PDFs made from scanned documents, where the underlying text sometimes has visible garbled words.
  • Digitised old newspapers and books, where library archive projects run correction passes over OCR output.
  • Receipt-scanning apps, extracting totals and item names from a photographed bill.

Remember this

  • OCR mistakes are visual confusions between similar-looking characters, not random typing errors.
  • Because the confusion patterns are known and repeat, targeted substitution rules catch far more errors than generic spell-checking alone.
  • Aggressive correction rules can break correctly-read words that happen to contain the same letter patterns as a known confusion.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install rapidfuzz

Fixing common OCR character confusions

ocr_fix_demo.py
from rapidfuzz import process, fuzz

# A small vocabulary. A real system would have thousands of domain words.
VOCAB = ["hello", "world", "invoice", "total", "amount", "date", "number", "modern", "model"]

# Typical OCR confusions: 0/O, 1/l, 5/S, 8/B, and "rn" misread as "m"
OCR_SUBS = str.maketrans({"0": "o", "1": "l", "5": "s", "8": "b"})

def fix_ocr_word(word, vocab):
    candidate = word.translate(OCR_SUBS).replace("rn", "m")
    match, score, _ = process.extractOne(candidate, vocab, scorer=fuzz.ratio)
    return match if score > 80 else word

scanned = ["1nvoice", "m0dern", "tota1", "arnount", "xyzabc"]
for w in scanned:
    print(f"{w!r:12} -> {fix_ocr_word(w, VOCAB)!r}")
Output
'1nvoice'    -> 'invoice'
'm0dern'     -> 'm0dern'
'tota1'      -> 'total'
'arnount'    -> 'amount'
'xyzabc'     -> 'xyzabc'

Line by line

Three of the five OCR errors got fixed cleanly. "1nvoice" and "tota1" had their digit-for-letter confusion swapped correctly. "arnount" — "rn" misread where "m" belongs, a real and common OCR error — got fixed by the explicit .replace("rn", "m") rule.

"m0dern" did not get fixed, and that failure is worth understanding. After substitution, "m0dern" becomes "modern" — but then .replace("rn", "m") fires again, on the "rn" inside "modern" itself, turning it into "modem." "modem" scores only 80.0 against the vocabulary's "model," right at the threshold, not above it — so the function correctly falls back to leaving the word unchanged rather than confidently returning the wrong word.

"xyzabc" correctly stayed unchanged. It matches nothing in the vocabulary closely enough, and the score > 80 threshold is exactly what stopped a wrong guess from being returned with false confidence.

Common mistakes

Chaining substitution rules without checking for unintended collisions. The "rn" to "m" rule, built to catch a real OCR error, quietly breaks the legitimate word "modern" by matching a coincidental "rn" inside it. Order and specificity of rules matters — this is a genuine trap, not a hypothetical one, as the output above shows directly.

Applying character substitutions before checking whether the original word is already correct. Always check the untouched word against the vocabulary first — only apply OCR-confusion substitutions if that direct check fails.

Setting the similarity threshold too low. A low threshold "fixes" more words, but also confidently returns wrong corrections for words that only vaguely resemble a real vocabulary entry. The threshold in this demo, 80, is what caught the "modem" near-miss and refused to guess.

Try it yourself

Change OCR_SUBS to also map "5" and "S" and test it against a word like "5tate" (should become "state," if "state" were in the vocabulary). Then check whether adding "modem" as a legitimate vocabulary word changes the outcome for "m0dern" — a good illustration of how correction quality depends entirely on what the vocabulary already contains.

What to learn next

Researcher — Mathematics and papers.

OCR errors as a structured, non-uniform noise channel

Unlike typing errors, which are reasonably well modelled as roughly uniform-probability single-character edits (the noisy-channel model discussed in spelling correction), OCR errors are dominated by a small set of highly-probable, font- and resolution-dependent visual confusions. A confusion matrix P(observed | true), estimated from a labelled corpus of scanned-and-corrected text, is sharply peaked on visually similar glyph pairs — l/1/I, O/0, rn/m, cl/d — rather than spread uniformly across the alphabet.

This structure is exactly what the hand-coded substitution table in the developer demo approximates crudely. Production OCR post-correction systems learn this confusion matrix directly from paired (OCR output, ground truth) data rather than hand-specifying it, since real confusion patterns vary by scanner, font, resolution and even document age (older, more degraded scans shift the confusion distribution further).

Formal correction as noisy-channel decoding

text
correct(w) = argmax over c in vocabulary of  P(c) * P(w | c)
  • P(c) is the language-model probability of candidate word c — identical in form to the spelling-correction formulation.
  • P(w | c) is now an OCR-specific confusion likelihood rather than a typing-error likelihood, learned from the confusion matrix described above.

This is the same noisy-channel structure as general spelling correction, with a different, more structured and more learnable error model — which is precisely why OCR-specific correction reliably outperforms generic spell-checking applied to OCR output: it exploits structure that generic edit-distance search cannot see.

Sequence-level correction

Character-level substitution, as in the developer demo, corrects one character confusion at a time and cannot easily fix errors that shift word boundaries — merged words ("theinvoice"), split words ("in voice"), or multi-character misreads that no longer resemble the true word closely enough for fuzzy matching to find it.

Modern OCR post-correction increasingly uses sequence-to-sequence neural models — the same architecture family as neural machine translation — trained on (raw OCR output, corrected text) pairs, learning to fix boundary errors and multi-character confusions jointly rather than one character rule at a time (Dong & Smith, 2018; Nguyen et al., 2021, survey OCR post-correction approaches directly).

Key references

  • Dong, R. & Smith, D. (2018). Multi-Input Attention for Unsupervised OCR Correction. ACL. — sequence-to-sequence OCR post-correction.
  • Nguyen, T. et al. (2021). Survey of Post-OCR Processing Approaches. ACM Computing Surveys. — a comprehensive review of both rule-based and neural OCR correction methods.
  • Kolak, O. & Resnik, P. (2002). OCR Error Correction Using a Noisy Channel Model. HLT. — the direct noisy-channel formulation applied to OCR specifically.
  • Smith, R. (2007). An Overview of the Tesseract OCR Engine. ICDAR. — the widely-used open-source OCR engine whose confidence scores and confusion patterns motivate much of this post-processing work.

Current state and open problems

Modern OCR engines, including deep-learning-based ones, have substantially lower baseline error rates than older template-matching engines, reducing the burden on post-correction for clean, modern scans. Correction remains genuinely important for degraded historical documents, low-quality photographs of documents (mobile receipt-scanning being a common real product case), and non-Latin scripts, where OCR error rates stay meaningfully higher.

Indic-script OCR specifically lags Latin-script OCR in accuracy, owing to the complexity of conjunct consonants and matra placement discussed in normalising Indic scripts — the same script complexity that makes text representation tricky also makes visual character recognition harder, and post-correction tooling for Indic-script OCR is correspondingly less mature than for English.

What to learn next