OCR and Document Vision

Tesseract in practice

Tesseract is free, offline and everywhere, and it repays a good scan and the right page segmentation mode far more than it repays clever code.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. The single most important setting
  5. What it needs from you
  6. Where you have already seen it
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Tesseract is a free OCR program that runs on your own machine, with no internet and no bill.

The analogy you have already lived

Think of an old, reliable scooter that has been repaired and improved for thirty years. It is not the fastest thing on the road. It starts every morning, it costs nothing to run, and every mechanic knows it.

Tesseract is that scooter. Newer engines beat it on hard pages. It still gets more documents read, every day, than anything shiny.

Why it exists

It began at HP in the 1980s, was open-sourced in 2005, and has been maintained ever since. Version 4 replaced its old letter-by-letter engine with a neural network that reads whole lines.

Its value is not accuracy. Its value is that it costs nothing, needs no server, and sends your documents nowhere.

For anything private — medical records, salary slips, legal papers — that last point decides the matter on its own.

The single most important setting

Tesseract will do a whole page, or one line, or one word. You tell it which with a setting called page segmentation mode, usually written --psm.

Getting this wrong is the most common reason people conclude "Tesseract is bad".

   a full scanned page      ->  --psm 3   (the default, finds columns itself)
   one cropped line         ->  --psm 7   ("this is a single line")
   one cropped word         ->  --psm 8   ("this is a single word")
   scattered text in a photo->  --psm 11  ("find text anywhere, any order")

Hand it a tight crop of one word while it is in page mode, and it hunts for paragraph structure. There is none, so it often returns nothing at all.

What it needs from you

Tesseract was built for scanned pages, and it expects to be treated like one.

Enough resolution. The official guidance is at least 300 DPI. Under that, letters lose the strokes that tell them apart.

A white margin. A letter touching the edge of the image is frequently dropped. Ten pixels of white border fixes it.

A straight, clean page. Everything in preprocessing scans for OCR pays off here more than for any other engine.

Where you have already seen it

  • Searchable PDFs made by a scanner's bundled software.
  • Archive projects digitising old newspapers and court records.
  • Offline receipt scanners inside accounting apps.
  • Countless small government and hospital systems that cannot send data outside.

The honest part

Tesseract is weak where documents are most interesting. Photographs at an angle, curved signboards, handwriting, and dense tables all defeat it.

It also gives you no layout understanding. You get words and their positions, and rebuilding the meaning of a form is entirely your job.

Knowing this in advance is what separates a working project from a frustrated one. Use Tesseract for clean printed pages. Reach for something else when your input is a phone photo of a crumpled bill.

Remember this

  • Tesseract is free, offline and private. That is why it is still everywhere.
  • The page segmentation mode is the setting that decides whether it works.
  • It needs 300 DPI, a straight page and a white margin. Give it those first.

What to learn next

Developer — Code and libraries.

Setup

Tesseract is a system program, not a Python package. The Python wrapper only calls it.

bash
# Ubuntu / Debian
sudo apt install tesseract-ocr tesseract-ocr-hin tesseract-ocr-tam

# macOS
brew install tesseract tesseract-lang

# Windows: install the UB-Mannheim build, then add its folder to PATH

pip install pytesseract==0.3.13 pillow==11.0.0

Written against Tesseract 5.5 and pytesseract 0.3.13. Check your install with tesseract --version and list installed languages with tesseract --list-langs.

Calling Tesseract properly

run_tesseract.py
import pytesseract
from PIL import Image

print("tesseract version:", pytesseract.get_tesseract_version())
print("languages installed:", pytesseract.get_languages())

page = Image.open("scan.png")

# --psm 3  = fully automatic page segmentation (the default)
# --oem 1  = LSTM engine only, which is what you want on Tesseract 4 and 5
text = pytesseract.image_to_string(page, lang="eng", config="--psm 3 --oem 1")
print(text)

# The same call, but asking for boxes and confidences instead of a blob of text.
tsv = pytesseract.image_to_data(page, lang="eng", config="--psm 3 --oem 1")
print(tsv)

# A tight crop of one line reads far better in single-line mode.
line = page.crop((36, 40, 356, 62))
print(pytesseract.image_to_string(line, config="--psm 7 --oem 1").strip())

# Restrict the alphabet when you know the field. This is the highest-value
# single flag on forms: an amount field cannot contain letters.
amount = page.crop((124, 96, 216, 114))
digits_only = "--psm 7 --oem 1 -c tessedit_char_whitelist=0123456789.,"
print(pytesseract.image_to_string(amount, config=digits_only).strip())

Page segmentation modes, in full

These are the values Tesseract's own manual page lists. The ones you will actually use are marked.

--psmMeaning
0Orientation and script detection only, no OCR
1Automatic page segmentation with orientation detection
2Automatic segmentation, no orientation, no OCR (not implemented)
3Fully automatic page segmentation, no orientation detectiondefault
4Assume a single column of text of variable sizesreceipts
5Assume a single uniform block of vertically aligned text
6Assume a single uniform block of textone paragraph
7Treat the image as a single text lineline crops
8Treat the image as a single wordword crops
9Treat the image as a single word in a circle
10Treat the image as a single character
11Sparse text, find as much as possible in no particular orderphotos
12Sparse text with orientation detection
13Raw line, bypassing Tesseract-specific processing

And the engine modes: --oem 0 original engine, --oem 1 LSTM only, --oem 2 both, --oem 3 whatever is available (default). On Tesseract 4 and 5, use 1. The original engine is present only for legacy models.

Measuring, and using the output

This block runs with no Tesseract installed. It does the two jobs nobody skips: scoring accuracy, and turning the TSV into usable lines.

score_and_parse.py
import csv
import io


def edit_distance(a, b):
    prev = list(range(len(b) + 1))
    for i, ca in enumerate(a, 1):
        cur = [i]
        for j, cb in enumerate(b, 1):
            cur.append(min(prev[j] + 1,           # deletion
                           cur[j - 1] + 1,        # insertion
                           prev[j - 1] + (ca != cb)))   # substitution
        prev = cur
    return prev[-1]


def cer(truth, guess):
    return edit_distance(truth, guess) / max(len(truth), 1)


def wer(truth, guess):
    t, g = truth.split(), guess.split()
    return edit_distance(t, g) / max(len(t), 1)


print("what different OCR mistakes actually cost you")
print(f"{'truth':<26}{'engine output':<26}{'CER':>7}{'WER':>7}")
cases = [
    ("INVOICE NO 4471", "INVOICE NO 4471", "perfect"),
    ("INVOICE NO 4471", "INVOICE NO 447l", "one digit-letter swap"),
    ("INVOICE NO 4471", "INVOICENO 4471", "one lost space"),
    ("INVOICE NO 4471", "1NVO1CE NO 4471", "two digit-letter swaps"),
    ("TOTAL 1250.00", "TOTAL 125O.OO", "zero read as capital O"),
    ("TOTAL 1250.00", "T0TAL 1250,00", "decimal point read as comma"),
]
for truth, guess, _ in cases:
    print(f"{truth:<26}{guess:<26}{cer(truth, guess):>7.3f}{wer(truth, guess):>7.3f}")

print("\nrow 2 costs 0.067 CER and 0.333 WER, for a single wrong character.")
print("row 6 costs 0.154 CER, and turns 1250.00 into something no parser accepts.")
print("CER is an average. It does not know which characters carry money.\n")

# Tesseract's image_to_data returns TSV in exactly this shape. These few rows are
# hand-written so the file runs without the Tesseract binary installed.
TSV = """level\tpage_num\tblock_num\tpar_num\tline_num\tword_num\tleft\ttop\twidth\theight\tconf\ttext
1\t1\t0\t0\t0\t0\t0\t0\t640\t480\t-1\t
2\t1\t1\t0\t0\t0\t36\t40\t320\t22\t-1\t
3\t1\t1\t1\t0\t0\t36\t40\t320\t22\t-1\t
4\t1\t1\t1\t1\t0\t36\t40\t320\t22\t-1\t
5\t1\t1\t1\t1\t1\t36\t40\t96\t18\t96\tINVOICE
5\t1\t1\t1\t1\t2\t140\t40\t34\t18\t95\tNO
5\t1\t1\t1\t1\t3\t182\t40\t62\t18\t31\t447l
5\t1\t1\t1\t1\t4\t250\t44\t8\t10\t12\t.
5\t1\t2\t1\t1\t1\t36\t96\t80\t18\t93\tTOTAL
5\t1\t2\t1\t1\t2\t124\t96\t92\t18\t88\t1250.00
"""

rows = list(csv.DictReader(io.StringIO(TSV), delimiter="\t", quoting=csv.QUOTE_NONE))
words = [r for r in rows if r["level"] == "5" and r["text"].strip()]
print(f"rows returned: {len(rows)}   of which words: {len(words)}")
print("the other rows are page, block, paragraph and line containers, with conf = -1\n")

for r in words:
    flag = "  <-- low confidence" if int(r["conf"]) < 60 else ""
    print(f"  conf {int(r['conf']):3d}   box ({r['left']},{r['top']},"
          f"{r['width']},{r['height']})   {r['text']!r}{flag}")

kept = [r for r in words if int(r["conf"]) >= 60]
print(f"\ndropping anything under 60 removes {len(words) - len(kept)} of {len(words)} words")

lines = {}
for r in kept:
    key = (r["block_num"], r["par_num"], r["line_num"])
    lines.setdefault(key, []).append(r)
print("\nrebuilt lines, sorted by block then line then word position:")
for key in sorted(lines):
    text = " ".join(w["text"] for w in sorted(lines[key], key=lambda w: int(w["left"])))
    print(f"  block {key[0]} line {key[2]}:  {text!r}")

truth = "INVOICE NO 4471 TOTAL 1250.00"
got = " ".join(" ".join(w["text"] for w in sorted(lines[k], key=lambda w: int(w["left"])))
                for k in sorted(lines))
print(f"\ntruth: {truth!r}")
print(f"kept:  {got!r}")
print(f"CER {cer(truth, got):.3f}   WER {wer(truth, got):.3f}")
print("filtering removed a stray '.', but it also removed the invoice number")
Output
what different OCR mistakes actually cost you
truth                     engine output                 CER    WER
INVOICE NO 4471           INVOICE NO 4471             0.000  0.000
INVOICE NO 4471           INVOICE NO 447l             0.067  0.333
INVOICE NO 4471           INVOICENO 4471              0.067  0.667
INVOICE NO 4471           1NVO1CE NO 4471             0.133  0.333
TOTAL 1250.00             TOTAL 125O.OO               0.231  0.500
TOTAL 1250.00             T0TAL 1250,00               0.154  1.000

row 2 costs 0.067 CER and 0.333 WER, for a single wrong character.
row 6 costs 0.154 CER, and turns 1250.00 into something no parser accepts.
CER is an average. It does not know which characters carry money.

rows returned: 10   of which words: 6
the other rows are page, block, paragraph and line containers, with conf = -1

  conf  96   box (36,40,96,18)   'INVOICE'
  conf  95   box (140,40,34,18)   'NO'
  conf  31   box (182,40,62,18)   '447l'  <-- low confidence
  conf  12   box (250,44,8,10)   '.'  <-- low confidence
  conf  93   box (36,96,80,18)   'TOTAL'
  conf  88   box (124,96,92,18)   '1250.00'

dropping anything under 60 removes 2 of 6 words

rebuilt lines, sorted by block then line then word position:
  block 1 line 1:  'INVOICE NO'
  block 2 line 1:  'TOTAL 1250.00'

truth: 'INVOICE NO 4471 TOTAL 1250.00'
kept:  'INVOICE NO TOTAL 1250.00'
CER 0.172   WER 0.200
filtering removed a stray '.', but it also removed the invoice number

Reading the output carefully

The TSV has a hierarchy, and most rows are not words. level runs 1 to 5 for page, block, paragraph, line and word. Container rows carry conf = -1 and empty text. Filtering on level == 5 is the first line of every TSV parser people write.

block_num, par_num, line_num, word_num are the reading order Tesseract inferred. Sorting by them, then by left inside a line, is how you rebuild readable text with the boxes intact. Sorting by top alone breaks on any two-column page.

Confidence filtering is a genuine trade, not a free win. Dropping everything under 60 removed a spurious . — and also removed 447l, the invoice number Tesseract was unsure about. Overall CER went from what it would have been to 0.172, because a deleted word costs as much as a wrong one. A low-confidence word is a signal to review, not a signal to delete.

0.154 CER and 1.000 WER on the same string. The comma-for-dot error hits one character out of thirteen, and every word in the line contains an error. Which number you report changes the story completely, so report both, and report field-level exact-match accuracy for anything structured.

Common mistakes

Leaving --psm at the default for cropped regions. This is the top cause of empty output. Crops need 7 or 8.

Not installing the language pack. lang="hin" fails unless tesseract-ocr-hin is installed. pytesseract.get_languages() tells you what you actually have.

Passing a numpy array where a path or PIL image is expected. pytesseract accepts PIL images, file paths and numpy arrays, but a numpy array must be a valid image dtype. Convert with Image.fromarray(arr.astype("uint8")).

Calling Tesseract once per word in a loop. Each call spawns a process and reloads the model. For a page of 400 words that is 400 process launches. Call it once per page with image_to_data, or batch crops into one tall image.

Expecting reading order on a two-column page. Tesseract's block detection handles simple columns and fails on sidebars, headers and captions. That is a job for document layout analysis.

Feeding a colour phone photo directly. Tesseract binarises internally with Otsu, and we saw what Otsu does under a shadow. Deskew, adaptive-threshold and upscale first.

Try it yourself

Take one of your own scanned pages. Run it at --psm 3, then crop a single line and run that at --psm 3 and --psm 7. Score all three with the cer function above against text you type by hand. The gap between the two crop runs is usually larger than any preprocessing change you will make that day.

What to learn next

Researcher — Mathematics and papers.

What is actually inside Tesseract 5

Three stages, and only the third is neural.

Page layout analysis uses the Tab-Stop detection algorithm (Smith, 2009, Hybrid page layout analysis via tab-stop detection). Connected components are grouped into columns by locating consistent left and right tab stops, then into blocks, paragraphs and lines. It is a purely geometric algorithm with no learned component, which is precisely why it fails on decorative layouts and succeeds on scanned books.

Line and word finding uses baseline fitting and a fixed-pitch detector, then splits words at estimated space widths.

Recognition since 4.0 is an LSTM trained with CTC — the architecture in CRNN text recognition and CTC loss. Language models are shipped as .traineddata files containing the network, a unicharset, and optionally a dictionary and word bigram data used to rescore.

The --oem 0 legacy engine is the pre-4.0 segment-and-classify system described in Smith (2007), An overview of the Tesseract OCR engine, ICDAR. It is retained only for legacy .traineddata and should not be used for new work.

Confidence, and what it is not

Tesseract's per-word conf derives from the recogniser's character-level certainty combined with dictionary agreement. Two consequences worth knowing:

  • It is not calibrated. A conf of 80 does not mean 80% of such words are correct. Empirically, the distribution is heavily skewed toward high values, and the useful signal is mostly at the low end. If you need calibrated numbers, fit an isotonic or Platt mapping on a labelled sample of your documents.
  • Dictionary agreement inflates it. A wrong word that happens to be a dictionary word scores higher than a correct word that is not. On invoice numbers, PAN numbers and product codes — the fields that matter — this pushes exactly the wrong way. -c load_system_dawg=0 -c load_freq_dawg=0 disables the dictionaries and often improves accuracy on code-like fields.

Where Tesseract sits in 2026

Tesseract 5PaddleOCR (PP-OCRv5)docTRVLM readers
InstallSystem packagepip, plus PaddlePaddlepip, PyTorch or TFGPU, model download
Scanned printed pageGoodGoodGoodGood
Phone photo, angledWeakStrongStrongStrong
HandwritingVery weakModerateModerateStrong
Tables and layoutNonePP-StructureV3PartialStrong
Cost per pageCPU millisecondsCPU tens of msCPU tens of msGPU seconds
Hallucination riskNoneNoneNoneReal

Two honest positions follow. If your inputs are clean scanned pages and privacy or cost matters, Tesseract is still the right default and the field has not moved past it for that case. If your inputs are photographs, a modern detection-plus-recognition stack will beat it by a wide margin, and no amount of --psm tuning closes that gap.

Fine-tuning

Tesseract supports fine-tuning .traineddata on your own line images with lstmtraining, starting from an existing model. It is worth doing for a single consistent font — an old typewriter archive, a specific government form, a historical typeface — where a few thousand line images can cut CER substantially.

It is not worth doing to handle general variation. The tooling is awkward, requires building Tesseract's training binaries, and the resulting model is narrower than the one you started from. For general improvement, changing engines beats fine-tuning this one.

Benchmarks to be careful with

Published "Tesseract versus X" comparisons are frequently unfair in one of three ways: default --psm used on cropped inputs, no preprocessing applied, or the wrong --oem. When reading such a comparison, check those three settings before believing the numbers. When publishing one, state them.

What to learn next

What to learn next

These follow on from what you just read.

  • OCR and Document Vision

    TrOCR and transformer OCR

    TrOCR reads a text crop with an image transformer and writes the answer with a language decoder, which makes it strong on messy text and capable of inventing words that were never there.

  • OCR and Document Vision

    Handwriting recognition

    Handwriting is hard because every writer has a different alphabet, letters connect, and the same person writes differently twice — and the fix is mostly data, not architecture.

  • OCR and Document Vision

    Document layout analysis

    Layout analysis works out which parts of a page are headings, paragraphs, figures and tables, and in what order a human would read them.