OCR and Document Vision

TrOCR and transformer OCR

TrOCR reads a text crop with an image transformer and writes the answer with a language decoder, which makes it strong on messy text and capable of inventing words that were never there.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. What it gains
  6. What it costs
  7. Where you have already seen it
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

TrOCR looks at a strip of text with an image model, then writes out the answer with a language model.

The analogy you have already lived

Think of a friend reading a smudged address aloud from a courier slip. Half the letters are unclear. They still say the whole address correctly, because they know how addresses are written.

They are not reading letter by letter. They are guessing a sensible sentence that fits what they can see.

That is the strength of TrOCR, and it is also the danger. A confident guesser sometimes guesses.

Why it exists

The older reader we met in CRNN text recognition had no idea what words are. It read column by column and never looked back.

That hurts on hard text. If the middle of a word is blurred, a column-by-column reader has nothing to fall back on.

TrOCR bolts a language model onto the end. The image half decides what the pixels look like. The text half decides what a plausible next character is, given everything written so far.

How it works

   crop of one line
        |
   cut into small square patches, like tiles
        |
   [ image transformer ]   -> a set of numbers describing the tiles
        |
   [ text decoder ]        -> writes one character at a time,
        |                     looking at the tiles AND at what it wrote before
   "PLATFORM 12"

The two halves are pretrained separately on huge amounts of data. Then they are trained together on images of text.

What it gains

It handles bad input better. Blur, unusual fonts and handwriting all improve, because the language half fills gaps.

It needs no separate alphabet design. It writes normal text, including punctuation and mixed scripts.

It gets language knowledge for free. The decoder already knows English before it sees a single scan.

What it costs

It is slow. The older reader takes milliseconds on a laptop. This one takes a second or more, and prefers a GPU.

It can invent text. This is the important one. A decoder that writes plausible text will do so even when the picture holds nothing readable.

We show that happening for real in the code below. Given a crop of pure random noise, the model confidently returns a shopping-receipt phrase. No pixel in that image was a letter.

Where you have already seen it

  • Phone apps that read handwritten notes surprisingly well.
  • Document tools that handle old, faded typewriter pages.
  • Systems reading forms filled in by hand.

The honest part

Never use a generative reader alone for numbers that matter. Account numbers, amounts, dosages, identity numbers.

For those, use a reader that cannot invent characters, or check the output against a second method. A wrong bank account number that looks perfectly formatted is worse than a blank field.

Remember this

  • TrOCR is an image model that sees, plus a language model that writes.
  • The language half makes it better on messy text.
  • The same language half lets it invent text that was never in the picture.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install transformers==5.6.2 torch==2.5.1 pillow==11.0.0 numpy==1.26.4

This downloads about 250 MB of weights on first run. It runs on CPU — the small checkpoint is 62 million parameters, and one crop takes a second or two on a laptop.

Reading real crops, and then reading nothing

trocr_demo.py
import numpy as np
import torch
from PIL import Image, ImageDraw, ImageFont
from transformers import TrOCRProcessor, VisionEncoderDecoderModel

MODEL = "microsoft/trocr-small-printed"          # 62M params, CPU-friendly
processor = TrOCRProcessor.from_pretrained(MODEL)
model = VisionEncoderDecoderModel.from_pretrained(MODEL)
model.eval()

print("parameters (millions):", round(sum(p.numel() for p in model.parameters()) / 1e6, 1))
print("encoder:", model.config.encoder.model_type, " decoder:", model.config.decoder.model_type)
print("every crop is resized to:", processor.image_processor.size["height"], "x",
      processor.image_processor.size["width"], "pixels\n")


def strip(text, size=48, pad=12):
    """Render one line with Pillow's bundled font, so no font file is needed."""
    font = ImageFont.load_default(size=size)
    box = ImageDraw.Draw(Image.new("RGB", (8, 8))).textbbox((0, 0), text, font=font)
    img = Image.new("RGB", (box[2] - box[0] + 2 * pad, box[3] - box[1] + 2 * pad), "white")
    ImageDraw.Draw(img).text((pad - box[0], pad - box[1]), text, font=font, fill="black")
    return img


def read(img):
    pixel_values = processor(images=img, return_tensors="pt").pixel_values
    with torch.no_grad():
        ids = model.generate(pixel_values, max_new_tokens=24)
    return processor.batch_decode(ids, skip_special_tokens=True)[0]


print("one line in, one line out:")
for text in ["PLATFORM 12", "Invoice No 4471", "Total 1250.00", "hello world"]:
    img = strip(text)
    print(f"   drew {text!r:<18} size {img.size}   read {read(img)!r}")

print("\nthis checkpoint was fine-tuned on printed receipts, so it upper-cases")
print("everything and has learned receipt-shaped spacing. That is training data")
print("showing through, not a bug in the architecture.\n")

print("what happens when there is no text to read:")
blank = Image.new("RGB", (300, 60), "white")
print("   blank white crop      ->", repr(read(blank)))
noise = Image.fromarray(
    np.random.default_rng(0).integers(0, 255, (60, 300, 3), dtype=np.uint8))
print("   random noise crop     ->", repr(read(noise)))
print("   a detector that finds no text still gets an answer back")
Output
parameters (millions): 61.6
encoder: deit  decoder: trocr
every crop is resized to: 384 x 384 pixels

one line in, one line out:
   drew 'PLATFORM 12'      size (342, 59)   read 'PLATFORM 12'
   drew 'Invoice No 4471'  size (374, 59)   read 'INVOICE NO 4471'
   drew 'Total 1250.00'    size (323, 61)   read 'TOTAL 1250.00'
   drew 'hello world'      size (262, 61)   read 'HELLOWORLD'

this checkpoint was fine-tuned on printed receipts, so it upper-cases
everything and has learned receipt-shaped spacing. That is training data
showing through, not a bug in the architecture.

what happens when there is no text to read:
   blank white crop      -> '1'
   random noise crop     -> 'TOTAL AMOUNT FOR A FREE SANDWICH'
   a detector that finds no text still gets an answer back

Reading the output carefully

Three of four lines came back exactly right, and the fourth is informative. 'hello world' became 'HELLOWORLD', upper-cased and with the space dropped. This checkpoint is fine-tuned on the SROIE receipt dataset, which is upper-case printed receipts. The model learned the distribution of its training text, not only its shapes.

That is the most practical thing to take from the run. Choosing a TrOCR checkpoint is choosing a text domain, not only an image domain. Use trocr-base-handwritten for handwriting, trocr-base-printed for receipts, and trocr-base-stage1 as the starting point when you fine-tune on your own documents.

384 x 384 is the input size, and every crop is squashed to it. A wide receipt line is stretched vertically and squeezed horizontally. That is by design, since the encoder uses fixed patches. It does mean very long lines lose detail, so split them before feeding them in.

The blank crop returned '1' and the noise crop returned a receipt phrase. These two lines are the whole risk argument. An autoregressive decoder always produces a token sequence. It has no way to represent "there is nothing here".

Compare with CTC loss. A CTC model that sees no letters emits blank at every frame, and the collapse rule returns the empty string. The difference is structural, not a matter of training quality.

Guarding a generative reader

None of these are optional in a production system.

Score the output. model.generate(..., output_scores=True, return_dict_in_generate=True) gives per-token scores. Take the mean log probability and reject crops below a threshold you calibrate on your own data.

Constrain the alphabet where you can. For a digits-only field, mask the logits so only digit tokens are reachable. generate accepts bad_words_ids, and a custom LogitsProcessor gives full control.

Cross-check against a non-generative reader. Running a CTC recogniser alongside and flagging disagreements catches most invented output at modest cost.

Never send a crop your detector was unsure about. Hallucination is worst on inputs containing no text, and a permissive detector supplies exactly those.

Common mistakes

Feeding a whole page. TrOCR is a line model with no layout stage. A full page returns one garbled line. Detect lines first, as in text detection models.

Forgetting model.eval() and torch.no_grad(). Without them you build a graph for every crop and memory climbs steadily.

Running one crop at a time. Batch them. processor(images=[a, b, c], return_tensors="pt") gives a batch, and throughput improves several times over.

Expecting Tesseract-like speed. Measure on your own hardware first. The gap against a CRNN is roughly two orders of magnitude per crop on CPU.

Assuming the checkpoint matches your script. TrOCR's decoder vocabulary comes from RoBERTa for the English models. Devanagari, Tamil or Arabic need a checkpoint whose tokeniser covers those characters, and most public TrOCR checkpoints do not.

Try it yourself

Render strip("O0O0 Il1l") and read it. Then try the same string upper-cased and lower-cased. The confusions on that one crop tell you more about the model's priors than any benchmark table, and they are the exact confusions that corrupt account numbers.

What to learn next

Researcher — Mathematics and papers.

The architecture

TrOCR (Li et al., 2023, TrOCR: Transformer-based optical character recognition with pre-trained models, AAAI; arXiv 2109.10282) is a plain encoder-decoder with no OCR-specific machinery.

Encoder. A vision transformer over $16 \times 16$ patches of a $384 \times 384$ crop, giving 576 patch tokens plus a class token. Initialised from DeiT or BEiT rather than trained from scratch.

Decoder. A standard transformer decoder initialised from RoBERTa. It attends to the encoder tokens by cross-attention, and to its own prefix by causal self-attention. It emits BPE tokens, not characters.

The paper's claim is that no convolutional backbone and no CTC head are needed, and that the gain comes from the two pretrained halves. Both initialisations are load-bearing; the ablations show substantial drops when either is trained from scratch.

CTC against autoregressive decoding

The distinction goes deeper than accuracy numbers.

CTC headAutoregressive decoder
Output lengthBounded by $T$ framesUnbounded until an end token
Empty outputNatural (all blank)Needs an end token generated first
Output dependenciesNone, by constructionFull
Decoding costOne forward passOne pass per token
AlignmentMonotonic, enforcedUnconstrained
Failure modeDeletion, substitutionInsertion, repetition loops, invention

The conditional independence assumption in CTC is exactly what makes it unable to hallucinate. Every emitted character must be supported by some frame's posterior. Removing that assumption buys language modelling and sells safety.

The noise-crop result above is the empirical face of this. It is not a defect of TrOCR specifically; Donut and every subsequent VLM reader share it. Document-VLM evaluations increasingly report a hallucination rate alongside CER for this reason.

Where the field went next

Donut (Kim et al., ECCV 2022, OCR-free document understanding transformer) removed the OCR stage entirely, mapping a page image straight to structured JSON with a Swin encoder and a BART decoder. It trades boxes for end-to-end training.

Nougat (Blecher et al., 2023) applied the same recipe to academic PDFs, emitting Markdown with LaTeX maths. Its documented dominant failure is the repetition loop: the decoder falls into a cycle and repeats a phrase until the token limit.

General VLMs. By 2026 the strongest open document readers are vision-language models doing full-page transcription to Markdown or HTML, with tables and reading order handled inside the model. The open landscape splits into pipeline engines and VLM readers, and choosing between them is a choice about provenance and cost more than about raw accuracy.

Feature-level hybrids. SVTR-style recognisers keep a single-stage transformer over image patches but train with CTC. They keep the speed and the non-hallucination while dropping the RNN. PP-OCRv5's recogniser is in this family, and it is a reasonable answer for anyone who wants transformer features without generative risk.

Evaluation caution

Comparing TrOCR against a CRNN on CER alone favours TrOCR in a way that hides the trade. Two additions make the comparison honest.

  • Report insertion rate separately from substitutions and deletions. Hallucination shows up as insertions, and a single CER averages it away.
  • Include a negative set — crops with no text, or text in an unsupported script. A CTC model produces near-zero output length on these. A generative model does not, and that gap is the number which should drive a deployment decision.

What to learn next