Document AI
Document AI reads a scanned page the way a person does, using where text sits on the page as much as what the words say.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Document AI reads scanned pages, forms and PDFs, using the position of the text as much as the text itself.
The analogy you have already lived
Think about looking at a restaurant bill. You do not read it left to right like a story. Your eye goes straight to the number at the bottom right, because that is where totals live.
Now think about a form you filled in at a bank. "Father's Name" printed on the left, an empty box on the right. You knew the box belonged to that label because of where it sat, not because an arrow connected them.
Position carries meaning. Strip a document down to its words and you throw that meaning away.
Why it exists
Ordinary text tools read a document as one long stream of words. That works for a novel and fails for almost everything else.
Feed a two-column magazine page to a plain text extractor and you get the two columns woven together into nonsense. Feed it a table and the rows dissolve. Feed it an invoice and "Total" arrives with no idea which number belongs to it.
Meanwhile, this is where enormous amounts of real work sit. Invoices, insurance claims, medical forms, land records, purchase orders and government paperwork are all pages, not databases.
Document AI exists because these pages need to become data, and the layout is part of the data.
How it works
Three things are read together.
scanned page
|
+--> [ OCR ] -> the words, and a box around each one
|
+--> [ layout ] -> where each box sits, and what kind of thing it is
|
+--> [ image ] -> the look: lines, stamps, logos, handwriting
|
v
[ model reading all three ]
|
v
"Total = Rs 4,820", "Invoice No = INV-2291"OCR stands for optical character recognition — turning a picture of text into text you can search. It is the first step, and everything downstream inherits its mistakes.
The important addition is the box. OCR does not only tell you the word; it tells you where the word was. Feeding those coordinates into the model is the whole idea behind modern document AI.
Where you have already used one
- Depositing a cheque by photographing it in a banking app.
- Scanner apps that turn a photo of a page into a searchable PDF.
- Expense apps that read a receipt total and date automatically.
- Government portals that read details straight off an uploaded ID card.
- Search inside a scanned book on an archive site.
What is honestly hard here
Bad inputs. Real documents are photographed at an angle, in poor light, with a thumb over one corner. A crease across a line can lose it entirely.
Handwriting. Printed text is largely solved. Handwriting is not, especially across the many scripts used in India.
Tables. Merged cells, rows that continue onto the next page, and headers that repeat all break naive extraction.
The stakes. A wrong digit in an invoice total is a real financial error. A confident wrong answer here causes immediate harm. So a confidence score and a human check are part of the design, not extras.
Remember this
- Document AI reads words plus their positions, not words alone.
- OCR turns picture-of-text into text, and everything downstream depends on it.
- Tables, handwriting and crumpled photographs are where systems still fail.
What to learn next
- Multimodal RAG — answering questions across a pile of documents.
- Cross-modal retrieval — finding the right page before you read it.
- Chat with your PDF — a full project built on these pieces.
Developer — Code and libraries.
Two things break document pipelines more often than model quality: reading order, and pairing labels with values. Both are pure geometry, and both are worth understanding before you reach for a model.
Setup
python3 --versionNo dependencies. The boxes below are the shape every OCR engine returns.
Reading order and field extraction
# What an OCR engine hands you: text plus a box, in no meaningful order.
# (x, y, width, height, text) with y growing downwards. The page is 600 wide.
BOXES = [
(40, 60, 220, 14, "Annual report 2025"),
(330, 60, 230, 14, "Section 2: Outlook"),
(40, 90, 220, 14, "Revenue grew by 18 percent"),
(330, 90, 230, 14, "We expect demand to hold"),
(40, 120, 220, 14, "across all four regions."),
(330, 120, 230, 14, "through the next two quarters."),
(40, 150, 220, 14, "Costs stayed flat."),
(330, 150, 230, 14, "Hiring will slow down."),
]
def naive_reading_order(boxes):
return [b[4] for b in sorted(boxes, key=lambda b: b[1])]
def column_reading_order(boxes, page_width=600, n_cols=2):
edges = [page_width * i / n_cols for i in range(n_cols + 1)]
out = []
for i in range(n_cols):
col = [b for b in boxes if edges[i] <= b[0] + b[2] / 2 < edges[i + 1]]
out += [b[4] for b in sorted(col, key=lambda b: b[1])]
return out
print("sorted top to bottom (what most people try first):")
print(" " + " ".join(naive_reading_order(BOXES)))
print()
print("column aware:")
print(" " + " ".join(column_reading_order(BOXES)))
# The other half of document AI: pull one field out by where it sits, not by wording.
FORM = [
(40, 40, 90, 14, "Invoice No"),
(150, 40, 80, 14, "INV-2291"),
(40, 70, 90, 14, "Date"),
(150, 70, 80, 14, "2026-03-11"),
(40, 100, 90, 14, "Total"),
(150, 100, 80, 14, "Rs 4,820"),
(330, 40, 90, 14, "Total pages"),
(450, 40, 80, 14, "3"),
]
def value_right_of(label, boxes, max_gap=80):
for x, y, w, h, text in boxes:
if text == label:
same_row = [b for b in boxes
if abs(b[1] - y) < h and 0 < b[0] - (x + w) < max_gap]
return same_row[0][4] if same_row else None
return None
print()
print("fields pulled out by position:")
for label in ("Invoice No", "Date", "Total", "Total pages"):
print(f"{label:>12}: {value_right_of(label, FORM)}")sorted top to bottom (what most people try first):
Annual report 2025 Section 2: Outlook Revenue grew by 18 percent We expect demand to hold across all four regions. through the next two quarters. Costs stayed flat. Hiring will slow down.
column aware:
Annual report 2025 Revenue grew by 18 percent across all four regions. Costs stayed flat. Section 2: Outlook We expect demand to hold through the next two quarters. Hiring will slow down.
fields pulled out by position:
Invoice No: INV-2291
Date: 2026-03-11
Total: Rs 4,820
Total pages: 3Read the first output line as a warning
"Annual report 2025 Section 2: Outlook Revenue grew by 18 percent We expect demand to hold..." — the two columns are shuffled together.
Send that string to a language model and it will answer questions from it, fluently and wrongly. Nothing downstream can detect the damage, because the text is grammatical.
This is the most common silent failure in document pipelines. The extraction step reported success. The content is scrambled.
The column-aware version reads correctly down the left column, then down the right. Ten lines of geometry fixed what no amount of prompt engineering would have.
Line by line, the parts that are not obvious
sorted(boxes, key=lambda b: b[1]) sorts by the top y coordinate. On a single-column page this is correct. On anything else it interleaves.
b[0] + b[2] / 2 is the box's horizontal centre. Assigning boxes to columns by their centre rather than their left edge handles boxes that slightly overhang a boundary.
abs(b[1] - y) < h accepts anything within one line height vertically. Requiring exact equality fails immediately on real OCR, where a value sits one or two pixels off its label.
0 < b[0] - (x + w) < max_gap requires the candidate to start after the label ends, and within a maximum gap. Without the gap limit, "Total" would reach across the page and claim "3" from the right-hand column.
if text == label uses exact equality on purpose. Switch it to label in text and "Total" matches "Total pages" first, returning 3 as your invoice amount. That substring bug has shipped in production more than once.
What to use in real life
OCR. Tesseract via pytesseract is free and works well on clean printed text. easyocr and doctr are neural and handle harder images, at a larger download. For Indian scripts, check language pack support before committing to one.
pip install pytesseract pillow # plus the tesseract binary from your OS package managerimport pytesseract
from PIL import Image
data = pytesseract.image_to_data(Image.open("page.png"), output_type=pytesseract.Output.DICT)
for i, word in enumerate(data["text"]):
if word.strip() and int(data["conf"][i]) > 60:
print(data["left"][i], data["top"][i], data["conf"][i], word)No output block: it depends entirely on your image. The point is the shape of what comes back — text, position and a confidence score per word. That confidence column is the most under-used field in document AI, and it is what you route to a human review queue.
Layout models. microsoft/layoutlmv3-base embeds text, 2-D position and image patches together, around 500 MB. naver-clova-ix/donut-base skips OCR entirely and reads pixels straight into structured output, around 800 MB. Both run on CPU, slowly. Check current sizes on the model cards.
General VLMs. A modern vision-language model can answer questions about a page directly. Convenient, expensive per page, and weaker on small print than a dedicated OCR path. For high-volume extraction, the OCR-plus-layout route is usually cheaper and more auditable.
Common mistakes
Assuming PDFs contain text. Many are pure images. Try a text extractor first, fall back to OCR, and log which path each document took.
Ignoring skew. A page photographed at three degrees breaks every row-alignment rule in this file. Deskew before you do geometry.
Throwing away confidence scores. OCR tells you when it is unsure. Route low-confidence numeric fields to a human, and you convert silent errors into visible work.
Testing on clean scans only. Collect the worst twenty documents your users actually send. That set is your real test suite.
Try it yourself
Change value_right_of to use label in text instead of text == label and rerun. Total now returns 3. Then fix it properly by preferring exact matches and falling back to prefix matches only when nothing exact is found.
What to learn next
- Multimodal RAG — answering questions across a pile of documents.
- Cross-modal retrieval — finding the right page before you read it.
- Chat with your PDF — a full project built on these pieces.
Researcher — Mathematics and papers.
The task family
Visually Rich Document Understanding covers several distinct problems that share an input:
- Document layout analysis — segmenting a page into regions with types (title, paragraph, table, figure).
- Key information extraction — mapping the page to a schema, as in invoice fields.
- Table structure recognition — recovering cell boundaries and row/column spans.
- Document VQA — free-form questions over a page (Mathew et al., 2020).
- Reading order prediction — a permutation task over text segments.
Each has separate benchmarks and separate failure modes, and treating them as one "read the document" problem is the most common architectural mistake.
Adding position to a language model
LayoutLM (Xu et al., 2019) added 2-D positional embeddings to BERT: each token receives learned embeddings for its normalised x0, y0, x1, y1 box coordinates, summed with the usual token and 1-D position embeddings. LayoutLMv2 added visual patch tokens and a spatially-aware attention bias. LayoutLMv3 (Huang et al., 2022) unified the pretraining objectives — masked language modelling, masked image modelling, and word-patch alignment — and dropped the CNN backbone in favour of linear patch projection.
The pretraining corpus for this family is IIT-CDIP, roughly 11 million scanned pages, which matters because it is genuinely noisy scanned material rather than clean digital documents.
An alternative encoding is used by TILT and by Pix2Struct: relative 2-D attention biases rather than absolute coordinate embeddings. Relative encodings transfer better across page sizes and resolutions.
OCR-free approaches
Donut (Kim et al., 2021) removes OCR entirely. A Swin encoder reads the page image; a BART decoder emits a structured sequence directly, trained with a synthetic-document generator for pretraining. Removing OCR removes error propagation and language-coverage limits, and it makes the model's failures harder to attribute.
Pix2Struct (Lee et al., 2022) pretrains on screenshot-to-simplified-HTML, with variable-resolution inputs that preserve aspect ratio. It transfers well to charts and UIs, where no OCR-friendly text layer exists.
The trade-off is stable across evaluations: OCR-based pipelines are more auditable and better on dense small text; OCR-free models are more robust to unusual layouts and scripts, and harder to debug.
Error propagation, quantified
For a pipeline of OCR then extraction, end-to-end accuracy is bounded above by OCR accuracy on the tokens that matter. Character error rate of 2 percent sounds acceptable and is not: a 12-character invoice number has roughly a 1 - 0.98^12, about 21 percent, chance of containing at least one wrong character under an independence assumption.
Errors are not independent in practice — they cluster on poor-quality regions — which makes the picture better in some documents and much worse in others. The operational consequence is the same: measure field-level exact-match accuracy, never character accuracy, when a field must be correct in full.
Reading order
Reading order is a permutation prediction problem, and it is genuinely hard on complex layouts. XY-cut (Nagy and Seth, 1984) recursively splits the page along whitespace gaps and remains a strong baseline for clean single- and multi-column pages. It fails on overlapping regions and floating elements.
LayoutReader (Wang, Xu and Cui, 2021) frames it as sequence-to-sequence over text segments, trained on ReadingBank — half a million pages with reading order derived from Word document XML. That derivation is the interesting part: the ground truth comes from the authoring format, not from human annotation.
Tables
Table recognition splits into detection (finding the table) and structure recognition (recovering cells). PubTabNet (Zhong et al., 2019) and FinTabNet provide large annotated sets, and TEDS (Tree-Edit-Distance-based Similarity) is the standard metric, comparing predicted and ground-truth HTML trees rather than cell text alone.
Merged cells, borderless tables and multi-page tables remain the hard cases. Reported accuracies on clean digital PDFs do not transfer to photographed pages.
Retrieval over documents
ColPali (Faysse et al., 2024) indexes document page images directly with a vision-language model and late-interaction scoring, skipping OCR and chunking altogether. On the ViDoRe benchmark it substantially outperforms OCR-plus-text-embedding pipelines on visually rich pages such as charts and infographics, at the cost of a much larger index — multiple vectors per page rather than one per chunk. This is the most relevant recent development for anyone building multimodal RAG over PDFs.
Benchmarks
- FUNSD (Jaume et al., 2019): 199 noisy scanned forms, entity labelling and linking. Small, and a genuine stress test.
- CORD: Indonesian receipts with a nested field schema.
- DocVQA (Mathew et al., 2020): 50,000 questions over 12,000 industry document images, scored with ANLS, a normalised edit-distance similarity that tolerates minor OCR differences.
- PubLayNet (Zhong et al., 2019): 360,000 automatically annotated pages for layout analysis.
- ChartQA and InfographicVQA: reasoning over rendered data graphics, where OCR is necessary and far from sufficient.
Papers
- Nagy and Seth, Hierarchical Representation of Optically Scanned Documents, ICPR 1984 — the XY-cut algorithm.
- Xu et al., LayoutLM, 2019 — arxiv.org/abs/1912.13318
- Zhong, Tang and Jimeno-Yepes, PubLayNet, 2019 — arxiv.org/abs/1908.07836
- Mathew, Karatzas and Jawahar, DocVQA, 2020 — arxiv.org/abs/2007.00398
- Kim et al., OCR-free Document Understanding Transformer (Donut), 2021 — arxiv.org/abs/2111.15664
- Wang, Xu and Cui, LayoutReader, 2021 — arxiv.org/abs/2108.11591
- Huang et al., LayoutLMv3, 2022 — arxiv.org/abs/2204.08387
- Lee et al., Pix2Struct, 2022 — arxiv.org/abs/2210.03347
- Faysse et al., ColPali, 2024 — arxiv.org/abs/2407.01449
What to learn next
- Multimodal RAG — answering questions across a pile of documents.
- Cross-modal retrieval — finding the right page before you read it.
- Chat with your PDF — a full project built on these pieces.