LayoutLM and form understanding
LayoutLM reads a form the way you do, using where a word sits as well as what it says, which is what lets it find the invoice number on a layout it has never seen.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
LayoutLM understands a form by using both the words and where those words sit on the page.
The analogy you have already lived
Think of glancing at an electricity bill you have never seen before. You find the amount due in about two seconds.
You did not read the whole bill. You knew the big bold number near the bottom right, next to a word like "payable", is the amount. Position told you as much as the words did.
An ordinary language model reads a form as one long string and throws that position away. LayoutLM keeps it.
Why it exists
Take a plain text model and paste in a form's text. You get something like:
Invoice No 4471 Date 12/03/26 ... Total 1250.00That reads acceptably. Now take a two-column form, or a form where the label sits above the value instead of beside it. The text ordering scrambles, and the relationship between "Total" and "1250.00" is gone.
The words alone are ambiguous. 4471 might be an invoice number, a PIN code or a quantity. What settles it is that it sits immediately to the right of the words "Invoice No".
How position gets in
Every word arrives from OCR with a box: left, top, right, bottom. LayoutLM turns those four numbers into part of the word's representation, alongside its meaning.
word "4471"
box x from 345 to 470, y from 120 to 152
|
rescale every page to a 0 to 1000 grid, so a big scan
and a small scan look the same
|
the model sees: this word, at this place, near those wordsThe rescaling step matters more than it looks. A form scanned at two different qualities produces different pixel numbers for the same layout. Rescaling makes them identical.
What it is used for
The task is usually tagging — labelling each word with what it is.
Invoice -> O (nothing)
No -> O
4471 -> INVOICE_ID
Date -> O
12/03/26 -> DATE
Total -> O
1250.00 -> TOTALOnce every word is tagged, pulling the fields out is straightforward. This is the same shape of problem as named entity recognition, with position added.
Why this beats writing rules
You could write rules. "The number after the words 'Invoice No' is the invoice number."
Then a supplier changes their template. Then a different supplier writes "Bill Ref". Then one prints the label above the value instead of beside it. Rule sets for document extraction grow forever and break constantly.
A trained model learns the shape of a form. A layout it has never seen still works, as long as it resembles what it saw in training.
Where you have already seen it
- Expense apps reading a photographed receipt into fields.
- Banks processing loan application forms in bulk.
- Insurance claim processing from scanned paperwork.
- Government systems reading standard forms at scale.
The honest part
These models need labelled examples of your forms. A public checkpoint trained on English business documents will not do well on an Indian government form in Hindi.
Labelling a few hundred forms is unglamorous and unavoidable. Anyone promising zero-shot extraction on your specific documents is describing a demo.
Remember this
- LayoutLM uses what a word says and where it sits, together.
- Every box is rescaled to a 0 to 1000 grid, so page size stops mattering.
- The output is a tag per word, which you then assemble into fields.
What to learn next
- Document AI — the wider picture this section feeds into.
- Named entity recognition — the same tagging task without the positions.
- BERT — the encoder LayoutLM is built on.
Developer — Code and libraries.
Setup
pip install transformers==5.6.2 numpy==1.26.4The code below downloads the LayoutLMv3 tokeniser only, a few megabytes. That is enough to see the two things people get wrong: box normalisation, and label alignment across sub-word pieces.
Fine-tuning the full model needs microsoft/layoutlmv3-base (about 500 MB) and a labelled dataset. The preparation below is the same either way.
Preparing a form for LayoutLM
from transformers import AutoTokenizer
PAGE_W, PAGE_H = 1240, 1754 # an A4 page scanned at 150 DPI
TAGS = ["O", "B-INVOICE_ID", "B-DATE", "B-TOTAL"]
# What OCR gives you: a word, where it sat in pixels, and the tag you want.
OCR = [
("Invoice", 90, 120, 260, 152, "O"),
("No", 275, 120, 330, 152, "O"),
("4471", 345, 120, 470, 152, "B-INVOICE_ID"),
("Date", 880, 120, 980, 152, "O"),
("12/03/26", 995, 120, 1160, 152, "B-DATE"),
("Total", 90, 980, 200, 1012, "O"),
("1250.00", 215, 980, 400, 1012, "B-TOTAL"),
]
def normalise(box, w, h):
"""LayoutLM wants every box scaled into a 0-1000 grid, whatever the page size."""
x0, y0, x1, y1 = box
return [int(1000 * x0 / w), int(1000 * y0 / h),
int(1000 * x1 / w), int(1000 * y1 / h)]
words = [r[0] for r in OCR]
pixel_boxes = [list(r[1:5]) for r in OCR]
boxes = [normalise(b, PAGE_W, PAGE_H) for b in pixel_boxes]
tag_ids = [TAGS.index(r[5]) for r in OCR]
print(f"page is {PAGE_W} x {PAGE_H} pixels; every box is rescaled to a 0-1000 grid\n")
print(f"{'word':<10}{'pixels':<28}{'0-1000 grid':<26}{'tag'}")
for w, pb, nb, r in zip(words, pixel_boxes, boxes, OCR):
print(f"{w:<10}{str(pb):<28}{str(nb):<26}{r[5]}")
print("\nrescaling is what makes a 150 DPI scan and a 300 DPI scan look alike to")
print("the model. Skip it and every position it learned during training is wrong.\n")
tok = AutoTokenizer.from_pretrained("microsoft/layoutlmv3-base")
print("tokeniser:", type(tok).__name__)
enc = tok(text=words, boxes=boxes, word_labels=tag_ids)
print("\nhow words become tokens, what box each carries, and what it is trained on:")
print(f"{'i':<4}{'token':<12}{'box':<26}{'from word':<12}{'training label'}")
for i, (tid, box, wid, lab) in enumerate(
zip(enc["input_ids"], enc["bbox"], enc.word_ids(), enc["labels"])):
src = words[wid] if wid is not None else "-"
shown = TAGS[lab] if lab != -100 else "-100 (ignored)"
print(f"{i:<4}{tok.decode([tid])!r:<12}{str(box):<26}{src:<12}{shown}")
print(f"\n{len(words)} words became {len(enc['input_ids'])} tokens")
print("every piece of a split word carries the SAME box")
print("only the FIRST piece gets a label; the rest are -100 so the loss skips them")
print("<s> and </s> get the box [0, 0, 0, 0], since they sit nowhere on the page")
print("\nwhy position is real information: one word, two places on the page")
for y in (120, 980):
print(f" 'Total' at pixel y={y:<5} -> grid box {normalise([90, y, 200, y + 32], PAGE_W, PAGE_H)}")
print(" same text, different 2-D embedding, so a different meaning to the model")page is 1240 x 1754 pixels; every box is rescaled to a 0-1000 grid word pixels 0-1000 grid tag Invoice [90, 120, 260, 152] [72, 68, 209, 86] O No [275, 120, 330, 152] [221, 68, 266, 86] O 4471 [345, 120, 470, 152] [278, 68, 379, 86] B-INVOICE_ID Date [880, 120, 980, 152] [709, 68, 790, 86] O 12/03/26 [995, 120, 1160, 152] [802, 68, 935, 86] B-DATE Total [90, 980, 200, 1012] [72, 558, 161, 576] O 1250.00 [215, 980, 400, 1012] [173, 558, 322, 576] B-TOTAL rescaling is what makes a 150 DPI scan and a 300 DPI scan look alike to the model. Skip it and every position it learned during training is wrong. tokeniser: LayoutLMv3Tokenizer how words become tokens, what box each carries, and what it is trained on: i token box from word training label 0 '<s>' [0, 0, 0, 0] - -100 (ignored) 1 ' Inv' [72, 68, 209, 86] Invoice O 2 'oice' [72, 68, 209, 86] Invoice -100 (ignored) 3 ' No' [221, 68, 266, 86] No O 4 ' 4' [278, 68, 379, 86] 4471 B-INVOICE_ID 5 '471' [278, 68, 379, 86] 4471 -100 (ignored) 6 ' Date' [709, 68, 790, 86] Date O 7 ' 12' [802, 68, 935, 86] 12/03/26 B-DATE 8 '/' [802, 68, 935, 86] 12/03/26 -100 (ignored) 9 '03' [802, 68, 935, 86] 12/03/26 -100 (ignored) 10 '/' [802, 68, 935, 86] 12/03/26 -100 (ignored) 11 '26' [802, 68, 935, 86] 12/03/26 -100 (ignored) 12 ' Total' [72, 558, 161, 576] Total O 13 ' 12' [173, 558, 322, 576] 1250.00 B-TOTAL 14 '50' [173, 558, 322, 576] 1250.00 -100 (ignored) 15 '.' [173, 558, 322, 576] 1250.00 -100 (ignored) 16 '00' [173, 558, 322, 576] 1250.00 -100 (ignored) 17 '</s>' [0, 0, 0, 0] - -100 (ignored) 7 words became 18 tokens every piece of a split word carries the SAME box only the FIRST piece gets a label; the rest are -100 so the loss skips them <s> and </s> get the box [0, 0, 0, 0], since they sit nowhere on the page why position is real information: one word, two places on the page 'Total' at pixel y=120 -> grid box [72, 68, 161, 86] 'Total' at pixel y=980 -> grid box [72, 558, 161, 576] same text, different 2-D embedding, so a different meaning to the model
Reading the output carefully
[345, 120, 470, 152] becomes [278, 68, 379, 86]. Same box, expressed as a fraction of the page times 1000. Rescan the page at 300 DPI and the pixel numbers double while the grid numbers stay put. That invariance is the entire reason for the convention, and forgetting it is the most common LayoutLM bug — the model runs, produces plausible tags, and is quietly wrong.
Seven words became eighteen tokens. Invoice split into ' Inv' and 'oice'; 12/03/26 split into five pieces. This is ordinary tokenization and it forces two decisions.
Every piece inherits the whole word's box. ' 12', '/', '03', '/', '26' all carry [802, 68, 935, 86]. The model has no sub-word position information, and does not need it — the word is the spatial unit.
Only the first piece gets a label. ' 4' is tagged B-INVOICE_ID; '471' is -100. -100 is PyTorch's ignore index for cross-entropy, so no loss is computed there. Passing word_labels makes the tokeniser do this alignment for you, which is safer than writing it yourself.
At prediction time you reverse this. Read the tag off the first token of each word, using word_ids() to map back. Averaging over pieces or taking the last piece both cause subtle errors on long numbers, which is exactly where accuracy matters.
<s> and </s> carry [0, 0, 0, 0]. A required convention. Giving them a real box places phantom text at the page corner and measurably degrades results.
Fine-tuning, in outline
from transformers import AutoModelForTokenClassification
model = AutoModelForTokenClassification.from_pretrained(
"microsoft/layoutlmv3-base", num_labels=len(TAGS))
# then a normal token-classification training loop over batches of
# input_ids, bbox, attention_mask, pixel_values and labelsLayoutLMv3 also takes pixel_values — the page image itself — through LayoutLMv3Processor. Set apply_ocr=False when you already have words and boxes from your own OCR, which you almost always should, since the processor's built-in OCR calls Tesseract with default settings.
Common mistakes
Forgetting to normalise, or normalising by the wrong page size. If your OCR ran on a resized image, normalise by that image's size, not the original file's.
Boxes outside 0 to 1000. Any value above 1000 indexes past the position embedding table and raises an index error, or worse, silently wraps in some implementations. Clamp after normalising.
Left and top greater than right and bottom. Some OCR engines emit rotated or reversed boxes. Sort the coordinates before use.
Aligning labels to the wrong sub-token. Use word_ids(). Writing your own alignment against the raw string breaks on punctuation and on any language with no spaces.
Expecting zero-shot field extraction. LayoutLMv3-base is pretrained, not task-trained. It knows nothing about your fields until you fine-tune it on labelled examples.
Feeding an unsorted word list. The model has 1-D position embeddings as well as 2-D ones, so reading order still matters. Sort words into reading order first, using document layout analysis.
Try it yourself
Move Total 1250.00 to the top of the page by changing its y from 980 to 200, and re-run. Only the grid boxes change. Now imagine that difference propagating through a trained model — a number at the top of a page and the same number at the bottom are different features, which is exactly what a text-only model cannot represent.
What to learn next
- Document AI — the wider picture this section feeds into.
- Named entity recognition — the same tagging task without the positions.
- BERT — the encoder LayoutLM is built on.
Researcher — Mathematics and papers.
The family
LayoutLM (Xu et al., KDD 2020) added 2-D position embeddings to BERT. Each word contributes four learned embeddings — for $x_0$, $y_0$, $x_1$, $y_1$ on the 0–1000 grid — plus width and height, summed into the input alongside the token and 1-D position embeddings. The image was used only as a separate per-token region feature, added late.
LayoutLMv2 (Xu et al., ACL 2021) brought the visual stream into the transformer as extra tokens from a CNN backbone, and added a spatial-aware self-attention bias: attention logits gain learned terms depending on the relative $x$ and $y$ offsets between two tokens. That relative bias is what lets the model express "the value is to the right of the label" without memorising absolute positions.
LayoutLMv3 (Huang et al., ACM MM 2022) removed the CNN entirely, using linear patch embeddings as in a vision transformer, and unified the pretraining objectives: masked language modelling, masked image modelling, and word-patch alignment (predicting whether a token's corresponding image patch was masked). Removing the CNN backbone is what makes it a single clean architecture and simplifies fine-tuning considerably.
Alternatives worth knowing. LiLT (Wang et al., ACL 2022) decouples the layout stream from the text stream so the layout model can be paired with any monolingual encoder — the practical route to a non-English form model without pretraining from scratch. Donut skips OCR entirely and generates structured output, trading boxes and provenance for end-to-end training.
The 0–1000 grid
Normalisation to a fixed integer grid serves three purposes:
- Resolution invariance. The same layout at 150 and 300 DPI produces identical inputs.
- A bounded embedding table. Position embeddings are a lookup of size 1001 per axis. Continuous coordinates would need a different mechanism.
- Quantisation as regularisation. A grid cell on an A4 page at 1000 divisions is about 0.21 mm wide, well below any meaningful layout distinction, so information loss is nil while spurious precision is discarded.
Benchmarks, and what they hide
- FUNSD (Jaume et al., ICDAR-OST 2019) — 199 scanned forms, entity labelling and entity linking. Small enough that reported F1 varies by a point or more between runs on seed alone; single-run comparisons should be treated sceptically.
- CORD — 1,000 Indonesian receipts, 30 field types.
- SROIE — receipt information extraction, four fields.
- XFUND — FUNSD extended to seven languages, the standard multilingual form benchmark.
- DocVQA — question answering over document images, a different framing that measures the same underlying capability.
Two structural cautions. All of these are small, and all are dominated by a few document templates, so a model can score well by memorising templates rather than learning layout. And entity linking — which value belongs to which label — is far harder than entity labelling, and headline numbers usually quote the easier one.
Where the field stands in 2026
The token-classification framing has been partly displaced by generative extraction: prompt a document VLM with a schema, receive JSON. It handles unseen layouts with no labelled data and no fine-tune, which is a genuine capability gain.
The trade is the one running through this whole section. A LayoutLM-style tagger can only output tokens that OCR actually found, and every extracted field comes with a box you can highlight for a human reviewer. A generative extractor can produce a well-formatted account number that appears nowhere on the page.
For low-stakes bulk extraction the generative route is usually the better engineering choice now. For anything audited — finance, insurance claims, medical records, legal filings — the ability to point at the pixels a value came from is not a nicety, and encoder taggers remain in production for that reason.
A common compromise is a VLM for extraction plus a grounding pass that locates each extracted string in the OCR output and rejects anything it cannot find. That recovers provenance and catches most invention, at the cost of one extra step.
What to learn next
- Document AI — the wider picture this section feeds into.
- Named entity recognition — the same tagging task without the positions.
- BERT — the encoder LayoutLM is built on.