OCR and Document Vision

Document layout analysis

Layout analysis works out which parts of a page are headings, paragraphs, figures and tables, and in what order a human would read them.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How the old method works
  5. Where it fails
  6. Where you have already seen it
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Layout analysis finds the parts of a page and works out the order you would read them in.

The analogy you have already lived

Think of opening a newspaper. Before reading a word, you already know the big line at the top is the headline. You know the block on the left continues at the top of the next column, not on the right.

Nobody taught you that with rules. You learned it from seeing thousands of pages. A computer given a plain grid of pixels knows none of it.

Why it exists

OCR gives you words and where they sit. That sounds like enough. It is not.

Take a two-column research paper. Sort the words top to bottom and you get the first line of the left column. Then the first line of the right column. Then the second line of the left. The text is complete and unreadable.

Now add a sidebar, a page number, a figure caption, and a table. The problem gets worse quickly.

Layout analysis answers three questions the words alone cannot.

What kind of thing is this block? Heading, paragraph, table, figure, caption, footer.

What belongs with what? This caption goes with that figure.

What comes after what? The reading order.

How the old method works

There is an idea from the 1980s that still works surprisingly well. It is called an XY-cut: cutting the page at the widest blank strip, over and over.

   look for a tall blank COLUMN running down the page
      found one  ->  split into left and right, do each side separately

   no blank column?
      look for a wide blank ROW
      found one  ->  split into top and bottom, do each separately

   no gap wide enough?  ->  this is one block. Report it.

Following the splits in order gives you the reading order for free. Left before right, top before bottom, at every level.

We run this in the code below and it recovers a two-column page exactly.

Where it fails

The same code fails the moment a figure spans both columns. There is now no blank column running the full height, so the gutter between the columns is invisible.

The cut goes horizontal instead, and glues both columns into one wide block. The whole page collapses.

That single weakness is why the field moved to trained models. A model that has seen a hundred thousand pages recognises a two-column layout by its look, not by its gaps.

Where you have already seen it

  • A PDF reader letting you select the left column without grabbing the right.
  • Reflowing a paper to fit your phone screen.
  • Screen readers reading a page aloud in a sensible order.
  • Tools that pull tables out of annual reports.

The honest part

Reading order is judged by humans, and humans disagree. Does a marginal note belong before or after the paragraph beside it? Does a page number come first or last?

So "correct" reading order is fuzzy at the edges, and benchmark scores hide that. For your own documents, write down the rule you want, then check against it.

Remember this

  • Layout analysis labels page parts and orders them for reading.
  • Cutting the page at the widest blank strip gets a long way, and breaks on wide figures.
  • Trained models handle real layouts; they learned what a page looks like.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy==1.26.4

Running a trained layout model means a download. Understanding what it replaced does not, and the classical algorithm is worth having in your hands, because it is still the right tool for clean, regular documents.

XY-cut, and the layout that defeats it

xy_cut.py
import numpy as np

W, H = 100, 60
# A two-column page. Each entry is (name, x0, y0, x1, y1) and the list is in
# the order a human reads it.
BLOCKS = [
    ("title",   6,  2, 94,  6),
    ("L-para1", 6,  10, 46, 22),
    ("L-para2", 6,  25, 46, 34),
    ("L-para3", 6,  37, 46, 50),
    ("R-para1", 54, 10, 94, 18),
    ("R-figure", 54, 21, 94, 33),
    ("R-caption", 54, 35, 94, 38),
    ("R-para2", 54, 41, 94, 50),
    ("footer",  6,  55, 94, 57),
]
TRUTH = ["title", "L-para1", "L-para2", "L-para3",
         "R-para1", "R-figure", "R-caption", "R-para2", "footer"]

page = np.zeros((H, W), bool)
for _, x0, y0, x1, y1 in BLOCKS:
    page[y0:y1, x0:x1] = True

print("the page (# is ink):")
for y in range(0, H, 2):
    print("  " + "".join("#" if page[y, x] else "." for x in range(0, W, 1)))

naive = [b[0] for b in sorted(BLOCKS, key=lambda b: (b[2], b[1]))]
print("\nreading order by sorting top-to-bottom then left-to-right:")
print("  ", naive)
print("   correct:", naive == TRUTH)


def xy_cut(mask, x0, y0, x1, y1, min_gap=3, depth=0):
    """Recursive XY-cut: split at the widest blank band, vertical splits first."""
    sub = mask[y0:y1, x0:x1]
    if sub.size == 0 or not sub.any():
        return []

    def bands(profile):
        """Runs of blank rows/columns, as (start, length), longest first."""
        runs, start = [], None
        for i, empty in enumerate(np.append(~profile, False)):
            if empty and start is None:
                start = i
            elif not empty and start is not None:
                if start > 0 and i < len(profile):        # ignore outer margins
                    runs.append((start, i - start))
                start = None
        return sorted(runs, key=lambda r: -r[1])

    vgaps = bands(sub.any(axis=0))                        # blank COLUMNS -> a gutter
    if vgaps and vgaps[0][1] >= min_gap:
        s, n = vgaps[0]
        return (xy_cut(mask, x0, y0, x0 + s, y1, min_gap, depth + 1)
                + xy_cut(mask, x0 + s + n, y0, x1, y1, min_gap, depth + 1))

    hgaps = bands(sub.any(axis=1))                        # blank ROWS -> paragraph break
    if hgaps and hgaps[0][1] >= 2:
        s, n = hgaps[0]
        return (xy_cut(mask, x0, y0, x1, y0 + s, min_gap, depth + 1)
                + xy_cut(mask, x0, y0 + s + n, x1, y1, min_gap, depth + 1))

    ys, xs = np.nonzero(sub)
    return [(x0 + xs.min(), y0 + ys.min(), x0 + xs.max() + 1, y0 + ys.max() + 1)]


regions = xy_cut(page, 0, 0, W, H)
name_of = {(x0, y0, x1, y1): n for n, x0, y0, x1, y1 in BLOCKS}
found = [name_of.get(r, f"unmatched{r}") for r in regions]
print("\nreading order from a recursive XY-cut:")
print("  ", found)
print("   correct:", found == TRUTH)

print("\nwhat the cut did, step by step:")
print("   1. found a blank column band and split the page into two columns")
print("   2. inside each column, split at the widest blank row band")
print("   3. repeated until no gap was wide enough, then emitted the block")

print("\nwhere it breaks: put a wide figure across both columns")
page2 = page.copy()
page2[21:33, 6:94] = True                                 # the figure now spans the gutter
regions2 = xy_cut(page2, 0, 0, W, H)
print("   blocks recovered:", len(regions2), "instead of", len(BLOCKS))
for r in regions2:
    print(f"     x {r[0]:>3}-{r[2]:<3} y {r[1]:>3}-{r[3]:<3}  width {r[2] - r[0]}")
print("   no blank column now runs the full height, so the gutter is invisible")
print("   whole bands of both columns get glued into one region")
Output
the page (# is ink):
  ....................................................................................................
  ......########################################################################################......
  ......########################################################################################......
  ....................................................................................................
  ....................................................................................................
  ......########################################........########################################......
  ......########################################........########################################......
  ......########################################........########################################......
  ......########################################........########################################......
  ......########################################......................................................
  ......########################################......................................................
  ......................................................########################################......
  ......................................................########################################......
  ......########################################........########################################......
  ......########################################........########################################......
  ......########################################........########################################......
  ......########################################........########################################......
  ....................................................................................................
  ......................................................########################################......
  ......########################################......................................................
  ......########################################......................................................
  ......########################################........########################################......
  ......########################################........########################################......
  ......########################################........########################################......
  ......########################################........########################################......
  ....................................................................................................
  ....................................................................................................
  ....................................................................................................
  ......########################################################################################......
  ....................................................................................................

reading order by sorting top-to-bottom then left-to-right:
   ['title', 'L-para1', 'R-para1', 'R-figure', 'L-para2', 'R-caption', 'L-para3', 'R-para2', 'footer']
   correct: False

reading order from a recursive XY-cut:
   ['title', 'L-para1', 'L-para2', 'L-para3', 'R-para1', 'R-figure', 'R-caption', 'R-para2', 'footer']
   correct: True

what the cut did, step by step:
   1. found a blank column band and split the page into two columns
   2. inside each column, split at the widest blank row band
   3. repeated until no gap was wide enough, then emitted the block

where it breaks: put a wide figure across both columns
   blocks recovered: 3 instead of 9
     x   6-94  y   2-6    width 88
     x   6-94  y  10-50   width 88
     x   6-94  y  55-57   width 88
   no blank column now runs the full height, so the gutter is invisible
   whole bands of both columns get glued into one region

Reading the output carefully

The naive sort produced title, L-para1, R-para1, R-figure, L-para2, ... It jumps between columns after every block. Extract text in that order and every sentence is broken. This is the single most common bug in home-grown PDF extractors, and it is silent — the output looks like text.

The XY-cut recovered the exact human order. Note that it was never told there are two columns. The vertical gap between x=46 and x=54 was the widest blank band, so it split there first, and left-before-right fell out of the recursion.

The recursion order is the reading order. That is the elegance of XY-cut: you get segmentation and ordering from one procedure, with no separate ordering model. Reverse the two recursive calls and you get a right-to-left language's order.

The failure case collapses to width 88 regions. Once the figure bridges the gutter, no blank column survives the full height of the page. The algorithm falls through to a horizontal cut and returns three full-width bands — the middle one containing both columns interleaved. Nine blocks became three, and the text order inside that middle band is now the naive one.

min_gap=3 is the whole tuning surface. Too high and real columns are missed. Too low and the space between two words becomes a column boundary. On real pages this number depends on resolution, font size and language, which is precisely the fragility that pushed the field to learned models.

What to use on real documents

ToolApproachGood for
XY-cut / projection profilesGeometricClean scans, regular multi-column, no downloads
Tesseract's own block detectionTab-stop geometryWhat you already get from Tesseract
DocLayout-YOLOYOLO detector, DocSynth-300K pretrainingFast, general, real-time layout boxes
Table TransformerDETR-style detectorTables specifically
PP-StructureV3Full pipeline: layout, tables, formulas, orderEnd-to-end document parsing
VLM page readersGenerate structured markup directlyComplex layouts, at GPU cost

The learned detectors treat layout as ordinary object detection over categories like title, text, list, table, figure. That is the whole idea: a page is a scene, and paragraphs are objects.

Common mistakes

Trusting the order a PDF text layer gives you. PDF stores drawing operations, not reading order. Text can be emitted in any sequence the generator chose. pdfplumber and PyMuPDF return it faithfully, which is not the same as correctly.

Sorting boxes by top alone. Two words on the same visual line rarely have the same top to the pixel. Bucket into lines with a tolerance of about half a line height, then sort within the bucket.

Ignoring headers, footers and page numbers. They repeat on every page and pollute downstream text badly. Detect and strip them by looking for blocks that appear at the same position across many pages.

Training on PubLayNet and deploying on business documents. PubLayNet is auto-annotated from PubMed and is overwhelmingly two-column scientific articles. DocLayNet is human-annotated across finance, law, patents, manuals and tenders, and transfers far better to real-world work.

Treating a caption as a paragraph. A figure caption that flows into body text corrupts a whole section. Captions are a distinct class in both DocLayNet and PubLayNet for good reason.

Try it yourself

Add a third column to BLOCKS and rerun. The XY-cut handles it with no code change, because the recursion finds the next-widest gutter on its own. Then narrow one gutter to two pixels and watch it fail — that is min_gap biting.

What to learn next

Researcher — Mathematics and papers.

The classical formulation

XY-cut (Nagy and Seth, 1984) treats layout as a recursive binary partition. At each node, compute horizontal and vertical projection profiles over the binarised region, find maximal zero-valued runs (valleys), and split at the widest valley exceeding a threshold. The result is the X-Y tree, whose in-order traversal is the reading order.

Two structural limitations follow directly from the definition.

  • It can only produce rectangular, recursively separable layouts. Any layout requiring a non-rectangular cut — an L-shaped text region flowing around a figure — is unrepresentable rather than difficult.
  • It is not robust to a single spanning element, as the code demonstrates. One full-width figure destroys the column structure of an entire page.

Docstrum (O'Gorman, TPAMI 1993) takes a bottom-up route instead, computing nearest-neighbour angle and distance histograms over connected components to estimate text orientation, within-line and between-line spacing, then clustering into lines and blocks. It handles skew and irregular layouts better and is slower.

Voronoi-based segmentation (Kise et al., 1998) builds an area Voronoi diagram over connected components and removes edges by a distance and area-ratio criterion. It remains competitive on degraded historical documents.

Tesseract's tab-stop detection (Smith, ICDAR 2009) is the survivor in production, and is what runs when you call Tesseract with --psm 3.

The learned formulation

Layout became an object detection problem once large annotated corpora existed.

PubLayNet (Zhong et al., ICDAR 2019) — 360k page images auto-annotated by matching PubMed Central XML to rendered PDFs. Five classes: text, title, list, table, figure. Its scale is its strength and its provenance is its weakness: annotations inherit XML errors, and the domain is almost entirely two-column biomedical articles.

DocLayNet (Pfitzmann et al., KDD 2022) — 80,863 pages, human-annotated, 11 classes, across six domains chosen for layout diversity: financial reports, scientific articles, laws and regulations, manuals, patents, government tenders. It reports inter-annotator agreement per class, which is unusual and valuable: agreement on text and table is high, and on section-header versus title it is not. That disagreement bounds achievable accuracy.

DocLayout-YOLO (Zhao et al., 2024, arXiv 2410.12628) is a strong current open detector: a YOLO-v10 base, pretrained on a synthetic DocSynth-300K corpus generated by a Mesh-candidate BestFit layout algorithm, with a global-to-local receptive field module for the wide scale range document elements span. It gives real-time layout boxes on CPU-class hardware.

Reading order as its own problem

Detection gives boxes and classes. Order is separate, and is where systems still diverge.

Three approaches:

  • Geometric rules — XY-cut over detected boxes rather than over pixels, which is far more robust than over raw ink because the boxes already suppress noise.
  • Sequence prediction — LayoutReader (Wang et al., EMNLP 2020) frames it as seq2seq over spatial tokens, trained on ReadingBank, 500K pages with order derived from Word document XML.
  • Generative — a VLM emits the page as ordered Markdown, with order implicit in the output.

The honest caveat is that ground-truth order is a human convention. ReadingBank's order comes from the underlying document's element order, which is not always what a reader would choose for marginalia, footnotes or multi-panel figures. Metrics such as ARD (average relative distance) and BLEU over the block sequence are useful for comparison and should not be read as truth.

Evaluation

Layout detection is scored with COCO-style mean average precision at IoU thresholds. Two caveats specific to documents:

  • Class confusion dominates the error budget more than localisation does. A section-header predicted as title costs full AP for both classes while being visually indistinguishable. Report a confusion matrix, not only mAP.
  • Large-object bias. A page has few, large regions. An IoU of 0.5 on a full-column paragraph still leaves a badly cropped block, so evaluate at IoU 0.75 and above for anything feeding text extraction.

Downstream, the metric that actually matters is end-to-end: does the extracted text, in the extracted order, match a human transcription of the page. That number is usually considerably worse than the layout mAP suggests.

What to learn next