OCR and Document Vision

Table extraction from documents

Getting a table out of a scan means finding the table, then recovering its rows, columns and merged cells — and the second half is where nearly everything fails.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. The two halves of the problem
  5. Two ways to find the grid
  6. Why the guessing breaks
  7. Where you have already seen it
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Table extraction turns a picture of a table into rows and columns you can put in a spreadsheet.

The analogy you have already lived

Think of copying a cricket scorecard from a newspaper into a notebook. You do not copy the numbers in the order your eye finds them. You look across a row, then move down.

Your eye reconstructs the grid, using the alignment of the numbers and the lines between them. Nothing on the page says "this is column three". You inferred it.

That inference is the whole job, and it is the part computers get wrong.

Why it exists

An invoice, a bank statement, a lab report, an annual report. The numbers you actually need live in tables, and a table is the one place where position is meaning.

Move a number from column two to column three and you have changed the fact. Losing a row breaks a total. Merging two cells changes a name.

So a table is not text with layout. It is data with a shape, and the shape must survive.

The two halves of the problem

Finding the table. Where on the page does the table start and end? This is ordinary detection and it is largely solved.

Recovering its structure. Which words belong to which row and column? This is the hard half.

   page  ->  [ find the table ]  ->  a box around it
                     |
             [ find the grid ]  ->  rows, columns, merged cells
                     |
             [ fill the cells ]  ->  each word into its cell
                     |
                a spreadsheet

Two ways to find the grid

If the table has printed lines, use them. Look for long straight ink runs, horizontal and vertical. Where they cross is where the cells are. This is exact and needs no guessing at all.

If it has no lines, guess from alignment. Words that share a top edge are probably one row. Words that share a left edge are probably one column. This is what the code below does, and it works on tidy tables.

Why the guessing breaks

Three things wreck it, and all three are common.

A cell that wraps onto two lines. "Steel rods" split across two lines looks like two rows.

A merged cell. A heading covering two columns has no boundary inside it, so the boundary is lost for the whole table.

Numbers aligned differently. Amounts aligned right, names aligned left. Guessing columns from left edges then fails for the amounts.

The code below demonstrates the first two failing, on a table that a moment earlier came out perfectly.

Where you have already seen it

  • Photographing a bank statement and getting transactions into an app.
  • Tools that pull financial tables out of company annual reports.
  • Converting a scanned price list into a spreadsheet.
  • Reading a lab report into a hospital system.

The honest part

Table extraction is the least reliable stage of document AI, and by some distance.

Everything else degrades gracefully. A misread word is one wrong word. A misrecovered table can shift a whole column of amounts by one position, and the result looks perfectly well-formed. That is the dangerous kind of wrong.

Any pipeline that reads money out of tables needs a check: totals that add up, columns that hold the type they should, row counts that match a stated count.

Remember this

  • Finding a table is easy. Recovering its rows and columns is not.
  • Printed lines give the grid exactly. Without them, you infer it from alignment.
  • Wrapped cells and merged cells break the inference, quietly.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install opencv-python==4.10.0.84 numpy==1.26.4

Trained table models need a download. The two classical methods do not, and you need both anyway: they are what you fall back to when the model gives an implausible grid.

Grid from alignment, and grid from ruling lines

table_grid.py
import cv2
import numpy as np

# What an OCR engine hands you: text plus a box. No table, no grid, no rows.
# (text, left, top, right, bottom)
WORDS = [
    ("Item",     20, 10,  70, 22), ("Qty",  200, 10, 240, 22), ("Price", 320, 10, 380, 22),
    ("Cement",   20, 40, 100, 52), ("12",   200, 40, 222, 52), ("380.00", 320, 40, 392, 52),
    ("Sand",     20, 70,  76, 82), ("4",    200, 70, 212, 82), ("120.50", 320, 70, 392, 82),
    ("Steel",    20, 100, 78, 112), ("30",  200, 100, 222, 112), ("2450.00", 320, 100, 404, 112),
]


def cluster(values, gap):
    """Group sorted 1-D positions into runs separated by more than `gap`."""
    groups = [[values[0]]]
    for v in sorted(values)[1:]:
        if v - groups[-1][-1] > gap:
            groups.append([v])
        else:
            groups[-1].append(v)
    return [sum(g) / len(g) for g in groups]


def build(words, row_gap=14, col_gap=40):
    rows = cluster(sorted((w[2] + w[4]) / 2 for w in words), row_gap)
    cols = cluster(sorted(w[1] for w in words), col_gap)
    grid = [["" for _ in cols] for _ in rows]
    for text, x0, y0, x1, y1 in words:
        r = int(np.argmin([abs((y0 + y1) / 2 - c) for c in rows]))
        c = int(np.argmin([abs(x0 - c) for c in cols]))
        grid[r][c] = (grid[r][c] + " " + text).strip()
    return rows, cols, grid


rows, cols, grid = build(WORDS)
print(f"{len(WORDS)} word boxes -> {len(rows)} rows x {len(cols)} columns")
print("row centres:   ", [round(r, 1) for r in rows])
print("column starts: ", [round(c, 1) for c in cols])
print()
for r in grid:
    print("   | " + " | ".join(f"{c:<9}" for c in r) + " |")

print("\nnow break it two ways that happen on every real invoice\n")

wrapped = [w for w in WORDS if w[0] != "Steel"] + [
    ("Steel", 20, 100, 78, 112), ("rods", 20, 116, 74, 128)]     # a cell wrapped to 2 lines
_, _, grid2 = build(wrapped)
print("1. one cell wraps onto a second line:")
for r in grid2:
    print("   | " + " | ".join(f"{c:<9}" for c in r) + " |")
print("   -> 'rods' became its own row, because rows are found by y position alone")

merged = [w for w in WORDS if w[0] not in ("Qty", "Price")] + [
    ("Quantity and price", 200, 10, 380, 22)]                    # a header spanning 2 columns
_, cols3, grid3 = build(merged)
print("\n2. one header cell spans two columns:")
for r in grid3:
    print("   | " + " | ".join(f"{c:<19}" for c in r) + " |")
print(f"   -> {len(cols3)} columns found, but the header row has a hole in it")
print("   -> nothing in the output records that one cell covers two columns")

print("\nwhen the table has printed lines, use them instead of guessing\n")
img = np.zeros((140, 440), np.uint8)
for y in [4, 30, 60, 90, 120]:
    cv2.line(img, (10, y), (430, y), 255, 2)
for x in [10, 190, 310, 430]:
    cv2.line(img, (x, 4), (x, 120), 255, 2)

h_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (60, 1))    # only long runs survive
v_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (1, 30))
h_lines = cv2.morphologyEx(img, cv2.MORPH_OPEN, h_kernel)
v_lines = cv2.morphologyEx(img, cv2.MORPH_OPEN, v_kernel)

h_at = [int(np.mean(g)) for g in np.split(np.flatnonzero(h_lines.any(axis=1)),
        np.where(np.diff(np.flatnonzero(h_lines.any(axis=1))) > 1)[0] + 1)]
v_at = [int(np.mean(g)) for g in np.split(np.flatnonzero(v_lines.any(axis=0)),
        np.where(np.diff(np.flatnonzero(v_lines.any(axis=0))) > 1)[0] + 1)]
print("horizontal rules found at y =", h_at)
print("vertical rules found at x   =", v_at)
print(f"that is a {len(h_at) - 1} x {len(v_at) - 1} grid of cells, with no guessing at all")
print("intersections give exact cell boundaries; drop each word into the cell it lands in")
Output
12 word boxes -> 4 rows x 3 columns
row centres:    [16.0, 46.0, 76.0, 106.0]
column starts:  [20.0, 200.0, 320.0]

   | Item      | Qty       | Price     |
   | Cement    | 12        | 380.00    |
   | Sand      | 4         | 120.50    |
   | Steel     | 30        | 2450.00   |

now break it two ways that happen on every real invoice

1. one cell wraps onto a second line:
   | Item      | Qty       | Price     |
   | Cement    | 12        | 380.00    |
   | Sand      | 4         | 120.50    |
   | Steel     | 30        | 2450.00   |
   | rods      |           |           |
   -> 'rods' became its own row, because rows are found by y position alone

2. one header cell spans two columns:
   | Item                | Quantity and price  |                     |
   | Cement              | 12                  | 380.00              |
   | Sand                | 4                   | 120.50              |
   | Steel               | 30                  | 2450.00             |
   -> 3 columns found, but the header row has a hole in it
   -> nothing in the output records that one cell covers two columns

when the table has printed lines, use them instead of guessing

horizontal rules found at y = [4, 30, 60, 90, 120]
vertical rules found at x   = [10, 190, 310, 430]
that is a 4 x 3 grid of cells, with no guessing at all
intersections give exact cell boundaries; drop each word into the cell it lands in

Reading the output carefully

The clean table came out exactly right, and that is the trap. Twelve boxes, four rows, three columns, every value in place. A demo built on this data would look finished.

The wrapped cell produced a fifth row containing only rods. No exception, no warning. Downstream this is a phantom line item with a missing quantity and a missing price. If the pipeline sums the price column, the total is still right — which means the check most people write would not catch it.

The merged header left a hole. Price vanished, Quantity and price sits in the second column, and the third header cell is empty. The rows below are still correct, so the values are fine and the column names are wrong. Every downstream mapping keyed on header text now fails.

The ruling-line method found the grid exactly: y = [4, 30, 60, 90, 120], x = [10, 190, 310, 430]. Four horizontal rules and three vertical rules bound a 4 x 3 grid, with no thresholds guessed from spacing. When a table has borders, use them — this is more accurate than any alignment heuristic and costs two morphology calls.

The kernels are the whole trick. MORPH_RECT (60, 1) opens the image with a 60-pixel-wide, 1-pixel-tall element. Only ink runs at least 60 pixels wide survive it, which is true of ruling lines and false of letters. The vertical kernel does the same job rotated. Set the lengths relative to your table size, not in absolute pixels.

What to reach for, in order

  1. Is it a digital PDF? Then do not do vision at all. pdfplumber gives you word boxes and, for bordered tables, the vector ruling lines directly. camelot wraps both the ruling-line ("lattice") and alignment ("stream") strategies.
  2. Is it a scan with borders? Morphology, as above.
  3. Is it a borderless scan? A trained model. Table Transformer (microsoft/table-transformer-detection for finding tables, microsoft/table-transformer-structure-recognition-v1.1-all for rows, columns and headers) is the standard open option, and is DETR applied to table structure.
  4. Is it complex — nested headers, spanning cells, no borders? PP-StructureV3 or a document VLM emitting HTML. Accept that you now need verification, since a generative model can invent a plausible number.

Note that Table Transformer gives you structure, not text. You still supply the words from OCR and assign them to the predicted cells.

Common mistakes

Clustering columns by left edge when numbers are right-aligned. Amounts share a right edge, not a left one. Cluster numeric columns by right edge, or by the box centre, and pick per column.

Assuming one row per text line. Wrapped cells break this, as shown. A more robust rule merges a line into the row above when its other columns are empty.

Losing merged cells silently. If your output format cannot express a spanning cell, you are discarding information rather than failing. Emit HTML with colspan and rowspan, or a cell list with explicit spans, so the loss is visible.

Skipping validation. Check that numeric columns parse, that a stated row count matches, and that a column of amounts sums to the stated total. Tables are the one place where a cheap arithmetic check catches almost every serious error.

Scoring with cell accuracy alone. A single missed row shifts every subsequent cell and scores near zero, which is correct but uninformative. Use a structure-aware metric — TEDS or GriTS — that separates structure errors from content errors.

Try it yourself

Change Sand to right-aligned by moving its box to start at 44 instead of 20. Rerun. It jumps into no column or a new one, depending on col_gap. Then switch build to cluster on box centres and watch the fix — and then break it again with a very wide cell.

What to learn next

Researcher — Mathematics and papers.

The two sub-problems, formally

Table detection: localise table regions on a page. This is standard object detection and is close to saturated on clean benchmarks — mAP above 0.95 is routine on ICDAR 2019 cTDaR and PubTables-1M.

Table structure recognition (TSR): given a table region, recover the logical grid. This is where the field lives. The output is not a set of boxes but a structure: an ordered set of rows and columns, plus a cell-to-grid assignment allowing spans. Two tables with identical cell boxes can have different structures if the spanning differs.

The dominant approaches

Object detection over structural elements. Table Transformer (Smock et al., CVPR 2022, PubTables-1M) casts TSR as DETR detection over six classes: table, column, row, column header, projected row header, and spanning cell. The grid comes from intersecting predicted rows and columns; spanning cells override the intersection. It is clean, and it inherits DETR's weakness on dense small objects — tables with many narrow columns.

The paper's larger contribution is PubTables-1M: 947,642 fully-annotated tables with a canonicalisation procedure that removes the oversegmentation ambiguity present in earlier datasets. That ambiguity — whether a multi-line header is one cell or several — was silently costing several points in every prior comparison.

Graph-based. Treat detected words as nodes and classify each pair as same-row, same-column or neither (Qasim et al., ICDAR 2019; GraphTSR). Naturally robust to spanning cells, and quadratic in word count.

Image-to-sequence. Emit HTML or LaTeX directly from the table image. TableMaster, EDD (Zhong et al., 2020) and modern VLMs sit here. It expresses arbitrary structure, and it can hallucinate a plausible cell.

Split-then-merge. Predict row and column separators over the image, forming a dense grid, then classify which adjacent cells merge (SPLERGE, Tensmeyer et al., ICDAR 2019). This decomposition maps well onto how humans describe tables and handles spans explicitly.

Metrics

TEDS (tree edit distance based similarity; Zhong et al., 2020) converts both tables to HTML trees and computes normalised tree edit distance:

$$ \mathrm{TEDS}(T_a, T_b) = 1 - \frac{\mathrm{EditDist}(T_a, T_b)}{\max(|T_a|, |T_b|)} $$

TEDS-Struct evaluates the tree with cell contents removed, isolating structure from OCR quality. Always report both: a system with 0.95 TEDS-Struct and 0.80 TEDS has a recognition problem, and the reverse has a structure problem.

GriTS (grid table similarity; Smock et al., 2023) compares the two tables as 2-D grids via a maximum-similarity substructure, in three variants: GriTS-Top for topology, GriTS-Con for content, GriTS-Loc for cell location. It avoids TEDS's sensitivity to the particular HTML serialisation chosen, which is a real problem when comparing systems that emit different but equivalent markup.

Cell-level precision and recall remain useful for error analysis and are misleading as headline numbers, because a single inserted row shifts every subsequent cell.

The honest state of the field

Three claims worth holding onto.

Bordered tables are close to solved. Ruling-line detection plus OCR reaches accuracy that no learned model meaningfully improves on, at a fraction of the cost. Reaching for a transformer here is a common and expensive mistake.

Borderless tables with spanning headers are not solved. Reported TEDS on the harder splits of FinTabNet and SciTSR sits well below what the headline numbers on PubTables-1M suggest, and the gap widens on scanned rather than born-digital input.

Benchmark leakage is a live concern. PubTables-1M, FinTabNet and SciTSR are all derived from public corpora that large VLMs have very likely seen during pretraining. A VLM scoring highly on them is weak evidence about your invoices. Build a small held-out set from your own documents before choosing.

The verification argument

For any pipeline that reads numbers out of tables, structural verification is worth more than a better model.

Cheap, effective checks: numeric columns must parse; a column of amounts must sum to a stated total; row counts must match a stated count; dates must be monotone in a statement. Each is a few lines, and together they catch the failure modes that matter — the ones where the output is well-formed and wrong.

What to learn next