Text Preprocessing

Getting clean text out of PDFs

A PDF stores painted characters with coordinates, not sentences — extraction is reconstruction, and knowing that explains every weird result you will meet.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A PDF does not store sentences. It stores characters painted at fixed positions on a page.

"Extracting text" means guessing the sentences back from the painting.

Think of a rangoli. The artist placed each grain of coloured powder at an exact spot, and your eye reads the whole as a peacock. Ask the floor "what does this mean?" and it can only answer: powder, positions, colours. A PDF is a rangoli of characters. The reading order in your head is not stored anywhere.

Why it exists

PDF was invented to make documents look identical everywhere — same fonts, same layout, on any printer, forever. It nails that goal precisely because it stores appearance: draw this glyph at this position, that one there.

The cost surfaces the moment machines need the text. Invoices, research papers, contracts, government circulars — enormous amounts of valuable data live in PDFs. Every "chat with your PDF" tool, every invoice-processing system, every legal-search product starts with the same unglamorous step: turning painted characters back into readable lines.

An extractor is the tool that does the guessing: it collects the characters, sorts them by position, and decides where words, lines and paragraphs begin.

How it works

what the PDF stores:            what you want back:
("Invoice", x=72,  y=720)
("Date:",   x=300, y=720)  →   "Invoice INV-2041   Date: 30-08-2026"
("INV-2041",x=130, y=720)

The extractor sorts characters top-to-bottom, left-to-right, and inserts spaces and line breaks where the position gaps suggest them. Two-column pages, tables and footers are where the guessing gets hard: read a two-column page straight across and you interleave unrelated sentences.

One more trap: some PDFs are scans — photographs of paper. They contain no characters at all, only pixels. Those need OCR (optical character recognition), a model that reads pictures of letters.

A real example you have seen

Tap "copy" on a PDF statement from your bank and paste it into notes — you get columns mashed together and amounts detached from their labels. That mess is the reconstruction problem. Tools that answer questions about a PDF you upload, like the chat with your PDF project, live or die on this step.

Remember this

  • A PDF stores positioned characters, not sentences; extraction is reconstruction.
  • Columns, tables and footers are where reconstruction goes wrong.
  • Scanned PDFs contain no text at all — they need OCR first.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pypdf fpdf2

Verified with pypdf 6.10 and fpdf2 2.8, CPU only. fpdf2 is used only to create a test PDF, so this lesson needs no downloads and no sample files.

Make a PDF, then get the text back out

pdf_extract.py
from fpdf import FPDF
from pypdf import PdfReader

# make a small two-column invoice PDF so this lesson needs no download
pdf = FPDF()
pdf.add_page()
pdf.set_font("helvetica", size=12)
pdf.cell(90, 10, "Invoice INV-2041")
pdf.cell(90, 10, "Date: 30-08-2026", new_x="lmargin", new_y="next")
pdf.cell(90, 10, "Wireless mouse")
pdf.cell(90, 10, "Rs. 1,200", new_x="lmargin", new_y="next")
pdf.cell(90, 10, "Mechanical keyboard")
pdf.cell(90, 10, "Rs. 3,300", new_x="lmargin", new_y="next")
pdf.output("invoice.pdf")

reader = PdfReader("invoice.pdf")
page = reader.pages[0]
print("pages:", len(reader.pages))
print(repr(page.extract_text()))

# layout mode tries to preserve the columns you saw on screen
print("---")
print(page.extract_text(extraction_mode="layout"))
Output
pages: 1
'Invoice INV-2041 Date: 30-08-2026\nWireless mouse Rs. 1,200\nMechanical keyboard Rs. 3,300'
---
Invoice INV-2041                                      Date: 30-08-2026

Wireless mouse                                        Rs. 1,200

Mechanical keyboard                                   Rs. 3,300

The walkthrough

Default mode glued the columns. 'Wireless mouse Rs. 1,200' — label and price mashed into one line with a single space. For a paragraph of prose that behaviour is right; for anything tabular it destroys the structure you needed.

extraction_mode="layout" preserved the geometry. The wide gaps reappear as runs of spaces, so a downstream parser (or a regex splitting on two-plus spaces) can recover label–value pairs. Layout mode is the first thing to reach for on invoices, statements and reports.

Know the tool tiers. pypdf is pure Python and fine for born-digital, mostly-linear documents. pdfplumber exposes per-character coordinates and has real table extraction. PyMuPDF (fitz) is the fast C-backed engine with block-level layout. For scanned pages, ocrmypdf (wrapping Tesseract) adds a text layer; modern vision-language models are covered in document AI.

Detect the scanned-PDF case in code. page.extract_text() returning an empty or near-empty string on a page that visibly has content is the tell. Route those pages to OCR instead of concluding the document is blank.

Common mistakes

Treating extraction output as clean text. Headers and footers repeat on every page and stitch themselves into sentences across page joins. Hyphenated line breaks split words. Ligatures arrive as odd characters. Budget a cleanup pass — the normalisation and Unicode lessons are exactly the toolkit.

Reading two-column papers straight down the page. Academic PDFs interleave badly in naive extractors. Use a layout-aware tool and check a page visually against its extraction before processing thousands.

Parsing tables from plain text output. If tables matter, use a table-aware extractor (pdfplumber, camelot) that works from ruling lines and coordinates — reconstructing a table from flattened text is misery with a deadline attached.

Assuming the text you see is the text stored. Some PDFs use font tricks where the stored character codes differ from the displayed glyphs — extraction yields gibberish despite a perfect-looking page. If output looks like nonsense for one specific file, inspect it before blaming the library.

Try it yourself

Add a second page with pdf.add_page() and a footer line on both pages, then extract all pages in a loop. Watch the footer interrupt the text flow — then write the two lines of Python that strip it.

What to learn next

Researcher — Mathematics and papers.

Why extraction is inherently heuristic

A PDF content stream is a PostScript-descended program: operators set transformation matrices and fonts, and text-showing operators (Tj, TJ) paint glyph runs at computed positions (ISO 32000-2). Three consequences: (1) reading order is unspecified — the spec's optional "tagged PDF" structure tree provides it, but a minority of real-world files carry useful tags; (2) word and paragraph boundaries do not exist — extractors infer spaces from inter-glyph advances exceeding thresholds, which is why kerned headlines sprout phantom spaces; (3) the mapping from character codes to Unicode is only as good as the font's /ToUnicode CMap — subsetted fonts without one make faithful extraction literally impossible from the drawing commands alone, the "looks perfect, extracts garbage" case.

The reconstruction pipeline, formalised

Layout analysis is the research name: segment the page into text blocks, figures and tables, then order the blocks. Classical geometric algorithms — recursive XY-cut (Nagy and Seth, 1984), Docstrum (O'Gorman, 1993) — cluster glyphs by whitespace valleys and nearest-neighbour angles. Modern systems treat it as detection/segmentation learning: LayoutLM and successors (Xu et al., 2020) jointly embed text, position and image; DocLayNet (Pfitzmann et al., 2022) and PubLayNet supply training corpora; end-to-end document parsers (Nougat — Blecher et al., 2023; Donut — Kim et al., 2022) map page images straight to markup, and current vision-language models absorb the task wholesale. Table structure recognition is its own subfield (ICDAR competitions; grid-and-span inference from ruling lines and alignment).

OCR

For scanned input, Tesseract 4+ (LSTM line recogniser; Smith, 2007 for the engine's lineage) remains the open-source workhorse, wrapped operationally by ocrmypdf, which adds an invisible text layer while preserving the image. Accuracy is dominated by binarisation and skew; on clean 300-dpi Latin-script scans, character error under 1% is typical, degrading sharply with compression artefacts and Indic scripts — where newer transformer OCRs (TrOCR — Li et al., 2021; Surya) lead. Always distinguish the two failure axes: recognition error (wrong characters) versus layout error (right characters, wrong order) — downstream RAG pipelines are far more damaged by the second.

Evaluation and practice

Benchmark extraction as you would any model: sample pages, produce gold text, compute character/word error rates and order-aware metrics. In production corpora, route by document class — born-digital single-column → fast extractor; multi-column or tabular → layout-aware; image-only → OCR — with the empty-text heuristic as the router's cheapest feature. The engineering reality worth stating: on real document collections, extraction quality sets the ceiling for every downstream task, and it is nearly always the highest-leverage place to spend effort.

What to learn next