Chunking and Long Documents

Chunking by document structure

Structure-aware chunking cuts a document at its own headings and sections, so every chunk keeps a label saying which part of the document it came from.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Structure-aware chunking cuts a document at its own headings, so every chunk keeps its section title attached.

Think about a printed textbook. Rip out one chapter, and you can still tell its subject, because the chapter heading tore out with it. Rip a page from the exact middle, with no heading in sight, and it tells you nothing.

Why it exists

The last two lessons cut text by character count or by meaning. Both ignore something a document often hands you for free: its own headings, sections, and structure.

A policy page or API reference already marks its own boundaries, in headings like "Eligibility" or "Authentication." Those are natural, author-written chunk boundaries. Ignoring them throws away a free signal.

Cutting at headings also solves a labelling problem. A chunk starting with "## Refund Timeline" tells both the search system and the model exactly what it covers.

How it works

# Refund Policy
## Eligibility
   Items can be returned within 30 days...
## Process
   Start a return from the Orders page...
## Refund Timeline
   Refunds are issued within 7 business days...

becomes three chunks, each carrying its own trail of headings:

Chunk 1: [Refund Policy > Eligibility]
         Items can be returned within 30 days...

Chunk 2: [Refund Policy > Process]
         Start a return from the Orders page...

Chunk 3: [Refund Policy > Refund Timeline]
         Refunds are issued within 7 business days...

That "[Refund Policy > Eligibility]" trail is a breadcrumb. It is the same idea as the "Home > Electronics > Phones" trail on a shopping website.

Where you have already seen it

  • Wikipedia search results. A snippet often shows the article title and the section heading above the quoted text.
  • Software documentation sites. A search for "reset password" links to the "Password Reset" section, not a 40-page manual's top.
  • Notion or Confluence search. Results show the page title and the specific heading the match came from.

Remember this

  • Headings are chunk boundaries the document's author already wrote for you.
  • Each chunk should carry a breadcrumb naming its section, so it still makes sense in isolation.
  • This method needs a document with real structure. Plain, unheaded text gets nothing from it.

What to learn next

Developer — Code and libraries.

This parses Markdown headings with a regular expression, then attaches a running breadcrumb to each section.

Setup

bash
python --version   # 3.9 or newer, no packages needed

Splitting on headings, with breadcrumbs

structure_chunk.py
import re

markdown = """# Refund Policy

## Eligibility

Items can be returned within 30 days of delivery if unused and in original packaging.

## Process

Start a return from the Orders page. A pickup is scheduled within 2 business days.

## Refund Timeline

Refunds are issued to the original payment method within 7 business days of pickup.
"""

def chunk_by_headings(md):
    pattern = re.compile(r'^(#{1,6})\s+(.*)$', re.MULTILINE)
    matches = list(pattern.finditer(md))
    chunks = []
    for i, m in enumerate(matches):
        level = len(m.group(1))
        heading = m.group(2).strip()
        body_start = m.end()
        body_end = matches[i + 1].start() if i + 1 < len(matches) else len(md)
        chunks.append({"level": level, "heading": heading, "body": md[body_start:body_end].strip()})
    return chunks

def with_breadcrumbs(chunks):
    trail = {}
    out = []
    for c in chunks:
        trail[c["level"]] = c["heading"]
        for lvl in list(trail):
            if lvl > c["level"]:
                del trail[lvl]
        breadcrumb = " > ".join(trail[lvl] for lvl in sorted(trail))
        out.append({**c, "breadcrumb": breadcrumb})
    return out

for c in with_breadcrumbs(chunk_by_headings(markdown)):
    print(f"[{c['breadcrumb']}]")
    print(" ", c["body"] or "(empty — title-only section)")
    print()
Output
[Refund Policy]
  (empty — title-only section)

[Refund Policy > Eligibility]
  Items can be returned within 30 days of delivery if unused and in original packaging.

[Refund Policy > Process]
  Start a return from the Orders page. A pickup is scheduled within 2 business days.

[Refund Policy > Refund Timeline]
  Refunds are issued to the original payment method within 7 business days of pickup.

Line by line

The regex ^(#{1,6})\s+(.*)$ matches one to six # characters at the start of a line, the Markdown heading levels. Group 1 captures the # marks, whose length gives the heading level. Group 2 captures the heading text.

body_end looks ahead to the next heading's start, or the end of the document for the last one. Everything between one heading and the next is that section's body.

with_breadcrumbs keeps a running trail dictionary keyed by heading level. A new level-2 heading clears any stale level-3-or-deeper entries left from an earlier branch. A breadcrumb then never shows a heading from an unrelated subsection.

The H1 chunk has an empty body. The document's title has no text directly under it before the first ## — a real, common edge case, handled here by printing a placeholder instead of an empty string.

Common mistakes

Feeding this raw HTML instead of Markdown. HTML headings look like <h2>Process</h2>, not ## Process. This regex only matches Markdown syntax — HTML documents need an HTML parser with the same breadcrumb idea applied to <h1>–<h6> tags.

Dropping the title-only chunk without checking it. An empty H1 chunk is expected and usually discarded. A non-empty section with real content but zero body text after cleaning, though, likely signals a parsing bug worth investigating.

Assuming every document has consistent heading levels. A document that jumps from # straight to ###, skipping ##, still parses here. The breadcrumb stays correct only because the trail dictionary keys on the actual level number, not on heading order. Verify this on your real documents.

Chunking by heading alone, with no size cap. A single section under one heading can still run to several pages. Structure-aware chunking answers "where is a good boundary," not "how big should a chunk be." Combine it with a size limit for very long sections.

Try it yourself

Add a ### Refunds for damaged items subsection under ## Refund Timeline, then re-run.

The new chunk's breadcrumb should read Refund Policy > Refund Timeline > Refunds for damaged items, three levels deep. The trail dictionary now holds a level-3 entry alongside the still-valid levels 1 and 2 above it.

What to learn next

Researcher — Mathematics and papers.

Structure as a free segmentation signal

Structure-aware chunking treats author-inserted markup — headings, list nesting, table boundaries — as ground-truth segmentation labels. This sidesteps the estimation problem that semantic chunking must solve, at the cost of only working on documents that carry usable structure in the first place. Plain OCR output, transcripts, and un-marked plaintext get nothing from this approach.

Breadcrumb prepending as a retrieval intervention

A breadcrumb prepended to chunk text is one case of a broader technique. It enriches a chunk's surface text with information true of it, but not literally present in its span. The next lesson in this section, contextual chunk headers, generalises this beyond headings to arbitrary document metadata. Both interventions are pre-retrieval: they change what gets embedded, not the retrieval algorithm. Gao et al. (2023), Retrieval-Augmented Generation for Large Language Models: A Survey (arXiv:2312.10997), uses this pre-retrieval, retrieval, post-retrieval split to organise the whole RAG optimisation literature.

Interaction with hierarchical retrieval

A breadcrumb-carrying chunk supports two retrieval strategies a flat chunk cannot. First, filtering: metadata search on the breadcrumb path (see Metadata filtering) can restrict a query to one section before semantic search runs. Second, hierarchical re-ranking: a coarse match at section level can gate a finer match at paragraph level, shrinking the candidate set a cross-encoder must score. Both strategies need the heading hierarchy preserved as structured metadata, not flattened into plain chunk text. Production systems typically store the breadcrumb as a separate field, alongside the chunk, in addition to or instead of prepending it.

Complexity

For a document with h headings and n total characters, the regex scan is O(n). Building breadcrumbs is O(h), since each heading's trail update is bounded by the fixed maximum depth of 6. That makes structure-aware chunking one of the cheapest methods in this section — cheaper than semantic chunking's per-sentence embedding cost.

Current state and open problems

Real-world document structure is far messier than clean Markdown. PDF-to-text extraction frequently loses heading formatting, collapsing a heading and its first paragraph into indistinguishable plain text. Production pipelines increasingly rely on layout-aware parsers — models that predict heading, body, or table from visual and typographic cues, not text alone. This lesson's regex approach is the right mental model for genuinely structured formats like Markdown, HTML, and DOCX. Treat it as a narrow case of a broader, unsolved layout-understanding problem for anything scanned or PDF-native.

Key references

  • Gao, Y. et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997

What to learn next