Chunking by document structure
Structure-aware chunking cuts a document at its own headings and sections, so every chunk keeps a label saying which part of the document it came from.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Structure-aware chunking cuts a document at its own headings, so every chunk keeps its section title attached.
Think about a printed textbook. Rip out one chapter, and you can still tell its subject, because the chapter heading tore out with it. Rip a page from the exact middle, with no heading in sight, and it tells you nothing.
Why it exists
The last two lessons cut text by character count or by meaning. Both ignore something a document often hands you for free: its own headings, sections, and structure.
A policy page or API reference already marks its own boundaries, in headings like "Eligibility" or "Authentication." Those are natural, author-written chunk boundaries. Ignoring them throws away a free signal.
Cutting at headings also solves a labelling problem. A chunk starting with "## Refund Timeline" tells both the search system and the model exactly what it covers.
How it works
# Refund Policy
## Eligibility
Items can be returned within 30 days...
## Process
Start a return from the Orders page...
## Refund Timeline
Refunds are issued within 7 business days...becomes three chunks, each carrying its own trail of headings:
Chunk 1: [Refund Policy > Eligibility]
Items can be returned within 30 days...
Chunk 2: [Refund Policy > Process]
Start a return from the Orders page...
Chunk 3: [Refund Policy > Refund Timeline]
Refunds are issued within 7 business days...That "[Refund Policy > Eligibility]" trail is a breadcrumb. It is the same idea as the "Home > Electronics > Phones" trail on a shopping website.
Where you have already seen it
- Wikipedia search results. A snippet often shows the article title and the section heading above the quoted text.
- Software documentation sites. A search for "reset password" links to the "Password Reset" section, not a 40-page manual's top.
- Notion or Confluence search. Results show the page title and the specific heading the match came from.
Remember this
- Headings are chunk boundaries the document's author already wrote for you.
- Each chunk should carry a breadcrumb naming its section, so it still makes sense in isolation.
- This method needs a document with real structure. Plain, unheaded text gets nothing from it.
What to learn next
- Chunking strategies compared — the character-count and sentence-based baselines.
- Chunking source code and tables — structure-aware splitting for two harder document types.
- LlamaIndex — a library with built-in structure-aware document parsers.
Developer — Code and libraries.
This parses Markdown headings with a regular expression, then attaches a running breadcrumb to each section.
Setup
python --version # 3.9 or newer, no packages neededSplitting on headings, with breadcrumbs
import re
markdown = """# Refund Policy
## Eligibility
Items can be returned within 30 days of delivery if unused and in original packaging.
## Process
Start a return from the Orders page. A pickup is scheduled within 2 business days.
## Refund Timeline
Refunds are issued to the original payment method within 7 business days of pickup.
"""
def chunk_by_headings(md):
pattern = re.compile(r'^(#{1,6})\s+(.*)$', re.MULTILINE)
matches = list(pattern.finditer(md))
chunks = []
for i, m in enumerate(matches):
level = len(m.group(1))
heading = m.group(2).strip()
body_start = m.end()
body_end = matches[i + 1].start() if i + 1 < len(matches) else len(md)
chunks.append({"level": level, "heading": heading, "body": md[body_start:body_end].strip()})
return chunks
def with_breadcrumbs(chunks):
trail = {}
out = []
for c in chunks:
trail[c["level"]] = c["heading"]
for lvl in list(trail):
if lvl > c["level"]:
del trail[lvl]
breadcrumb = " > ".join(trail[lvl] for lvl in sorted(trail))
out.append({**c, "breadcrumb": breadcrumb})
return out
for c in with_breadcrumbs(chunk_by_headings(markdown)):
print(f"[{c['breadcrumb']}]")
print(" ", c["body"] or "(empty — title-only section)")
print()[Refund Policy] (empty — title-only section) [Refund Policy > Eligibility] Items can be returned within 30 days of delivery if unused and in original packaging. [Refund Policy > Process] Start a return from the Orders page. A pickup is scheduled within 2 business days. [Refund Policy > Refund Timeline] Refunds are issued to the original payment method within 7 business days of pickup.
Line by line
The regex ^(#{1,6})\s+(.*)$ matches one to six # characters at the start of a line, the Markdown heading levels. Group 1 captures the # marks, whose length gives the heading level. Group 2 captures the heading text.
body_end looks ahead to the next heading's start, or the end of the document for the last one. Everything between one heading and the next is that section's body.
with_breadcrumbs keeps a running trail dictionary keyed by heading level. A new level-2 heading clears any stale level-3-or-deeper entries left from an earlier branch. A breadcrumb then never shows a heading from an unrelated subsection.
The H1 chunk has an empty body. The document's title has no text directly under it before the first ## — a real, common edge case, handled here by printing a placeholder instead of an empty string.
Common mistakes
Feeding this raw HTML instead of Markdown. HTML headings look like <h2>Process</h2>, not ## Process. This regex only matches Markdown syntax — HTML documents need an HTML parser with the same breadcrumb idea applied to <h1>–<h6> tags.
Dropping the title-only chunk without checking it. An empty H1 chunk is expected and usually discarded. A non-empty section with real content but zero body text after cleaning, though, likely signals a parsing bug worth investigating.
Assuming every document has consistent heading levels. A document that jumps from # straight to ###, skipping ##, still parses here. The breadcrumb stays correct only because the trail dictionary keys on the actual level number, not on heading order. Verify this on your real documents.
Chunking by heading alone, with no size cap. A single section under one heading can still run to several pages. Structure-aware chunking answers "where is a good boundary," not "how big should a chunk be." Combine it with a size limit for very long sections.
Try it yourself
Add a ### Refunds for damaged items subsection under ## Refund Timeline, then re-run.
The new chunk's breadcrumb should read Refund Policy > Refund Timeline > Refunds for damaged items, three levels deep. The trail dictionary now holds a level-3 entry alongside the still-valid levels 1 and 2 above it.
What to learn next
- Chunking source code and tables — structure-aware rules for two document types with no headings.
- Giving each chunk its context back — turning a breadcrumb into text the embedding model actually reads.
- LlamaIndex — production-grade structure parsers for PDFs, HTML, and Markdown.
Researcher — Mathematics and papers.
Structure as a free segmentation signal
Structure-aware chunking treats author-inserted markup — headings, list nesting, table boundaries — as ground-truth segmentation labels. This sidesteps the estimation problem that semantic chunking must solve, at the cost of only working on documents that carry usable structure in the first place. Plain OCR output, transcripts, and un-marked plaintext get nothing from this approach.
Breadcrumb prepending as a retrieval intervention
A breadcrumb prepended to chunk text is one case of a broader technique. It enriches a chunk's surface text with information true of it, but not literally present in its span. The next lesson in this section, contextual chunk headers, generalises this beyond headings to arbitrary document metadata. Both interventions are pre-retrieval: they change what gets embedded, not the retrieval algorithm. Gao et al. (2023), Retrieval-Augmented Generation for Large Language Models: A Survey (arXiv:2312.10997), uses this pre-retrieval, retrieval, post-retrieval split to organise the whole RAG optimisation literature.
Interaction with hierarchical retrieval
A breadcrumb-carrying chunk supports two retrieval strategies a flat chunk cannot. First, filtering: metadata search on the breadcrumb path (see Metadata filtering) can restrict a query to one section before semantic search runs. Second, hierarchical re-ranking: a coarse match at section level can gate a finer match at paragraph level, shrinking the candidate set a cross-encoder must score. Both strategies need the heading hierarchy preserved as structured metadata, not flattened into plain chunk text. Production systems typically store the breadcrumb as a separate field, alongside the chunk, in addition to or instead of prepending it.
Complexity
For a document with h headings and n total characters, the regex scan is O(n). Building breadcrumbs is O(h), since each heading's trail update is bounded by the fixed maximum depth of 6. That makes structure-aware chunking one of the cheapest methods in this section — cheaper than semantic chunking's per-sentence embedding cost.
Current state and open problems
Real-world document structure is far messier than clean Markdown. PDF-to-text extraction frequently loses heading formatting, collapsing a heading and its first paragraph into indistinguishable plain text. Production pipelines increasingly rely on layout-aware parsers — models that predict heading, body, or table from visual and typographic cues, not text alone. This lesson's regex approach is the right mental model for genuinely structured formats like Markdown, HTML, and DOCX. Treat it as a narrow case of a broader, unsolved layout-understanding problem for anything scanned or PDF-native.
Key references
- Gao, Y. et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997
What to learn next
- Chunking source code and tables — structure without headings.
- Metadata filtering — using a breadcrumb to narrow a search before ranking begins.
- Checking that an answer really came from the source — where breadcrumb metadata becomes a citation.