Chunking and Long Documents

Chunking source code and tables

Code and tables break under text-chunking rules built for prose, so each needs its own boundaries — function or class edges for code, and a repeated header row for tables.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Code and tables need their own chunking rules, because cutting them like prose breaks them.

Think about a train made of coupled bogies. You can uncouple it at the joints between bogies, never through the middle of one. Cut through the middle of a bogie and both halves are useless. A function in code, or a row in a table, works the same way.

Why it exists

Sentence-based chunking looks for a full stop. Code rarely has one — a function can run fifty lines without a single period. A cut by character count can slice through a function body, an open bracket, or a half-finished loop.

Tables have a different problem. A table's column names sit once, in the header row. Split it into batches of ten rows, and every batch after the first loses those labels. A chunk reading "22, 1900" gives no clue those mean "revenue" and "units."

Both problems share one fix: cut at the structure's own natural joints. Repeat whatever context each piece needs to stand alone.

How it works

Code, cut by function boundary (a natural "joint"):

  def add(a, b):        <-- chunk 1 starts
      return a + b       <-- chunk 1 ends

  def subtract(a, b):    <-- chunk 2 starts
      return a - b       <-- chunk 2 ends


Table, cut by row batch, header repeated in every chunk:

  Chunk 1                     Chunk 2
  product, revenue             product, revenue   <- header repeated
  Widget A, 12.5                Widget C, 3.1
  Widget B, 8.2

Where you have already seen it

  • GitHub Copilot and code-search tools. Jump-to-definition and code search work on whole functions, not arbitrary line ranges.
  • Spreadsheet formulas that reference a header row. Excel and Google Sheets keep column names pinned at the top for exactly this reason.
  • Financial report chatbots. Ask "what was Widget C's revenue?" and a good one knows which column that is, even in a three-row chunk.

Remember this

  • Code should be cut at function or class boundaries, never through the middle of one.
  • Tables need their header row repeated in every chunk, or the numbers lose their meaning.
  • Both problems come from the same cause: text-shaped chunking rules applied to non-text-shaped data.

What to learn next

Developer — Code and libraries.

Two small functions: one splits Python source at function and class boundaries using the ast module, the other splits a CSV table into row batches with the header repeated.

Setup

bash
python --version   # 3.9 or newer, ast and csv are both in the standard library

Splitting code at function boundaries

code_chunk.py
import ast

source = '''
def add(a, b):
    """Add two numbers."""
    return a + b


def subtract(a, b):
    """Subtract b from a."""
    return a - b


class Calculator:
    def multiply(self, a, b):
        return a * b
'''

tree = ast.parse(source)
lines = source.splitlines()

chunks = []
for node in ast.iter_child_nodes(tree):
    if isinstance(node, (ast.FunctionDef, ast.ClassDef)):
        start = node.lineno - 1
        end = node.end_lineno
        chunks.append("\n".join(lines[start:end]))

for i, c in enumerate(chunks):
    print(f"--- chunk {i} ---")
    print(c)
Output
--- chunk 0 ---
def add(a, b):
    """Add two numbers."""
    return a + b
--- chunk 1 ---
def subtract(a, b):
    """Subtract b from a."""
    return a - b
--- chunk 2 ---
class Calculator:
    def multiply(self, a, b):
        return a * b

Splitting a table, header repeated in every chunk

table_chunk.py
import csv
import io

csv_text = """product,quarter,revenue_crore,units_sold
Widget A,Q1,12.5,4200
Widget A,Q2,14.1,4700
Widget B,Q1,8.2,1900
Widget B,Q2,9.0,2100
Widget C,Q1,3.1,900
Widget C,Q2,3.9,1050
"""

rows = list(csv.reader(io.StringIO(csv_text)))
header, data_rows = rows[0], rows[1:]

def chunk_table(header, data_rows, rows_per_chunk):
    chunks = []
    for i in range(0, len(data_rows), rows_per_chunk):
        batch = data_rows[i:i + rows_per_chunk]
        lines = [",".join(header)] + [",".join(row) for row in batch]
        chunks.append("\n".join(lines))
    return chunks

for i, c in enumerate(chunk_table(header, data_rows, rows_per_chunk=2)):
    print(f"--- chunk {i} ---")
    print(c)
Output
--- chunk 0 ---
product,quarter,revenue_crore,units_sold
Widget A,Q1,12.5,4200
Widget A,Q2,14.1,4700
--- chunk 1 ---
product,quarter,revenue_crore,units_sold
Widget B,Q1,8.2,1900
Widget B,Q2,9.0,2100
--- chunk 2 ---
product,quarter,revenue_crore,units_sold
Widget C,Q1,3.1,900
Widget C,Q2,3.9,1050

Line by line

ast.parse(source) builds a syntax tree, the structured representation Python itself uses to understand code. ast.iter_child_nodes(tree) walks the top-level statements, and isinstance(node, (ast.FunctionDef, ast.ClassDef)) keeps only functions and classes.

node.lineno and node.end_lineno give the exact first and last line of each function or class, provided directly by the parser. This is why the cut never lands mid-function — the boundary comes from Python's own grammar, not a guess.

Note class Calculator stayed whole, methods and all. ast.iter_child_nodes only walks top-level nodes, so a method inside a class is not split out on its own. That is deliberate here — a class chunk that is too large is a separate problem, solved by walking one level deeper for large classes.

The table function rebuilds header as the first line of every chunk. Every chunk is now readable in isolation, at the cost of repeating a few bytes of header text per chunk.

Common mistakes

Chunking code by line count, the way you would chunk prose. A 20-line chunk boundary has no relationship to where a function starts or ends. It is close to guaranteed to split at least one function in half across a real file.

Forgetting module-level imports and constants. A function chunk with no import numpy as np above it may reference np with nothing in the chunk explaining where it comes from. Production code chunkers often prepend the file's import block to every function chunk.

Splitting a table without deciding what "one chunk" means for aggregate questions. "What was total 2023 revenue?" needs every row, not one batch of two. Row-batch chunking works for lookup questions about a specific row. It answers aggregate questions badly, or not at all — see the next lesson in this section, plus Answering questions about tables.

Assuming every function fits in one chunk. A 300-line function still becomes one chunk here, with no size cap. Combine structural splitting with a maximum size limit, falling back to a sub-split within an oversized function when needed.

Try it yourself

Add a def divide(a, b): function to source, after Calculator, and re-run.

A new chunk 3 should appear, cleanly separated from Calculator. Function boundaries in the syntax tree are unaffected by whatever comes before or after them in the file — that isolation is what a syntax-aware splitter buys you over a character-count splitter.

What to learn next

Researcher — Mathematics and papers.

Code as a structured, not linear, object

Prose chunking treats a document as a 1-dimensional character stream. Source code is not linear in the sense that matters — its true structure is a syntax tree, and any chunk boundary that ignores the tree risks producing a syntactically invalid fragment. Splitting on ast node boundaries, as in the developer block, guarantees every chunk is independently parseable. Production code-search and code-RAG systems generally go further, using a language-agnostic parser generator — most commonly tree-sitter — so the same boundary-respecting approach works across languages without one hand-written parser per language.

Retrieval implications for code

A retrieved function without its surrounding class context, or without the imports it depends on, is frequently uninterpretable to a model even though it is syntactically valid on its own. This motivates hierarchical retrieval for code specifically: retrieve at function granularity, but supply file-level or module-level context alongside it, mirroring the parent-document retrieval pattern covered later in this section.

Tables and structured-data retrieval

Naive row-batch chunking, as shown in the developer block, is workable for point-lookup questions — "what is Widget C's Q1 revenue" — and fails for aggregate questions spanning many rows, since no single chunk contains the full column needed to sum or average. Herzig et al. (2020), TAPAS: Weakly Supervised Table Parsing via Pre-training (arXiv:2004.02349), addresses this differently: rather than chunking a table for retrieval, TAPAS encodes the entire table jointly with the question, using row and column embeddings, and predicts cell selections and aggregation operations directly. This is a fundamentally different strategy than retrieve-then-read — it treats the table as one structured unit, small enough to fit in a single model call, rather than a document to be chunked at all. That strategy stops working once the table exceeds the model's input limit, which real spreadsheets and database exports routinely do — at that point chunking again becomes necessary, and the aggregate-question problem returns.

Complexity

Parsing source into an AST is O(n) in file length for a well-formed file. Table chunking is O(rows). Both are cheaper than semantic chunking's per-unit embedding cost, since neither requires a model call.

Key references

  • Herzig, J. et al. (2020). TAPAS: Weakly Supervised Table Parsing via Pre-training. arXiv:2004.02349

Current state and open problems

Code retrieval increasingly uses embedding models trained specifically on code, rather than general-purpose text embedding models, since code identifiers and structure carry different statistical regularities than prose. Table retrieval remains split between two unresolved strategies: chunk-and-retrieve, which handles arbitrarily large tables but struggles with aggregation, and whole-table encoding, which handles aggregation but hits a hard size ceiling. No single deployed system fully unifies both, and picking between them, per document type, remains a manual design decision.

What to learn next