Chunking source code and tables
Code and tables break under text-chunking rules built for prose, so each needs its own boundaries — function or class edges for code, and a repeated header row for tables.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Code and tables need their own chunking rules, because cutting them like prose breaks them.
Think about a train made of coupled bogies. You can uncouple it at the joints between bogies, never through the middle of one. Cut through the middle of a bogie and both halves are useless. A function in code, or a row in a table, works the same way.
Why it exists
Sentence-based chunking looks for a full stop. Code rarely has one — a function can run fifty lines without a single period. A cut by character count can slice through a function body, an open bracket, or a half-finished loop.
Tables have a different problem. A table's column names sit once, in the header row. Split it into batches of ten rows, and every batch after the first loses those labels. A chunk reading "22, 1900" gives no clue those mean "revenue" and "units."
Both problems share one fix: cut at the structure's own natural joints. Repeat whatever context each piece needs to stand alone.
How it works
Code, cut by function boundary (a natural "joint"):
def add(a, b): <-- chunk 1 starts
return a + b <-- chunk 1 ends
def subtract(a, b): <-- chunk 2 starts
return a - b <-- chunk 2 ends
Table, cut by row batch, header repeated in every chunk:
Chunk 1 Chunk 2
product, revenue product, revenue <- header repeated
Widget A, 12.5 Widget C, 3.1
Widget B, 8.2Where you have already seen it
- GitHub Copilot and code-search tools. Jump-to-definition and code search work on whole functions, not arbitrary line ranges.
- Spreadsheet formulas that reference a header row. Excel and Google Sheets keep column names pinned at the top for exactly this reason.
- Financial report chatbots. Ask "what was Widget C's revenue?" and a good one knows which column that is, even in a three-row chunk.
Remember this
- Code should be cut at function or class boundaries, never through the middle of one.
- Tables need their header row repeated in every chunk, or the numbers lose their meaning.
- Both problems come from the same cause: text-shaped chunking rules applied to non-text-shaped data.
What to learn next
- Chunking by document structure — the same natural-boundary idea, applied to headings.
- Answering questions about tables — what happens after a table chunk gets retrieved.
- Giving each chunk its context back — repeating headers is one case of a bigger pattern.
Developer — Code and libraries.
Two small functions: one splits Python source at function and class boundaries using the ast module, the other splits a CSV table into row batches with the header repeated.
Setup
python --version # 3.9 or newer, ast and csv are both in the standard librarySplitting code at function boundaries
import ast
source = '''
def add(a, b):
"""Add two numbers."""
return a + b
def subtract(a, b):
"""Subtract b from a."""
return a - b
class Calculator:
def multiply(self, a, b):
return a * b
'''
tree = ast.parse(source)
lines = source.splitlines()
chunks = []
for node in ast.iter_child_nodes(tree):
if isinstance(node, (ast.FunctionDef, ast.ClassDef)):
start = node.lineno - 1
end = node.end_lineno
chunks.append("\n".join(lines[start:end]))
for i, c in enumerate(chunks):
print(f"--- chunk {i} ---")
print(c)--- chunk 0 ---
def add(a, b):
"""Add two numbers."""
return a + b
--- chunk 1 ---
def subtract(a, b):
"""Subtract b from a."""
return a - b
--- chunk 2 ---
class Calculator:
def multiply(self, a, b):
return a * bSplitting a table, header repeated in every chunk
import csv
import io
csv_text = """product,quarter,revenue_crore,units_sold
Widget A,Q1,12.5,4200
Widget A,Q2,14.1,4700
Widget B,Q1,8.2,1900
Widget B,Q2,9.0,2100
Widget C,Q1,3.1,900
Widget C,Q2,3.9,1050
"""
rows = list(csv.reader(io.StringIO(csv_text)))
header, data_rows = rows[0], rows[1:]
def chunk_table(header, data_rows, rows_per_chunk):
chunks = []
for i in range(0, len(data_rows), rows_per_chunk):
batch = data_rows[i:i + rows_per_chunk]
lines = [",".join(header)] + [",".join(row) for row in batch]
chunks.append("\n".join(lines))
return chunks
for i, c in enumerate(chunk_table(header, data_rows, rows_per_chunk=2)):
print(f"--- chunk {i} ---")
print(c)--- chunk 0 --- product,quarter,revenue_crore,units_sold Widget A,Q1,12.5,4200 Widget A,Q2,14.1,4700 --- chunk 1 --- product,quarter,revenue_crore,units_sold Widget B,Q1,8.2,1900 Widget B,Q2,9.0,2100 --- chunk 2 --- product,quarter,revenue_crore,units_sold Widget C,Q1,3.1,900 Widget C,Q2,3.9,1050
Line by line
ast.parse(source) builds a syntax tree, the structured representation Python itself uses to understand code. ast.iter_child_nodes(tree) walks the top-level statements, and isinstance(node, (ast.FunctionDef, ast.ClassDef)) keeps only functions and classes.
node.lineno and node.end_lineno give the exact first and last line of each function or class, provided directly by the parser. This is why the cut never lands mid-function — the boundary comes from Python's own grammar, not a guess.
Note class Calculator stayed whole, methods and all. ast.iter_child_nodes only walks top-level nodes, so a method inside a class is not split out on its own. That is deliberate here — a class chunk that is too large is a separate problem, solved by walking one level deeper for large classes.
The table function rebuilds header as the first line of every chunk. Every chunk is now readable in isolation, at the cost of repeating a few bytes of header text per chunk.
Common mistakes
Chunking code by line count, the way you would chunk prose. A 20-line chunk boundary has no relationship to where a function starts or ends. It is close to guaranteed to split at least one function in half across a real file.
Forgetting module-level imports and constants. A function chunk with no import numpy as np above it may reference np with nothing in the chunk explaining where it comes from. Production code chunkers often prepend the file's import block to every function chunk.
Splitting a table without deciding what "one chunk" means for aggregate questions. "What was total 2023 revenue?" needs every row, not one batch of two. Row-batch chunking works for lookup questions about a specific row. It answers aggregate questions badly, or not at all — see the next lesson in this section, plus Answering questions about tables.
Assuming every function fits in one chunk. A 300-line function still becomes one chunk here, with no size cap. Combine structural splitting with a maximum size limit, falling back to a sub-split within an oversized function when needed.
Try it yourself
Add a def divide(a, b): function to source, after Calculator, and re-run.
A new chunk 3 should appear, cleanly separated from Calculator. Function boundaries in the syntax tree are unaffected by whatever comes before or after them in the file — that isolation is what a syntax-aware splitter buys you over a character-count splitter.
What to learn next
- Chunking by document structure — natural boundaries for prose, not code.
- Answering questions about tables — what to do once a table chunk is retrieved.
- Hugging Face — libraries with production-grade code and document parsers.
Researcher — Mathematics and papers.
Code as a structured, not linear, object
Prose chunking treats a document as a 1-dimensional character stream. Source code is not linear in the sense that matters — its true structure is a syntax tree, and any chunk boundary that ignores the tree risks producing a syntactically invalid fragment. Splitting on ast node boundaries, as in the developer block, guarantees every chunk is independently parseable. Production code-search and code-RAG systems generally go further, using a language-agnostic parser generator — most commonly tree-sitter — so the same boundary-respecting approach works across languages without one hand-written parser per language.
Retrieval implications for code
A retrieved function without its surrounding class context, or without the imports it depends on, is frequently uninterpretable to a model even though it is syntactically valid on its own. This motivates hierarchical retrieval for code specifically: retrieve at function granularity, but supply file-level or module-level context alongside it, mirroring the parent-document retrieval pattern covered later in this section.
Tables and structured-data retrieval
Naive row-batch chunking, as shown in the developer block, is workable for point-lookup questions — "what is Widget C's Q1 revenue" — and fails for aggregate questions spanning many rows, since no single chunk contains the full column needed to sum or average. Herzig et al. (2020), TAPAS: Weakly Supervised Table Parsing via Pre-training (arXiv:2004.02349), addresses this differently: rather than chunking a table for retrieval, TAPAS encodes the entire table jointly with the question, using row and column embeddings, and predicts cell selections and aggregation operations directly. This is a fundamentally different strategy than retrieve-then-read — it treats the table as one structured unit, small enough to fit in a single model call, rather than a document to be chunked at all. That strategy stops working once the table exceeds the model's input limit, which real spreadsheets and database exports routinely do — at that point chunking again becomes necessary, and the aggregate-question problem returns.
Complexity
Parsing source into an AST is O(n) in file length for a well-formed file. Table chunking is O(rows). Both are cheaper than semantic chunking's per-unit embedding cost, since neither requires a model call.
Key references
- Herzig, J. et al. (2020). TAPAS: Weakly Supervised Table Parsing via Pre-training. arXiv:2004.02349
Current state and open problems
Code retrieval increasingly uses embedding models trained specifically on code, rather than general-purpose text embedding models, since code identifiers and structure carry different statistical regularities than prose. Table retrieval remains split between two unresolved strategies: chunk-and-retrieve, which handles arbitrarily large tables but struggles with aggregation, and whole-table encoding, which handles aggregation but hits a hard size ceiling. No single deployed system fully unifies both, and picking between them, per document type, remains a manual design decision.
What to learn next
- Answering questions about tables — the aggregation problem this lesson sets up.
- Retrieve small, return big — supplying surrounding context to a retrieved code chunk.
- FAISS — indexing many small chunks efficiently at scale.