AI glossary

Chunking

In one sentence Chunking is splitting documents into retrieval-sized pieces for RAG — the unglamorous decision that quietly determines answer quality.

By Updated

Chunking is cutting documents into smaller pieces before embedding them, so a RAG system can retrieve the relevant piece instead of the whole file.

A pharmacist does not hand you the entire medical encyclopedia when you ask about one tablet — she reads you the relevant paragraph from the right page. Chunks are those paragraphs. Retrieval works passage-by-passage, so before anything is searched, every document must be cut into pieces, and where you cut decides what can ever be found whole.

The tension is size. Chunks too small orphan their meaning — "the second dose doubles it" retrieved without the sentence naming the dose is noise. Chunks too large dilute the embedding (one vector must summarise many topics, matching none sharply) and waste context-window on padding. Common practice lands between 200 and 800 tokens, with 10-20% overlap between neighbours so sentences at a boundary exist intact in at least one chunk.

The craft that separates good pipelines from bad: cut along the document's own structure — headings, paragraphs, list items — rather than every N characters mid-sentence. Keep tables and code blocks whole. Attach metadata (title, section, source) to every chunk. Two refinements worth knowing: parent-child retrieval (search small precise chunks, hand the model the larger surrounding section) and contextual chunking (prepend a one-line summary of where the chunk came from). When answers degrade, teams instinctively swap models; inspecting what retrieval actually returned — and re-chunking — is more often the fix.

Where to go next