LlamaIndex
LlamaIndex loads your documents, cuts them into searchable pieces and finds the right piece when a question arrives, so a model can answer from your files.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
LlamaIndex takes your own documents and makes them searchable, so a model can answer questions about them.
The analogy you have already lived
You have flipped to the back of a textbook and run your finger down the index. You found "photosynthesis — page 214" and turned straight to that page.
Nobody read the whole book to find that page. The index did the work. Somebody went through every page in advance and wrote down what was where.
LlamaIndex builds that index for your files. Your PDFs, your class notes, your company's policy documents. Then when a question arrives, it finds the right pages before the model answers.
Why it exists
RAG is the technique of looking things up before answering. The idea is simple. Building it is not.
Between "I have a folder of PDFs" and "the bot answers correctly" sit six separate jobs.
- Read the files. PDF, Word, Markdown, a website, a database — each needs different code.
- Cut them into pieces. A whole 80-page manual will not fit into one question.
- Turn each piece into numbers so pieces can be compared by meaning.
- Store those numbers somewhere you can search fast.
- Find the best few pieces when a question comes in.
- Paste them into a prompt with the question, and ask the model.
Every team was writing those six steps again from scratch. LlamaIndex packages all six, with sensible defaults, and lets you replace any one of them.
The five words that unlock everything
The documentation uses these constantly. They are not hard, but nobody defines them.
- Document — one whole file you loaded, with its text and where it came from.
- Node — one small piece of a document. Also called a chunk. This is the unit that gets searched.
- Index — the searchable structure built from all your nodes.
- Retriever — the part that takes a question and returns the closest nodes.
- Query engine — retriever plus model. Question in, written answer out.
How it works
your files (pdf, txt, docx, a website)
|
[ load ] -> Documents
|
[ split ] -> Nodes: small pieces, each remembering its source
|
[ index ] -> a structure you can search quickly
|
a question arrives
|
[retrieve] -> the two or three closest Nodes
|
[ answer ] -> the model writes, using those Nodes
|
an answer, plus the file and page it came fromNotice the last line. Because every node remembers where it came from, the answer can cite its source. That is the difference between a bot you can check and a bot you have to trust.
Where you have already seen it
- "Chat with your PDF" tools that answer questions about a file you uploaded.
- A college portal that answers admission questions from its own rulebook.
- Documentation sites with a search box that answers in sentences, not links.
The honest part
The cutting-up step decides everything, and there is no right answer. Cut too small and a chunk says "must be renewed annually" with no idea what "it" is. Cut too big and the useful sentence drowns among four irrelevant paragraphs. You will tune this by testing, not by reading.
PDFs are genuinely painful. A PDF stores where each letter is drawn, not what the sentences are. Two-column layouts, tables and scanned pages come out scrambled. Whole companies exist to solve this one problem.
The five-line demo hides real work. Every tutorial shows a tiny script that answers questions about a folder. It works. Then you point it at your actual documents and the answers are wrong. The reason is almost always upstream: the loading or the cutting, not the model.
Read that as encouragement, not warning. Once you know where the failure lives, you know where to look.
Remember this
- LlamaIndex turns your files into something a model can search.
- A node is one small piece of a document, and it is what gets found.
- When answers are wrong, suspect the loading and cutting, not the model.
What to learn next
- What is RAG? — the idea this library implements, from first principles.
- Vector databases — where the numbers get stored and searched.
- Chat with your PDF — build the whole thing end to end.
Developer — Code and libraries.
Setup
Install the core package only. It is small and pure Python.
pip install llama-index-coreAvoid pip install llama-index for now. That is a meta-package pulling in OpenAI clients, a vector store and more, and it will fail at import time if you do not have an API key.
Integrations are separate packages, one per thing: llama-index-llms-ollama, llama-index-embeddings-huggingface, llama-index-readers-file.
Step one: watch a document become nodes
Chunking is where RAG succeeds or fails, so look at it directly before anything else.
from llama_index.core import Document
from llama_index.core.node_parser import SentenceSplitter
policy = Document(
text=(
"Refunds are sent back to the original payment method. "
"A refund takes five working days to reach a UPI app. "
"Delivery inside India takes three to five working days. "
"You may return an unused item within thirty days of delivery. "
"Every laptop carries a one year warranty on manufacturing defects."
),
metadata={"source": "policy.txt"},
)
# Counting words instead of tokens keeps this offline and identical on every machine.
splitter = SentenceSplitter(
chunk_size=24,
chunk_overlap=8,
tokenizer=lambda text: text.split(),
)
nodes = splitter.get_nodes_from_documents([policy])
print("chunks:", len(nodes))
for i, node in enumerate(nodes):
print(f"--- chunk {i} (source={node.metadata['source']}) ---")
print(node.get_content())Metadata length (2) is close to chunk size (24). Resulting chunks are less than 50 tokens. Consider increasing the chunk size or decreasing the size of your metadata to avoid this. chunks: 3 --- chunk 0 (source=policy.txt) --- Refunds are sent back to the original payment method. A refund takes five working days to reach a UPI app. --- chunk 1 (source=policy.txt) --- Delivery inside India takes three to five working days. You may return an unused item within thirty days of delivery. --- chunk 2 (source=policy.txt) --- Every laptop carries a one year warranty on manufacturing defects.
The warning on the first line is real and harmless here. It fires because the toy chunk_size of 24 is unusually small next to the metadata. In production you would use a few hundred.
Two things in that output matter.
No sentence was cut in half. SentenceSplitter respects sentence boundaries, then packs whole sentences up to the limit. A splitter that chops at exactly N characters will slice through the middle of the one fact you needed.
source=policy.txt survived into every chunk. That metadata is what lets an answer cite its origin. Attach page numbers, section headings and dates here, at load time. You cannot recover them later.
Step two: retrieve
Now find the right chunk for a question. This uses BM25, a word-matching search that needs no model and no network, so the numbers below are reproducible.
pip install llama-index-retrievers-bm25from llama_index.core import Document
from llama_index.core.node_parser import SentenceSplitter
from llama_index.retrievers.bm25 import BM25Retriever
policy = Document(
text=(
"Refunds are sent back to the original payment method. "
"A refund takes five working days to reach a UPI app. "
"Delivery inside India takes three to five working days. "
"You may return an unused item within thirty days of delivery. "
"Every laptop carries a one year warranty on manufacturing defects."
),
metadata={"source": "policy.txt"},
)
splitter = SentenceSplitter(chunk_size=24, chunk_overlap=8, tokenizer=lambda t: t.split())
nodes = splitter.get_nodes_from_documents([policy])
retriever = BM25Retriever.from_defaults(nodes=nodes, similarity_top_k=2)
for question in ["How long does delivery take?", "Can I get my money back?"]:
print("Q:", question)
for hit in retriever.retrieve(question):
print(f" {hit.score:.3f} {hit.node.get_content()[:60]}...")
print()Metadata length (2) is close to chunk size (24). Resulting chunks are less than 50 tokens. Consider increasing the chunk size or decreasing the size of your metadata to avoid this. Q: How long does delivery take? 0.688 Delivery inside India takes three to five working days. You ... 0.183 Refunds are sent back to the original payment method. A refu... Q: Can I get my money back? 0.382 Refunds are sent back to the original payment method. A refu... 0.000 Every laptop carries a one year warranty on manufacturing de...
Scores may shift in the third decimal with a different BM25 version. The ranking is the point.
Read the second question carefully
"Can I get my money back?" found the refunds chunk. That looks like the system understood the question. It did not.
BM25 matched on the word "back", which appears in "sent back to the original payment method". A lucky word overlap, not comprehension. Rephrase it as "Will I be reimbursed?" and the score drops to zero, because "reimbursed" shares no words with anything.
Look at the second hit too: 0.000. The retriever returned a chunk with a score of zero, because it was told to return two. A query engine will happily paste that irrelevant chunk into the prompt.
Both problems have known fixes. Word matching is replaced by embeddings, which put "reimbursed" and "refund" near each other in number space. Zero-score results are dropped with a score floor. Neither is automatic — you add them.
Step three: the whole pipeline
With an embedding model and a local LLM, the six steps collapse into a handful of lines.
pip install llama-index-embeddings-huggingface llama-index-llms-ollamafrom llama_index.core import SimpleDirectoryReader, VectorStoreIndex, Settings
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.ollama import Ollama
# Set these BEFORE building the index, or LlamaIndex reaches for OpenAI by default.
Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")
Settings.llm = Ollama(model="llama3.2:1b", request_timeout=120.0)
documents = SimpleDirectoryReader("./docs").load_data()
index = VectorStoreIndex.from_documents(documents)
engine = index.as_query_engine(similarity_top_k=3)
answer = engine.query("How long does a refund take?")
print(answer)
for node in answer.source_nodes:
print(f" source: {node.metadata.get('file_name')} score: {node.score:.3f}")No output block here. The answer text depends on your files, your model and the run. bge-small-en-v1.5 is a 130 MB download the first time, and it runs on CPU.
answer.source_nodes is the part worth keeping. It tells you which chunks produced the answer. Print it during development, always. A confident answer built from the wrong chunks looks identical to a correct one until you look.
Common mistakes
Forgetting that OpenAI is the default. Build an index without setting Settings.embed_model and you get:
ImportError: `llama-index-embeddings-openai` package not found, please run `pip install llama-index-embeddings-openai`
That message sounds like a missing dependency. It is really telling you that you never chose an embedding model, and the default is a paid API.
Rebuilding the index on every run. Embedding a large folder takes minutes and, with a paid API, money. Persist it:
index.storage_context.persist(persist_dir="./storage")Then reload with load_index_from_storage.
Trusting SimpleDirectoryReader on complicated PDFs. It is fine for text-based PDFs and Markdown. Two-column layouts, scanned pages and tables come out as scrambled text, and your retrieval quality collapses for a reason invisible in the code. Print the loaded text before you index it.
Leaving similarity_top_k high. More context is not better. Accuracy on facts placed in the middle of a long prompt measurably drops. Three good chunks beat twenty mediocre ones.
Losing the heading path. A chunk reading "must be renewed annually" is useless alone. Put the section title into the chunk text itself, not only into metadata, so the retriever can match on it.
Try it yourself
Run chunks.py with chunk_overlap=0 and compare against chunk_overlap=8. Then extend the document with a fact deliberately split across a sentence boundary, and check which setting keeps it findable.
After that, take ten real questions in the words a real user would type, run them through retrieve.py, and count how many return the right chunk. That fraction is your retrieval recall. It is the honest health metric for any RAG system, and it is measurable before you involve a model at all.
What to learn next
- What is RAG? — the idea this library implements, from first principles.
- Vector databases — where the numbers get stored and searched.
- Chat with your PDF — build the whole thing end to end.
Researcher — Mathematics and papers.
The data model
LlamaIndex's core abstraction is the Node: a chunk of content with an id, metadata, and typed relationships (SOURCE, PREVIOUS, NEXT, PARENT, CHILD). Documents are nodes; chunks are nodes with a SOURCE relationship. The graph is what enables retrieval strategies that fetch a small node and expand to its neighbours or parent.
Metadata participates in retrieval in two configurable ways: excluded_embed_metadata_keys controls what is included in the embedded text, and excluded_llm_metadata_keys controls what reaches the prompt. These are distinct on purpose — a file path may help a human verify a citation while adding noise to the embedding vector.
Chunking strategies, and what the evidence supports
- Fixed-size with overlap. The baseline. Overlap of 10–20% bounds the probability that a fact spanning a boundary is lost from every chunk.
- Sentence-aware (
SentenceSplitter). Packs whole sentences to a budget. Strictly better than character slicing at equal cost. - Structure-aware (Markdown headings, HTML sections, code AST). Uses the document's own hierarchy. Best when it exists.
- Semantic chunking (
SemanticSplitterNodeParser). Embeds sentences, cuts where adjacent-sentence similarity drops below a percentile threshold. Intuitively appealing; published comparisons show inconsistent gains over sentence splitting, at substantially higher indexing cost. - Small-to-big (
SentenceWindowNodeParser,AutoMergingRetriever). Embed a small unit for retrieval precision, return a larger window for generation context. This decoupling is the most reliably useful idea in the list, because retrieval and generation genuinely want different granularities.
Response synthesis modes
Once k nodes are retrieved, they must be turned into an answer. The mode determines cost and failure behaviour.
| Mode | Calls | Behaviour |
|---|---|---|
compact | 1 (usually) | Pack nodes into as few prompt windows as fit. Default. |
refine | k | Answer from node 1, then iteratively revise with each further node. |
tree_summarize | O(k) | Hierarchical bottom-up summarisation. Best for "summarise everything". |
accumulate | k | Answer per node independently, concatenate. No cross-node reasoning. |
refine degrades characteristically: later nodes can overwrite a correct early answer, because the model is asked to revise rather than to arbitrate. Prefer compact unless the context genuinely will not fit.
Position effects constrain k
Liu et al. (2023) documented a U-shaped accuracy curve over the position of relevant information in a long context: strong at the beginning and end, measurably weaker in the middle. Two operational consequences.
Order retrieved nodes so the highest-scoring sit at the edges of the context, not in reading order. And treat "increase similarity_top_k" as a change that can reduce accuracy — it must be measured, not assumed.
Where quality actually comes from
Ranked by observed effect on end-to-end accuracy in deployed systems:
- Document parsing quality. Everything downstream is bounded by it. PDF table and multi-column extraction is the dominant unsolved practical problem.
- Chunk boundaries and metadata enrichment. Including the heading path in chunk text is a small change with a large measured effect.
- Hybrid retrieval. BM25 plus dense, fused with reciprocal rank fusion (Cormack et al., 2009). Sparse handles identifiers, product codes and rare proper nouns; dense handles paraphrase. They fail on disjoint query sets.
- Cross-encoder reranking. Retrieve
k ≈ 50–100cheaply, rerank to 3–5 with a joint query-passage scorer. Usually the highest quality-per-engineering-hour change available. - Query transformation. HyDE (Gao et al., 2022) embeds a hypothetical answer rather than the question, closing the vocabulary gap between short queries and long passages. Multi-query expansion fires several rewrites and fuses results.
- The generator model. Last, and by a distance. Retrieval failure dominates the error budget in almost every deployed system.
Evaluation
Measure the stages separately or you will debug the wrong one.
- Retrieval: hit rate and MRR against labelled question-to-node pairs. LlamaIndex ships
RetrieverEvaluatorand agenerate_question_context_pairshelper that synthesises a labelled set from your own corpus. - Faithfulness: is every claim entailed by the retrieved context?
FaithfulnessEvaluator, or RAGAS (Es et al., 2023). - Relevancy: does the answer address the question at all?
Reference-free LLM-judged metrics are useful for regression detection and weak as ground truth. Calibrate them against a human-labelled subset before trusting a number.
LlamaIndex versus LangChain
Both now cover both territories, but the design centres differ and it shows in the defaults. LlamaIndex's primitives are documents, nodes, indices and retrievers — the vocabulary of information retrieval. LangChain's are runnables and graphs — the vocabulary of pipeline composition.
For a document question-answering system, LlamaIndex's defaults are closer to correct out of the box. For a multi-step agent with tools and branching control flow, LangGraph is the better-shaped abstraction. Mixing them is common and supported.
References
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, 2020 — arxiv.org/abs/2005.11401
- Liu et al., Lost in the Middle, 2023 — arxiv.org/abs/2307.03172
- Gao et al., Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE), 2022 — arxiv.org/abs/2212.10496
- Es et al., RAGAS: Automated Evaluation of Retrieval Augmented Generation, 2023 — arxiv.org/abs/2309.15217
- Robertson and Zaragoza, The Probabilistic Relevance Framework: BM25 and Beyond, 2009
- Cormack et al., Reciprocal Rank Fusion, SIGIR 2009
What to learn next
- What is RAG? — the idea this library implements, from first principles.
- Vector databases — where the numbers get stored and searched.
- Chat with your PDF — build the whole thing end to end.