Multimodal AI

Multimodal RAG

Multimodal RAG lets a model look things up in pictures, charts, slides and video before answering, not only in text.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The lesson people learn the hard way
  6. Where you have already used it
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Multimodal RAG searches your pictures, charts, slides and video before the model answers.

Ordinary RAG searches only your text files. This one does not stop there.

The analogy you have already lived

Think about an open-book exam where the book has diagrams. You do not read the whole book. You flip to the right page, and half of what you need is a labelled drawing rather than a paragraph.

Now imagine somebody retyped that book for you, but skipped every diagram and wrote "there is a figure here" instead. You could still find the page. You could not answer the question.

That is exactly the difference between ordinary RAG and multimodal RAG.

Why it exists

Plain RAG searches text and answers from text. It works well when the answer is written in words.

A lot of the time it is not. Product manuals answer with exploded diagrams. Financial reports answer with charts. Training material answers with screenshots and video.

Run ordinary RAG over those files and the pipeline quietly drops the important part. Text extraction returns the captions and page numbers, and the answer is gone.

Multimodal RAG keeps the non-text content in the system, so it can be found and used.

How it works

   your documents (PDFs, images, slides, video)
        |
        +--> text passages ---------------\
        +--> chart and table contents -----\
        +--> image descriptions ------------>-- one searchable index
        +--> video moments with transcripts /
                                           |
   your question --------------------------+
                                           |
                                     top few pieces
                                           |
                                    [ model answers ]
                                           |
                            an answer, with what it used

Two designs are common, and they are worth telling apart.

Turn everything into text first. Describe images, read charts into numbers, transcribe video. Then run ordinary text search. Simple, cheap and it fails when the description missed the point.

Search the pictures directly. Put page images and photos into a shared space and search them with your typed question. Nothing is thrown away at indexing time, and the index is far bigger.

Most real systems mix the two.

The lesson people learn the hard way

The retrieval step decides everything, and it fails quietly.

If the right chart never comes back, no model on earth will produce the right answer. It will write something fluent from whatever it did receive.

So the first thing to measure is not answer quality. It is whether the correct piece of evidence was retrieved at all. Get that wrong and every other improvement is wasted work.

Where you have already used it

  • Support bots that answer with a diagram from the manual.
  • Search over your own notes and screenshots in a note-taking app.
  • Meeting tools that answer questions about what was shown on a shared screen.
  • Study apps that find the right slide from a lecture recording.

What is honestly hard here

Charts. A bar chart's meaning is in the numbers, and the numbers are drawn, not written. Extracting them correctly is unreliable.

Deciding what to store. One description per image loses detail. Many descriptions per image bloat the index and slow everything down. There is no setting that avoids the trade.

Showing your work. With text you can quote the sentence you used. With an image, honest citation means showing the picture and the region, and most systems skip that.

Remember this

  • Multimodal RAG looks things up in images and video, not only text.
  • Either describe everything as text, or search the images directly — usually both.
  • Retrieval failure is the main cause of bad answers, and it is silent.

What to learn next

Developer — Code and libraries.

The most valuable experiment here compares two things: what the system can retrieve, and what it can actually answer from. Those are different, and confusing them is the standard mistake.

Setup

bash
python3 --version

Standard library only. Runs instantly.

Retrieval succeeded, the answer is still impossible

mm_rag.py
import re
from collections import Counter

# A tiny corpus with three kinds of item. Each one carries what we can search.
CORPUS = [
    {"id": "policy.txt#3", "kind": "text",
     "caption": "Refunds are processed within five working days of approval.",
     "extracted": ""},
    {"id": "revenue.png", "kind": "image",
     "caption": "a bar chart of quarterly revenue",
     "extracted": "Q1 120 crore Q2 140 crore Q3 175 crore Q4 160 crore"},
    {"id": "wiring.jpg", "kind": "image",
     "caption": "a photo of a wiring diagram with red and black cables",
     "extracted": "warning disconnect mains before servicing"},
    {"id": "demo.mp4#02:14", "kind": "video",
     "caption": "a person resetting the router by holding the button",
     "extracted": "hold reset for ten seconds until the light blinks amber"},
    {"id": "manual.pdf#9", "kind": "text",
     "caption": "The router supports both 2.4 GHz and 5 GHz bands.",
     "extracted": ""},
]

STOP = {"a", "an", "the", "of", "in", "is", "are", "to", "was", "for", "with", "and",
        "what", "how", "do", "i", "my", "it", "on", "by", "before", "until"}

def words(text):
    return [w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP]

def search(query, fields):
    q = Counter(words(query))
    scored = []
    for item in CORPUS:
        bag = Counter(words(" ".join(item[f] for f in fields)))
        scored.append((sum(min(n, bag[w]) for w, n in q.items()), item["id"], item["kind"]))
    scored.sort(key=lambda row: (-row[0], row[1]))     # id breaks ties, so runs repeat
    return scored[:3]

for query in ["what was Q3 revenue", "how long do I hold the reset button"]:
    print(f'query: "{query}"')
    print("  caption only        :", search(query, ["caption"]))
    print("  caption + extracted :", search(query, ["caption", "extracted"]))
    print()

def build_prompt(query, fields, use_extracted):
    hits = [h for h in search(query, fields) if h[0] > 0]
    if not hits:
        return "No evidence found. Refuse to answer."
    lines = []
    for _, item_id, kind in hits:
        item = next(c for c in CORPUS if c["id"] == item_id)
        body = (item["extracted"] or item["caption"]) if use_extracted else item["caption"]
        lines.append(f"[{item_id}] ({kind}) {body}")
    return ("Answer using ONLY the evidence. Cite the id in square brackets.\n\n"
            + "EVIDENCE:\n" + "\n".join(lines) + f"\n\nQUESTION: {query}\nANSWER:")

Q = "what was Q3 revenue"
print("--- prompt built from captions alone ---")
print(build_prompt(Q, ["caption"], use_extracted=False))
print("\n--- prompt built with the chart's numbers extracted ---")
print(build_prompt(Q, ["caption", "extracted"], use_extracted=True))
Output
query: "what was Q3 revenue"
  caption only        : [(1, 'revenue.png', 'image'), (0, 'demo.mp4#02:14', 'video'), (0, 'manual.pdf#9', 'text')]
  caption + extracted : [(2, 'revenue.png', 'image'), (0, 'demo.mp4#02:14', 'video'), (0, 'manual.pdf#9', 'text')]

query: "how long do I hold the reset button"
  caption only        : [(1, 'demo.mp4#02:14', 'video'), (0, 'manual.pdf#9', 'text'), (0, 'policy.txt#3', 'text')]
  caption + extracted : [(3, 'demo.mp4#02:14', 'video'), (0, 'manual.pdf#9', 'text'), (0, 'policy.txt#3', 'text')]

--- prompt built from captions alone ---
Answer using ONLY the evidence. Cite the id in square brackets.

EVIDENCE:
[revenue.png] (image) a bar chart of quarterly revenue

QUESTION: what was Q3 revenue
ANSWER:

--- prompt built with the chart's numbers extracted ---
Answer using ONLY the evidence. Cite the id in square brackets.

EVIDENCE:
[revenue.png] (image) Q1 120 crore Q2 140 crore Q3 175 crore Q4 160 crore

QUESTION: what was Q3 revenue
ANSWER:

Compare the two prompts, not the two retrieval scores

Caption-only retrieval found the right file. revenue.png came first, scoring 1 on the word "revenue". By every retrieval metric, this query succeeded.

Now read the prompt it produced. The evidence is "a bar chart of quarterly revenue". The number 175 is nowhere in it. A model given that prompt has two options: refuse, or invent a number that looks right.

The second prompt contains the actual values, and the question becomes answerable.

Retrieval recall and answerability are different measurements. A dashboard showing 95 percent recall tells you nothing about whether the retrieved evidence contained the answer. Measure both, separately, on the same query set.

Line by line, the parts that are not obvious

Counter[missing_key] returns 0 without inserting the key. Without that, scoring a query would silently grow every document's bag.

sum(min(n, bag[w]) for w, n in q.items()) clamps each query word's contribution to how often it appears in the document. Repeating a word in the query cannot inflate a score.

scored.sort(key=lambda row: (-row[0], row[1])) sorts by score descending, then by id ascending. The tie-break is not decoration. Without it, two zero-scoring items return in insertion order, and your tests break whenever you add a document.

hits = [h for h in search(...) if h[0] > 0] is the score floor. The search function returns three results regardless of quality. The filter turns "no match" into a refusal instead of a hallucination.

The caption and extracted split models the real design choice. caption is what a describer wrote. extracted is what a chart parser, OCR pass or transcript produced. They serve different purposes: captions help you find the item, extracted content lets you answer from it.

What to store per item, in practice

For each non-text item, store several fields and search across them:

  • A short description for retrieval — one or two sentences, generated by a vision-language model.
  • Extracted literal content — OCR text, chart data series, table cells, transcript with timestamps.
  • Surrounding context — the caption printed beside the figure, the section heading, the slide title. This is the cheapest quality gain available and it is routinely skipped.
  • A stable id that can be shown to the user: file, page, and region or timestamp.

At answer time, hand the model the extracted content plus, if your model accepts images, the image itself. Descriptions are for finding; pixels and extracted values are for answering.

Two indexing strategies, honestly compared

Describe then search textSearch page images directly
Index sizeSmallLarge, often many vectors per page
Setup costOne caption pass per itemOne embedding pass per page
Loses informationYes, at describe timeLess
Good forPhotos, mixed corporaSlides, scanned reports, charts
CitationPoints at the filePoints at the page, sometimes the region

The image-native route matured with ColPali-style models, and it removes the OCR and chunking stages entirely. Budget for the index size before choosing it.

Common mistakes

Measuring the answer instead of the retrieval. Build a small set of questions with known correct evidence ids. Report how often the right id is in the top k. That number bounds everything downstream.

One description per image. A photo answers different questions for different users. Store the description, the OCR text and the surrounding context as separate searchable fields rather than concatenating them into one blob.

Retrieving too much. Accuracy on facts placed in the middle of a long context measurably drops, and images are expensive tokens. Three good pieces beat twenty mediocre ones, and with images the cost difference is much larger than with text.

No citation the user can check. Show the page image or the video frame you used. With text you can quote; with images, showing is the only honest equivalent.

Reindexing everything on every change. Hash each source file and re-embed only what moved. Image embedding is far more expensive than text embedding, so this matters sooner than it does in a text pipeline.

Try it yourself

Add a sixth item: an image of a warranty table with caption set to "a table of warranty periods" and extracted left empty. Then ask "how long is the laptop warranty". Retrieval will find it and the prompt will be unanswerable — the same failure as the revenue chart, now created by you. Then fill in extracted and watch it work. That gap is where most multimodal RAG projects lose their time.

What to learn next

Researcher — Mathematics and papers.

What changes from text-only RAG

The generation stage is unchanged from what is RAG: condition on retrieved evidence, generate, cite. Three things change materially.

The index is heterogeneous. Items differ in modality, in granularity and in embedding cost. A page image, a paragraph and a 10-second video segment are not comparable units, and score normalisation across them is not solved by any standard method.

Retrieval can be modality-native or modality-collapsed. Either encode every item into a shared space and search it directly, or convert everything to text and search that. These have different failure modes, and hybrid systems need a fusion rule.

Evidence is expensive at generation time. An image costs hundreds of tokens. The context budget that holds twenty text chunks holds two or three images.

Two retrieval designs

Modality-collapsed. Generate a textual surrogate for each non-text item — caption, OCR, transcript, chart-to-table — then use a standard text retriever. The advantages are that all existing text retrieval machinery applies, including BM25, hybrid fusion and reranking, and the index stays small. The disadvantage is an information bottleneck at surrogate-generation time that no downstream component can undo.

Modality-native. Embed items in a shared space. MuRAG (Chen et al., 2022) built a memory of image-text pairs retrieved by a joint encoder and trained the reader on it. ColPali (Faysse et al., 2024) embeds page images with a vision-language model and uses ColBERT-style late interaction, outperforming OCR-plus-text pipelines on visually rich documents by a wide margin on ViDoRe.

The cost of late interaction is index size: roughly 1,000 vectors per page rather than one per chunk. Binary quantisation and pooling reduce this substantially, with a measurable recall cost worth quantifying on your own corpus.

Score fusion across modalities

Cosine similarities from different encoders are not calibrated to each other, and the modality gap (Liang et al., 2022) means image-text scores and text-text scores occupy different ranges. Two workable approaches:

  • Rank-based fusion. Reciprocal rank fusion, RRF(d) = sum over lists of 1 / (k + rank_i(d)) with k conventionally 60. Requires no calibration because it discards scores entirely.
  • Per-modality calibration. Fit a monotone map from raw score to estimated relevance probability using a labelled sample per modality, then merge in probability space. More work, and it preserves score magnitude information that RRF throws away.

Do not merge raw scores from different encoders without one of these. It is the most common silent bug in hybrid multimodal indexes.

Evaluation, staged

Evaluate three things separately, in this order:

  1. Retrieval recall@k against labelled query-to-evidence-id pairs, reported per modality. Aggregate recall hides the fact that images are usually retrieved much worse than paragraphs.
  2. Answerability of the retrieved set. Given only the retrieved evidence, can a human answer the question? This isolates the surrogate-generation bottleneck from the retrieval bottleneck, and it is the step almost every team skips.
  3. Grounding of the generated answer. Is every claim entailed by the retrieved evidence? For image evidence this needs a human or a strong VLM judge, and both need a calibrated reference set.

Benchmarks: WebQA (Chang et al., 2021) requires reasoning over image and text sources with a "which sources did you use" component. MMLongBench-Doc and M3DocVQA test long multi-page document QA. ViDoRe targets page-image retrieval specifically. InfographicVQA and ChartQA isolate the chart-understanding failure mode.

Chart and table understanding

This is the sharpest current weakness. Reading a bar chart requires recovering the data series from rendered marks, which is an inverse rendering problem, not an OCR problem. DePlot (Liu et al., 2022) reframes it as chart-to-table translation, then hands the table to a language model for reasoning — a decomposition that measurably beats end-to-end approaches, because the reasoning step is then a text problem.

The practical implication for a pipeline: store the recovered table, not a description of the chart, and store a confidence for the recovery.

Context economics

With n_img images at roughly 500 to 1,500 tokens each and text chunks at 300 to 800, the retrieval budget is dominated by images. Two consequences follow.

First, the top-k for images should be smaller than for text, tuned separately. Second, a rerank stage matters more here than in text RAG, because the cost of admitting a wrong image is far higher than the cost of admitting a wrong paragraph.

Position effects apply as they do in text: Liu et al. (2023) documented a U-shaped accuracy curve over evidence position, so ordering matters, and increasing k can reduce accuracy.

Papers

What to learn next