Multimodal AI

Vision-language models

A vision-language model turns an image into tokens a language model can read, so you can chat about a picture in ordinary words.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The cost nobody mentions first
  6. Where you have already used one
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A vision-language model is a language model with an extra front door for pictures.

It reads an image and a question together, then replies in words.

Short name: VLM — when a chat app answers questions about a screenshot you pasted, this is what it is.

The analogy you have already lived

Think about a friend on the phone describing a photo to you. You cannot see it. They translate what they see into words, and you reason about those words.

Now imagine the translation is not into English but into a private shorthand only your brain understands. You still reason perfectly, because the shorthand arrives in the format your thinking already uses.

A VLM does exactly that. A vision encoder converts the picture into the language model's private shorthand, and the language model takes it from there.

Why it exists

CLIP can match a picture to a sentence. It cannot write you a paragraph, follow an instruction, or answer a follow-up question. Scoring is all it does.

Meanwhile language models had become genuinely good at instructions and conversation, and had never seen anything.

So the natural move was to bolt them together. Keep the vision encoder that already understands pictures. Keep the language model that already understands instructions. Build a small bridge between them, and train only the bridge at first.

That bridge is the whole invention. The two big parts already existed.

How it works

A language model reads tokens — small chunks of text, each turned into a number list. A VLM converts the picture into number lists of exactly the same shape, then slots them into the sequence.

   picture -> [vision encoder] -> patch vectors -> [projector] -> image tokens
                                                                       |
   "What is in this picture?"  ->  text tokens  ------------------\    |
                                                                   v   v
                                            [ language model reads one sequence ]
                                                          |
                                                    "A dog on grass."

The projector is a small piece of maths that reshapes vision output into the language model's token width. It is the only genuinely new component.

The language model never learns that some tokens are "images". They arrive in the same format as words, so it treats them the same way.

The cost nobody mentions first

An image is not one token. It is hundreds.

A picture is cut into small squares called patches, and each patch becomes a token. A moderate-sized image becomes several hundred tokens before you have typed a single word.

That is why sending a photo costs far more than sending a sentence. Higher resolution costs much more again. The bill and the delay both come from this.

Where you have already used one

  • Chat assistants that read a screenshot and explain an error message.
  • Accessibility apps that describe what a phone camera is pointed at.
  • Shopping apps that answer "will this fit a 15-inch laptop?" from a product photo.
  • Homework helpers that read a photographed maths problem.

What is honestly hard here

Fine detail. The image is squeezed into a few hundred tokens. Small text and small objects can vanish before the language model sees anything.

Confident invention. A VLM will describe objects that are not in the picture. The language half is a fluent writer, and fluency does not stop at the edge of the evidence. Researchers call this object hallucination, and it is measured, not anecdotal.

Size. These models are large. The smallest useful open ones run from a few hundred megabytes to a couple of gigabytes. The good ones are much bigger. This is not a quick install on a metered connection.

Remember this

  • A VLM = vision encoder + projector + language model.
  • The projector turns image patches into tokens the language model can read.
  • One image costs hundreds of tokens, which drives both price and latency.

What to learn next

Developer — Code and libraries.

The interesting part of a VLM is the seam. This example builds the seam end to end with NumPy, on an image small enough to print. You see exactly what the language model receives.

Setup

bash
pip install numpy

No downloads. Runs instantly.

Patches, projector, sequence

vlm_seam.py
import numpy as np

rng = np.random.default_rng(3)

# A 32x32 greyscale "photo": a bright square sitting in a dark field.
image = np.zeros((32, 32), dtype=np.float32)
image[8:20, 6:18] = 1.0

PATCH = 8
grid = image.shape[0] // PATCH
patches = (image.reshape(grid, PATCH, grid, PATCH)     # cut the picture into tiles
                .transpose(0, 2, 1, 3)
                .reshape(grid * grid, PATCH * PATCH))
print("patches:", patches.shape, "->", grid * grid, "image tokens")

# Stand-in vision encoder: patch pixels -> 12-dim patch features.
W_vision = rng.normal(scale=0.1, size=(PATCH * PATCH, 12))
patch_features = np.tanh(patches @ W_vision)

# The projector. This tiny matrix is the whole bridge between the two worlds.
LLM_DIM = 20
W_proj = rng.normal(scale=0.1, size=(12, LLM_DIM))
image_tokens = patch_features @ W_proj
print("image tokens after the projector:", image_tokens.shape)

# The language side: ordinary word embeddings, same width.
prompt = ["What", "is", "in", "this", "picture", "?"]
text_tokens = rng.normal(scale=0.1, size=(len(prompt), LLM_DIM))

sequence = np.vstack([image_tokens, text_tokens])
print("sequence handed to the language model:", sequence.shape)
print(f"  {len(image_tokens)} image tokens + {len(text_tokens)} text tokens")

# Which patches actually carry the square? Feature energy is a rough proxy.
energy = np.linalg.norm(image_tokens, axis=1)
print("\npatch energy grid (bright patches should stand out):")
print(np.round(energy.reshape(grid, grid), 2))

print("\nreal-model token cost of one image:")
for side, patch in [(224, 14), (336, 14), (672, 14), (1024, 14)]:
    n = (side // patch) ** 2
    print(f"  {side}x{side} at patch {patch}: {n:>4} tokens, "
          f"attention pairs ~ {n * n:,}")
Output
patches: (16, 64) -> 16 image tokens
image tokens after the projector: (16, 20)
sequence handed to the language model: (22, 20)
  16 image tokens + 6 text tokens

patch energy grid (bright patches should stand out):
[[0.   0.   0.   0.  ]
 [0.37 0.61 0.46 0.  ]
 [0.33 0.41 0.31 0.  ]
 [0.   0.   0.   0.  ]]

real-model token cost of one image:
  224x224 at patch 14:  256 tokens, attention pairs ~ 65,536
  336x336 at patch 14:  576 tokens, attention pairs ~ 331,776
  672x672 at patch 14: 2304 tokens, attention pairs ~ 5,308,416
  1024x1024 at patch 14: 5329 tokens, attention pairs ~ 28,398,241

What that output is telling you

The energy grid marks where the bright square is. Patches with no square in them produce all-zero features, and every patch with part of the square lights up. That is the crude beginning of grounding: some tokens carry the object and most do not.

Then look at the token table. Moving from 224 to 1024 pixels turns 256 tokens into 5,329. Attention cost grows with the square of that. The same image costs roughly 430 times more attention work at the higher resolution.

That single table explains almost every design choice in modern VLMs. Pooling image tokens, cropping into tiles, resampling to a fixed small number — all of it exists to fight this table.

Line by line, the parts that are not obvious

The reshape / transpose / reshape chain is the standard patch extraction trick. The first reshape splits both axes into (tile index, position within tile). The transpose groups the two tile indices together. The final reshape flattens each tile into a row. No pixel is copied twice.

W_proj has shape (12, 20) and that is the entire projector. In LLaVA the equivalent object maps 1024 vision dimensions to 4096 language dimensions. That is around 21 million parameters, tiny beside a 7-billion-parameter language model. Training only this piece is what made VLMs cheap to build.

np.vstack([image_tokens, text_tokens]) is the actual fusion. There is no special image attention mechanism here. The image becomes a prefix, and ordinary self-attention does the rest.

np.tanh stands in for a real vision transformer. Swap in a ViT and every shape in this file stays the same.

The three ways real VLMs bridge the gap

Linear or MLP projector. LLaVA's approach. Cheapest, keeps one token per patch, and the token count is whatever the encoder produces.

Resampler with learned queries. Flamingo's Perceiver Resampler and BLIP-2's Q-Former compress any number of patches into a fixed small set, often 32 or 64 tokens. Much cheaper for the language model, and it loses fine detail.

Cross-attention layers. Extra layers inside the language model attend to vision features directly, rather than inserting tokens into the sequence. Keeps the text sequence short, and requires editing the language model.

Running a real one

Open VLMs range from a few hundred megabytes to tens of gigabytes. A 7-billion-parameter model needs roughly 14 GB in half precision, more than most laptops have. On CPU it runs at seconds per token.

Start with the smallest checkpoint you can find, in the few-hundred-megabyte class. The API is the same whatever the size:

bash
pip install transformers torch pillow accelerate
run_vlm.py
from transformers import AutoProcessor, AutoModelForVision2Seq
from PIL import Image

MODEL = "HuggingFaceTB/SmolVLM-256M-Instruct"   # check the current size before you pull it
processor = AutoProcessor.from_pretrained(MODEL)
model = AutoModelForVision2Seq.from_pretrained(MODEL)

image = Image.open("your_photo.jpg")
messages = [{"role": "user", "content": [{"type": "image"},
                                         {"type": "text", "text": "Describe this image."}]}]
text = processor.apply_chat_template(messages, add_generation_prompt=True)
inputs = processor(text=text, images=[image], return_tensors="pt")

out = model.generate(**inputs, max_new_tokens=60)
print(processor.decode(out[0], skip_special_tokens=True))

No output block, deliberately. The text depends on your image, the checkpoint version and the sampling settings. Check the model card for the exact download size before you run this on a metered connection.

Common mistakes

Sending a huge image and paying for it. Most processors resize for you, and some tile large images into many crops, multiplying the token count. Print inputs["pixel_values"].shape and count what you are actually sending.

Expecting exact text reading. A VLM at 336 pixels cannot read six-point text. For documents, use the pipeline in document AI instead of asking a general VLM to squint.

Trusting a description with no verification. Ask the model to answer "not visible" when it cannot see something, and test that it does. Untested, it will invent.

Fine-tuning the whole stack on a small dataset. Train the projector first with both encoders frozen. That is the standard first stage precisely because it needs the least data.

Try it yourself

Change PATCH from 8 to 4 in the NumPy example. The token count goes from 16 to 64 and the energy grid gets sharper. Then work out what you gained, what it cost, and where the break-even sits for a photo of a printed page.

What to learn next

Researcher — Mathematics and papers.

Architecture, stated precisely

A VLM composes three functions. A vision encoder E_v maps an image to n patch features in R^{d_v}. A connector P maps those to m vectors in R^{d_l}, the language model's embedding width. A decoder LM autoregressively models p(y | [P(E_v(x)); embed(prompt)]).

The design space is m versus n, whether P is linear or attentive, which components are frozen, and whether the visual tokens enter as a prefix or through cross-attention.

The lineage

Frozen (Tsimpoukelli et al., 2021) showed that training only a vision encoder to emit a prefix for a frozen 7B language model yields few-shot multimodal behaviour. The key result was that the language model's in-context learning ability transfers to images without touching its weights.

Flamingo (Alayrac et al., 2022) added two mechanisms: a Perceiver Resampler compressing variable patch counts to 64 latents, and gated cross-attention layers interleaved into a frozen Chinchilla LM. The tanh gate is initialised at zero, so the model starts exactly as the original text LM and learns how much vision to admit. This made training stable on interleaved image-text web documents.

BLIP-2 (Li et al., 2023) introduced the Q-Former: 32 learnable queries attending to frozen ViT features, trained in two stages — representation learning against text, then generative pretraining against a frozen LLM. Compute-efficient because both large components stay frozen.

LLaVA (Liu et al., 2023) replaced all of that with a linear projection, later a two-layer MLP, and trained on GPT-generated instruction data. It matched or beat far more complex designs. The lesson the field took from it: connector sophistication matters less than instruction-tuning data quality.

Resolution and the token budget

Attention over n visual tokens costs O(n^2 d) per layer. With patch size p, n = (H/p)(W/p), so cost is quartic in the side length. Three families of fix are in use:

  • Tiling / AnyRes. Split a high-resolution image into fixed-size crops, encode each, and concatenate — LLaVA-1.6 and InternVL. Linear in area, and it breaks global context across tile boundaries.
  • Token compression. Resamplers, pooling, or merging redundant tokens. Bolya et al. (2022), Token Merging, shows large speedups with small accuracy loss.
  • Native dynamic resolution. Qwen2-VL (Wang et al., 2024) processes images at their native size with 2D rotary position embeddings, so token count scales with actual content rather than a fixed grid.

Training stages

The standard recipe has two or three stages:

  1. Alignment. Freeze encoder and LM; train the connector on caption data. Cheap, and it teaches the connector to speak the LM's embedding distribution.
  2. Instruction tuning. Unfreeze the LM (often with LoRA — see LoRA); train on multi-turn visual instruction data. This is where nearly all task ability appears.
  3. Optional high-resolution or preference stage. Unfreeze the vision encoder, or run DPO against hallucination-labelled preferences.

Object hallucination, measured

Rohrbach et al. (2018) defined CHAIR, the fraction of mentioned objects absent from the ground-truth annotation, computed per instance and per sentence. Li et al. (2023) introduced POPE, which reframes hallucination as balanced yes/no questions about object presence and samples negatives three ways: at random, from frequent objects, and from objects that co-occur with the present ones. Accuracy drops sharply under the co-occurrence sampling, which is the diagnostic result: the language prior is generating the object, not the image.

Mitigations with published gains include contrastive decoding against a distorted image (VCD, Leng et al., 2023), preference optimisation on hallucination-labelled pairs, and raising input resolution. None of them solves it.

Evaluation

Benchmarks in common use: MMMU (college-level multi-discipline reasoning), MMBench, SEED-Bench, MathVista, ChartQA, DocVQA, TextVQA and POPE. Treat single-number leaderboard positions with suspicion — contamination is widespread, and several benchmarks are answerable at above-chance rates from the question text alone.

The honest evaluation for a product is a held-out set of your own images with your own questions, scored by humans at least once.

Cost

A 7B-parameter VLM in bf16 needs roughly 14 GB for weights plus KV cache. With 576 image tokens, the prefill is dominated by the image, and time-to-first-token is where users feel the resolution decision. Quantisation to 4 bits brings weights to roughly 4 GB, which is the practical route to running one on consumer hardware — see quantisation in practice.

Papers

What to learn next