Multimodal AI

Image captioning

Image captioning is a model writing a sentence about a picture, one word at a time, guided by what the picture contains.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The failure you will see immediately
  6. Where you have already seen it
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Image captioning is a model looking at a picture and writing a sentence about it, one word at a time.

The analogy you have already lived

Think about describing a photo to someone over the phone. You do not produce the whole sentence at once. You start with "there's a...", see where that leads, and continue.

Each word you pick narrows what can come next. After "there's a red", you will not say "sitting". The sentence builds itself under two pressures: what you can see, and what sounds like proper language.

A captioning model works under exactly those two pressures. The picture votes for words. Grammar votes for words. The next word is where the votes agree.

Why it exists

The first reason is access. Millions of people use screen readers, and a photo with no description is a blank. Alt text is the written description attached to an image. It is often missing, and automatic captions fill some of that gap.

The second reason is search. Text is searchable and pictures are not. Turning a photo library into sentences makes it findable with ordinary words.

The third reason is that captioning was the first real test of whether a machine could connect seeing to saying. Answering a fixed question is easier than composing a free sentence.

How it works

   photo -> [encoder] -> what is in it
                             |
       "a"  <---------------- |
        |                     |
      "dog" <---------------- |   each word is chosen using both
        |                     |   the picture and the words so far
      "runs" <--------------- |
        |
      "." (stop)

The model repeatedly answers one question: given this picture and the words so far, what word comes next?

Picking the single highest-scoring word each time is called greedy decoding. It is fast and it commits early — a bad first word cannot be taken back.

The alternative is beam search, where several possible sentences are kept alive at once and the best complete one wins. Better sentences, more computation.

The failure you will see immediately

Captioning models drift towards safe, generic sentences. "A man riding a horse." "A group of people standing in a room."

The reason is honest and structural. Common sentences were common in training, so they score well on average. A caption that is bland is rarely badly wrong, and the training objective rewards being rarely wrong.

Getting specific captions out of a model is a real research problem, not a settings change.

Where you have already seen it

  • Alt text suggestions on social platforms for uploaded photos.
  • Screen readers describing an unlabelled image on a webpage.
  • Photo apps grouping pictures by generated descriptions.
  • Camera apps for blind users that narrate the scene aloud.

What is honestly hard here

Judging a caption is harder than writing one. Two people describing the same photo write different sentences, and both are correct. So a scoring system that compares against one "right answer" is measuring the wrong thing.

Models describe things that are not there. A photo of a kitchen counter attracts the word "sink", because kitchens usually have sinks. This is called object hallucination, and it is measured in research, not only complained about.

Remember this

  • A caption is built one word at a time, using the picture and the words so far.
  • Greedy decoding is fast and commits early; beam search keeps options open.
  • Captions drift towards generic sentences, and scoring them automatically is unreliable.

What to learn next

Developer — Code and libraries.

Captioning has two halves worth separating: the decoding loop, and the metric that judges it. This example implements both with the standard library so you can see where each one goes wrong.

Setup

bash
python3 --version

No dependencies at all. Runs instantly.

A decoder and a metric you can read

tiny_captioner.py
import math
from collections import Counter

VOCAB = ["a", "dog", "bus", "runs", "waits", "on", "the", "grass", "road", "<end>"]

# Stand-in image encoder output: how strongly the picture votes for each word.
IMAGE_VOTES = {
    "dog_on_grass.jpg": {"dog": 3.0, "grass": 3.0, "runs": 1.5, "a": 0.5},
    "bus_on_road.jpg":  {"bus": 3.0, "road": 3.0, "waits": 1.5, "a": 0.5},
}

# Stand-in language model: how well each word follows the previous one.
BIGRAM = {
    "<start>": {"a": 3.0},
    "a":       {"dog": 1.0, "bus": 1.0},
    "dog":     {"runs": 2.0, "waits": 1.0},
    "bus":     {"waits": 2.0, "runs": 0.5},
    "runs":    {"on": 2.0},
    "waits":   {"on": 2.0},
    "on":      {"the": 3.0},
    "the":     {"grass": 1.0, "road": 1.0},
    "grass":   {"<end>": 3.0},
    "road":    {"<end>": 3.0},
}

def caption(image, max_len=8):
    words, prev = [], "<start>"
    for _ in range(max_len):
        votes = IMAGE_VOTES[image]
        scores = {w: BIGRAM.get(prev, {}).get(w, -5.0) + votes.get(w, 0.0) for w in VOCAB}
        nxt = max(scores, key=scores.get)          # greedy: take the best word, never look back
        if nxt == "<end>":
            break
        words.append(nxt)
        prev = nxt
    return " ".join(words)

def bleu(candidate, references, max_n=2):
    cand = candidate.split()
    precisions = []
    for n in range(1, max_n + 1):
        cand_ngrams = Counter(tuple(cand[i:i + n]) for i in range(len(cand) - n + 1))
        best = Counter()
        for ref in references:
            r = ref.split()
            ref_ngrams = Counter(tuple(r[i:i + n]) for i in range(len(r) - n + 1))
            for g, c in ref_ngrams.items():
                best[g] = max(best[g], c)
        clipped = sum(min(c, best[g]) for g, c in cand_ngrams.items())
        precisions.append(clipped / max(sum(cand_ngrams.values()), 1))
    ref_len = min((abs(len(r.split()) - len(cand)), len(r.split())) for r in references)[1]
    bp = 1.0 if len(cand) > ref_len else math.exp(1 - ref_len / max(len(cand), 1))
    if min(precisions) == 0:
        return 0.0
    return bp * math.exp(sum(math.log(p) for p in precisions) / max_n)

for image in IMAGE_VOTES:
    print(f"{image}: {caption(image)!r}")

generated = caption("dog_on_grass.jpg")
references = ["a dog runs on the grass", "a brown dog running across a lawn"]
print(f"\nBLEU-2 for {generated!r}: {bleu(generated, references):.3f}")
print(f"BLEU-2 for 'a dog on the grass': {bleu('a dog on the grass', references):.3f}")
print(f"BLEU-2 for 'a bus waits on the road': {bleu('a bus waits on the road', references):.3f}")
Output
dog_on_grass.jpg: 'a dog runs on the grass'
bus_on_road.jpg: 'a bus waits on the road'

BLEU-2 for 'a dog runs on the grass': 1.000
BLEU-2 for 'a dog on the grass': 0.709
BLEU-2 for 'a bus waits on the road': 0.316

The last line is the point of this lesson

"A bus waits on the road" is completely wrong for a photo of a dog on grass. It scores 0.316.

It earns that score from "a", "on", "the" and the bigram "on the". Function words are shared by almost every English caption, so a totally wrong sentence starts from a non-zero floor.

Now compare the middle line. "A dog on the grass" is correct, if less detailed, and scores 0.709 rather than 1.000. It was punished for not matching the reference wording.

That is the honest state of automatic caption scoring. It punishes correct paraphrase and rewards shared filler. Any BLEU comparison between two systems that differ by less than a few points is noise.

Line by line, the parts that are not obvious

scores = {w: bigram + image_votes} is the whole model in one line. Real captioners add a language-model logit to an image-conditioned logit in exactly this shape, after a softmax rather than in raw score space.

The default of -5.0 for unseen bigrams is what keeps output grammatical. Without a penalty for impossible transitions, the image votes dominate and you get "dog grass runs".

max(scores, key=scores.get) is greedy decoding, and it is where beam search would go. Greedy takes the best word now; beam search keeps the k best partial sentences and compares completed ones. Beam search fixes cases where a strong first word leads into a dead end.

In bleu, best[g] = max(best[g], c) is clipping: a candidate n-gram can score at most as many times as it appears in any single reference. Without clipping, the caption "the the the the" scores highly.

The brevity penalty bp stops the model gaming precision by emitting two words. It only applies when the candidate is shorter than the closest reference.

Running a real captioner

Two mainstream open checkpoints, both CPU-runnable:

  • nlpconnect/vit-gpt2-image-captioning — a ViT encoder with a GPT-2 decoder, roughly 1 GB to download.
  • Salesforce/blip-image-captioning-base — stronger captions, also around the 1 GB mark.

Check the file sizes on the model card before pulling either on a metered connection. Both take a few seconds per image on CPU.

bash
pip install transformers torch pillow
real_caption.py
from transformers import BlipProcessor, BlipForConditionalGeneration
from PIL import Image

MODEL = "Salesforce/blip-image-captioning-base"
processor = BlipProcessor.from_pretrained(MODEL)
model = BlipForConditionalGeneration.from_pretrained(MODEL)

image = Image.open("your_photo.jpg").convert("RGB")
inputs = processor(image, return_tensors="pt")

greedy = model.generate(**inputs, max_new_tokens=25)
beam = model.generate(**inputs, max_new_tokens=25, num_beams=5)

print("greedy:", processor.decode(greedy[0], skip_special_tokens=True))
print("beam-5:", processor.decode(beam[0], skip_special_tokens=True))

No output block, deliberately. The caption depends on your photo and the checkpoint version, and an invented example would set a false expectation. Run it on five of your own photos and compare the two decoding modes.

Common mistakes

Reporting BLEU alone. Use CIDEr and SPICE alongside it, and read a random sample of outputs by hand. Metrics catch regressions; they do not establish quality.

Comparing BLEU across papers. Tokenisation, casing and the number of references all change the number. Only within-experiment comparisons mean anything.

Long beams producing shorter captions. Beam search favours high-probability sequences, and short sequences have higher probability. Add a length penalty or you will get terser captions as you raise the beam width.

Ignoring hallucination. Count how often a mentioned object is absent from the picture, on a sample you label yourself. This number correlates poorly with BLEU, which is the reason to measure it separately.

Try it yourself

Add "cat": 2.8 to the votes for dog_on_grass.jpg and rerun. The image now votes almost equally for two animals, and greedy decoding still commits to one with no hedging. Then add a bigram entry making "a cat" grammatical and see how a small language-side change flips the whole sentence.

What to learn next

Researcher — Mathematics and papers.

The objective

Captioning is conditional language modelling. Given image x and caption tokens y = (y_1, ..., y_T):

p(y | x) = product over t of  p(y_t | y_1..y_{t-1}, x)

L = - sum over t of  log p(y_t | y_1..y_{t-1}, x)

Training uses teacher forcing: the ground-truth prefix is fed at every step regardless of what the model predicted. This makes training parallel over t, and it creates exposure bias — at inference the model conditions on its own outputs, a distribution it never saw during training. Ranzato et al. (2015) named the problem and proposed sequence-level training as the fix.

The lineage

Show and Tell (Vinyals et al., 2015) fed a CNN's final pooled feature as the initial hidden state of an LSTM. Show, Attend and Tell (Xu et al., 2015) replaced that with attention over a spatial feature grid, and produced the attention maps that made grounding visible. Bottom-Up and Top-Down attention (Anderson et al., 2017) attended over region proposals from an object detector rather than a uniform grid, which held the state of the art for years.

Transformer decoders then replaced LSTMs, and the modern form is a vision encoder feeding a pretrained language decoder — BLIP (Li et al., 2022) and the vision-language models family. BLIP's contribution was as much data as architecture: CapFilt bootstraps a captioner and a filter to clean noisy web captions, and the cleaned data drives most of the gain.

Metrics, and what each one misses

BLEU-n (Papineni et al., 2002): clipped n-gram precision with a brevity penalty. Designed for translation with multiple references. Insensitive to meaning and heavily influenced by function words, as the developer block demonstrates.

METEOR (Banerjee and Lavie, 2005): unigram alignment with stemming and synonym matching, and a fragmentation penalty. Correlates better with human judgement than BLEU; slower and language-resource dependent.

ROUGE-L: longest common subsequence, recall-oriented. Rarely the right primary metric for captioning.

CIDEr (Vedantam et al., 2015): TF-IDF weighted n-gram cosine similarity across reference sets. The IDF term downweights n-grams common across the whole corpus, which is a direct answer to BLEU's function-word floor. This is why CIDEr is the standard headline metric on COCO.

SPICE (Anderson et al., 2016): parses candidate and references into scene graphs of objects, attributes and relations, and computes F1 over graph tuples. It measures propositional content rather than wording, and it is bounded by parser quality.

CLIPScore (Hessel et al., 2021): reference-free, uses a CLIP-style model's image-text similarity. Useful where references are unavailable, and it inherits every compositional weakness of the underlying model.

The consensus practice is to report CIDEr and SPICE together, with BLEU-4 for comparability with older work, and to treat any of them as a regression detector rather than as ground truth.

Decoding

Beam search maximises sequence log-probability, which is not the same as maximising quality. Two well-documented consequences:

  • Length bias. Log-probability is a sum of negative terms, so longer sequences score lower. A length penalty ((5 + |y|) / 6)^alpha (Wu et al., 2016) is the standard correction.
  • Diversity collapse. Beams share prefixes, so the k candidates are often minor variants. Diverse beam search (Vijayakumar et al., 2016) adds a dissimilarity term across groups.

Sampling methods used for open text — see temperature and sampling — produce more specific captions and more hallucinations. The trade-off is direct and worth measuring on your own data rather than assuming.

Sequence-level training

Cross-entropy optimises token-level likelihood while evaluation is sequence-level CIDEr. Self-critical sequence training (Rennie et al., 2016) closes the gap with REINFORCE, using the greedy decode's own CIDEr as the baseline:

grad L = - ( CIDEr(y_sampled) - CIDEr(y_greedy) ) * grad log p(y_sampled | x)

The greedy baseline requires no learned value function and is nearly free to compute. SCST reliably raises CIDEr by several points, and it also increases repetitive, metric-gaming phrasing. Inspect samples before shipping it.

Hallucination

CHAIR (Rohrbach et al., 2018) reports the fraction of mentioned objects not present in the ground truth, at instance and sentence level. Its central finding is that CHAIR is only weakly related to CIDEr — a model can improve on the standard metric while hallucinating more. Any captioning system used for accessibility should report both.

Datasets

MS COCO Captions (Chen et al., 2015) provides five references per image for 123,000 images and remains the default benchmark. Conceptual Captions (Sharma et al., 2018) and its 12M successor supply web-scale noisy alt text for pretraining. NoCaps (Agrawal et al., 2019) tests objects absent from COCO training, exposing how much of a captioner's ability is memorised vocabulary.

Papers

What to learn next