Multimodal AI

CLIP

CLIP learns one shared space for images and text by playing a matching game on four hundred million picture-caption pairs.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The clever part at the end
  6. Where you have already seen it
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

CLIP is a pair of encoders — one that turns pictures into numbers, one that turns words into numbers — trained by playing a matching game.

You hand it a batch of photos and a shuffled batch of captions. Its job is to pair them up correctly.

The analogy you have already lived

Picture a school memory game. Thirty-two photos face down on one side. The same thirty-two captions face down on the other. Turn over one card from each side and see if they belong together.

Play that game a few hundred million times and you stop guessing. You start noticing what makes a photo and a sentence belong together.

Now the important twist. Nobody ever told you the rules of "belonging". You worked them out from which pairings were marked right and which were marked wrong.

That is the entire training procedure.

CLIP stands for Contrastive Language-Image Pre-training. Contrastive means it learns by comparing right answers against wrong ones.

Why it exists

Before CLIP, teaching a model to recognise something meant collecting labelled photos of it. A thousand pictures of a rickshaw, hand-labelled "rickshaw". Then a thousand of something else.

That does not scale, and it locks the model into a fixed list. A model trained on a thousand categories can never name the one-thousand-and-first thing.

CLIP's move was to stop using labels and start using the captions people already write. The internet is full of pictures with text beside them. That text is free, plentiful and open-ended.

The payoff is zero-shot classification — sorting pictures into categories the model was never trained on. You write the categories as sentences at the moment you need them.

How it works

Take a batch of photos with their real captions. Encode all of them. Then score every photo against every caption.

                 caption 1   caption 2   caption 3
   photo 1          HIGH        low         low
   photo 2          low         HIGH        low
   photo 3          low         low         HIGH

The correct pairs sit on the diagonal. Training pushes those scores up and pushes everything off the diagonal down.

Two things happen at once. The right caption is pulled closer. Every wrong caption in the batch is pushed away. That second half is what "contrastive" means, and it is why bigger batches teach more per step.

The clever part at the end

Once trained, you can classify without any classifier.

Want to sort photos into "cat", "dog" and "neither"? Write three sentences: "a photo of a cat", "a photo of a dog", "a photo of something else". Encode them. Score each photo against all three. Take the highest.

You never trained on cats. You wrote a sentence.

Where you have already seen it

  • Photo apps that search your gallery by description.
  • Image generators, which use a CLIP-style text encoder to understand your prompt.
  • Content filters that flag pictures matching a written policy.
  • Stock and shopping search that matches a phrase to a picture.

What is honestly hard here

CLIP is a matching machine, not a reasoning machine. It scores a whole picture against a whole sentence, with no notion of parts.

So it struggles with counting, with left versus right, and with word order. "A cat under a table" and "a table under a cat" often score almost the same. That is a real limitation of the design, not a bug someone forgot to fix.

It also inherits whatever the internet's captions contain. Studies found CLIP-style models carry the biases of their caption sources, and those biases show up in retrieval results. Any system you build on top inherits them too.

Remember this

  • CLIP learns by matching photos to captions inside a batch, right pairs up and wrong pairs down.
  • After training, an image and a sentence live in one shared space.
  • Zero-shot classification comes free: write the categories as sentences.

What to learn next

Developer — Code and libraries.

The CLIP loss is short enough to write from scratch. Doing that once removes all the mystery. The version below trains real projections with real gradients, on data small enough to run instantly.

Setup

bash
pip install numpy

Nothing is downloaded. The whole training run takes under two seconds on a laptop CPU.

The contrastive loss, written out

tiny_clip.py
import numpy as np

rng = np.random.default_rng(7)
N, D_IMG, D_TXT, D = 8, 10, 7, 6      # 8 pairs per batch, 6-dim shared space
TEMP = 0.07                            # CLIP learns this; pinned here so results repeat

meaning = rng.normal(size=(N, D))
img_feats = meaning @ rng.normal(size=(D, D_IMG)) + 0.2 * rng.normal(size=(N, D_IMG))
txt_feats = meaning @ rng.normal(size=(D, D_TXT)) + 0.2 * rng.normal(size=(N, D_TXT))

W_img = 0.1 * rng.normal(size=(D_IMG, D))   # the two projection heads we train
W_txt = 0.1 * rng.normal(size=(D_TXT, D))

def unit(u):
    return u / np.linalg.norm(u, axis=1, keepdims=True)

def softmax(x, axis):
    e = np.exp(x - x.max(axis=axis, keepdims=True))
    return e / e.sum(axis=axis, keepdims=True)

def unit_backward(u, dz):
    # gradient through L2 normalisation: strip the component pointing along z
    n = np.linalg.norm(u, axis=1, keepdims=True)
    z = u / n
    return (dz - (dz * z).sum(axis=1, keepdims=True) * z) / n

eye = np.eye(N)
for step in range(1, 2001):
    Ui, Ut = img_feats @ W_img, txt_feats @ W_txt
    Zi, Zt = unit(Ui), unit(Ut)
    logits = Zi @ Zt.T / TEMP              # every image scored against every caption
    P_rows, P_cols = softmax(logits, 1), softmax(logits, 0)
    loss = -0.5 * (np.log(np.diag(P_rows)).mean() + np.log(np.diag(P_cols)).mean())

    dlogits = 0.5 * ((P_rows - eye) + (P_cols - eye)) / N
    W_img -= 5.0 * img_feats.T @ unit_backward(Ui, dlogits @ Zt / TEMP)
    W_txt -= 5.0 * txt_feats.T @ unit_backward(Ut, dlogits.T @ Zi / TEMP)

    if step in (1, 100, 500, 2000):
        acc = (logits.argmax(axis=1) == np.arange(N)).mean()
        print(f"step {step:>4}   loss {loss:.4f}   image->caption top-1 {acc:.2f}")

S = unit(img_feats @ W_img) @ unit(txt_feats @ W_txt).T
print("\ncaption->image top-1:", (S.argmax(axis=0) == np.arange(N)).mean())
print("similarity matrix (rows = images, columns = captions):")
print(np.round(S, 2))
Output
step    1   loss 8.7444   image->caption top-1 0.12
step  100   loss 1.4098   image->caption top-1 0.38
step  500   loss 0.6720   image->caption top-1 0.88
step 2000   loss 0.3402   image->caption top-1 1.00

caption->image top-1: 1.0
similarity matrix (rows = images, columns = captions):
[[ 0.58  0.37 -0.29 -0.14  0.14  0.    0.34  0.02]
 [ 0.22  0.58 -0.72 -0.28  0.08  0.08  0.52  0.21]
 [ 0.06 -0.42  0.66  0.32  0.06 -0.02 -0.34 -0.34]
 [-0.39 -0.12  0.23  0.57  0.34  0.44 -0.02 -0.51]
 [ 0.22  0.32 -0.12  0.42  0.5   0.46  0.39 -0.5 ]
 [-0.02  0.28 -0.17  0.44  0.46  0.52  0.34 -0.44]
 [ 0.34  0.51 -0.42  0.13  0.37  0.31  0.53 -0.24]
 [-0.45 -0.16 -0.12 -0.34 -0.39 -0.25 -0.24  0.47]]

Small differences in the last decimal are possible across NumPy builds, because matrix multiplication may sum in a different order. The trend and the accuracies will match.

Read the matrix carefully

Every row's largest value sits on the diagonal, so both accuracies reached 1.00. That is the target, and it is worth noticing what the target is not.

The off-diagonal values are far from zero. Row 1 has 0.52 in column 6, against 0.58 on the diagonal. The loss never asked for a clean identity matrix — it asked for the diagonal to win. Ranking is the objective, not separation.

That distinction explains a production surprise. A CLIP model with excellent top-1 accuracy can still return near-identical scores for the top few candidates. A fixed cut-off then behaves erratically.

Line by line, the parts that are not obvious

logits = Zi @ Zt.T / TEMP builds the full N-by-N score grid in one multiply. Dividing by a small temperature sharpens the softmax. At 0.07 a similarity difference of 0.1 becomes a logit difference of about 1.4, which is what makes the gradient decisive. Real CLIP learns this value and clamps it to keep training stable.

P_rows and P_cols are two softmaxes over the same grid — one along rows, one along columns. The loss averages the two, which is why CLIP is called symmetric. Drop one and the space degrades in that direction.

The dlogits line is the gradient of softmax cross-entropy: predicted probability minus the one-hot target. Both directions share one grid, so their gradients add.

unit_backward is the piece most hand-written implementations get wrong. After normalising, any gradient component pointing along the vector itself changes only its length, which normalisation immediately undoes. Removing that component is required, not an optimisation.

Using the real CLIP

The smallest official checkpoint is openai/clip-vit-base-patch32, around 150 million parameters. Expect a download of roughly 600 MB the first time, cached afterwards. It runs on CPU in a second or two per image. That is a real download on a mobile connection, so start it when you have the bandwidth.

bash
pip install transformers torch pillow
real_clip.py
from transformers import CLIPModel, CLIPProcessor
from PIL import Image

model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")

image = Image.open("your_photo.jpg")
labels = ["a photo of a dog", "a photo of a cat", "a photo of a bus"]

inputs = processor(text=labels, images=image, return_tensors="pt", padding=True)
probs = model(**inputs).logits_per_image.softmax(dim=1)

for label, p in zip(labels, probs[0].tolist()):
    print(f"{p:.3f}  {label}")

No output block here, deliberately. The numbers depend on your photo, and printing an invented set would teach you to expect something that will not happen. On a clear photo of a dog, expect the dog line above 0.9 and the others near zero.

Common mistakes

Batch size too small. Every other item in the batch is a negative example. A batch of 8 gives 7 negatives per image, which teaches very little. CLIP itself used 32,768. If you train your own, batch size is the first knob, ahead of model size.

Forgetting the prompt template. "a photo of a dog" beats a bare "dog" on almost every benchmark, because training captions are sentences, not single words. Averaging the embeddings of several templates gains another point or two.

Comparing scores across images. logits_per_image.softmax(dim=1) normalises across your label list. Feed a picture of a bicycle with only cat and dog labels and one of them still gets a high probability. Add a rejection label, or set a floor on the raw similarity.

Treating CLIP as an object detector. It scores whole images. For "is there a helmet in this photo", crop candidate regions first or use a detector. Whole-image similarity misses small objects entirely.

Try it yourself

Change N from 8 to 3 in the tiny example and rerun. The loss falls faster and the final accuracy is meaningless, because there are only two negatives to learn from. Then raise it to 32. Watch how many more steps the same accuracy needs, and why that result is worth more.

What to learn next

Researcher — Mathematics and papers.

The objective

For a batch of N pairs, with L2-normalised embeddings z_i for images and z_t for texts and a learned temperature tau:

logits[i][j] = ( z_i[i] . z_t[j] ) / tau

L_i2t = -(1/N) * sum_i  log(  exp(logits[i][i]) / sum_j exp(logits[i][j])  )
L_t2i = -(1/N) * sum_j  log(  exp(logits[j][j]) / sum_i exp(logits[i][j])  )

L = (L_i2t + L_t2i) / 2

z_i[i] is the projected, normalised image embedding of pair i. z_t[j] is the projected, normalised text embedding of pair j. tau is a scalar temperature, parameterised in CLIP as exp(t) with t learned and clamped to keep 1/tau below 100. N is the batch size. The diagonal entries are the positives; everything else is treated as a negative.

This is InfoNCE (van den Oord et al., 2018) applied symmetrically across two views. Minimising it maximises a lower bound on the mutual information between the two modalities, and the bound is capped at log N — which is the formal reason batch size matters so much.

What CLIP actually did

Radford et al. (2021) trained on WIT, 400 million image-text pairs collected from the public web, using 500,000 queries and capping each query at 20,000 pairs to balance the distribution. Neither the dataset nor the collection code was released.

Architectures: a ResNet-50 family and a ViT family for images, and a 63M-parameter Transformer for text, with the end-of-sequence token's final layer taken as the text embedding. Both towers project to a shared 512-dimensional space (768 for the largest ViT). Batch size was 32,768, sharded across GPUs with an all-gather so every device sees the full negative set.

The headline result: ViT-L/14 at 336 pixels matches a fully supervised ResNet-50 on ImageNet without seeing a single ImageNet training label. The more important result is the robustness gap — on distribution shifts such as ImageNet-A, ImageNet-R and ObjectNet, CLIP degrades far less than supervised models with equal in-distribution accuracy.

Prompt engineering as a measurable effect

The paper reports that switching from a bare class name to the template "a photo of a {label}." gains roughly 1.3 points of ImageNet accuracy, and ensembling 80 templates by averaging their normalised text embeddings gains about 3.5 points in total. This is a training-free change that costs one extra forward pass per class.

The mechanism is a distribution-shift correction. Web captions are sentences; a bare noun sits outside the text encoder's training distribution.

The temperature, in detail

A logit scale of 1/tau = 100 is CLIP's clamp. Larger scales sharpen the softmax and concentrate gradient on the hardest negative; smaller scales spread it. Wang and Liu (2021) show the temperature controls a hardness-aware property of the loss — low temperature makes it approach a hard-negative-mining objective, which improves discrimination but hurts tolerance to semantically similar negatives.

That trade-off has a practical face. A very low temperature pushes apart two near-duplicate captions in the batch that genuinely describe the same image. These false negatives are not modelled by the loss at all, and their frequency rises with batch size.

Known limitations, measured

  • Compositionality. ARO (Yuksekgonul et al., 2022) and Winoground (Thrush et al., 2022) show near-chance or below-chance performance on relation and order tasks.
  • Counting. Accuracy drops sharply beyond three or four instances. A pooled global embedding has no counting mechanism.
  • OCR and fine text. Base variants read large text and fail on small text; this drove the resolution increases in later models.
  • Bias. The paper's own Section 7 reports demographic disparities in zero-shot classification and cautions against surveillance use. Birhane et al. (2021), auditing LAION-400M, documented explicit and harmful content surviving automated filtering.

What came after

ALIGN (Jia et al., 2021) used 1.8 billion noisier pairs with no curation and matched or beat CLIP, arguing scale substitutes for filtering. OpenCLIP (Ilharco et al., 2021) reproduced the recipe on public LAION data and published the first honest scaling curves for this family. SigLIP (Zhai et al., 2023) replaced the batch softmax with a pairwise sigmoid loss, removing the all-gather and making small-batch training viable. EVA-CLIP and DFN pushed data filtering further, consistent with DataComp's finding that data selection dominates architecture at fixed compute.

Cost

Original CLIP training: 592 V100 GPUs for 18 days for the largest ViT model, and 256 V100s for 12 days for the largest ResNet. Inference is cheap by comparison — one ViT-B/32 forward pass is roughly 4.4 GFLOPs per image, which is why the gallery side is precomputed and only the query is encoded live.

Papers

What to learn next