Video Understanding and Tracking

Re-identification embeddings

Turn each crop of a person into a short list of numbers so that two crops of the same person land close together, even from different cameras minutes apart.

On this page 8
  1. Why it exists
  2. How it works
  3. Two words you will see everywhere
  4. The trap nobody mentions
  5. Where you have already seen this
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A re-identification embedding turns a picture of a person into a short list of numbers. Two pictures of the same person produce similar lists.

Think about spotting a friend across a crowded railway platform. You do not read their face from that distance. You recognise the red jacket, the height, the way the bag hangs off one shoulder.

Ten minutes later, on a different platform, you spot them again by the same handful of cues. Your brain compressed a whole person into a few features and matched on those.

That compression is an embedding — a list of numbers standing in for something complicated. You met the idea in embeddings, where words became numbers. Here, people become numbers.

Why it exists

Trackers built on box positions have a hard limit. Two people walking close together cannot be told apart by position. Neither can someone who vanishes behind a wall for ten seconds.

When the boxes stop helping, only appearance is left. And a system needs a way to compare appearances that survives a change of camera, lighting and viewing angle.

Comparing raw pixels does not survive any of that. Two photos of the same person under different lights differ enormously pixel by pixel. So you need numbers that describe the person rather than the photograph.

How it works

   crop of a person          embedding
   ┌────────┐
   │   O    │      ──►     [0.31, -0.08, 0.55, ... ]     128 numbers
   │  /|\   │
   │   |    │
   │  / \   │
   └────────┘

   two crops of the SAME person  ->  the two lists point in a similar direction
   two crops of DIFFERENT people ->  the two lists point apart

The network is trained with a rule that sounds almost too simple. Show it three pictures: two of the same person, one of somebody else. Push the matching pair closer together, push the odd one further away. Repeat millions of times.

Nothing tells it which features to use. It settles on clothing colour, build, posture and gait because those are what survive a change of camera.

Two words you will see everywhere

Query is the picture you are searching with. Gallery is the collection you are searching in.

Ask "where else does this person appear?" and the system ranks the whole gallery by similarity. A good system puts every other picture of that person at the top of the list.

The trap nobody mentions

Different cameras produce systematically different pictures. One has a warm tint, another a cold one. One is overhead, another at eye level.

An untreated system learns the camera instead of the person. Every picture from camera one clusters together, and the ranking becomes worthless. The code below shows this failure in numbers, and shows a very cheap fix.

Where you have already seen this

  • Photo apps grouping pictures of the same person across years.
  • Shops measuring how many visitors return, without knowing any names.
  • Lost-child systems searching camera footage across a large station.
  • Sports analytics following a player across camera angles.

What is honestly hard here

Be clear-eyed about what this technology is.

It works, but only within its limits. It fails when people change clothes, and it struggles when a uniform makes everyone look alike. Most published systems are trained mostly on adults in one country's clothing, and generalise worse elsewhere.

It is also surveillance. The most widely used benchmark in this field was withdrawn by its own creators in 2019. People had been recorded without proper consent. The data was also being used for surveillance research. That withdrawal is part of the field's history and should be part of any decision to build with it.

Remember this

  • An embedding turns a person crop into a short list of numbers built for comparison.
  • Training pushes same-person pairs together and different-person pairs apart.
  • Cameras add their own signature, and an untreated system will match cameras, not people.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy torch

Run against numpy 1.26.4 and torch 2.5.1 on CPU. We use synthetic features so the metrics and the failure modes are visible without downloading a person dataset.

Building the metrics, and finding the camera bias

Re-identification is scored with two numbers. Rank-1 is the fraction of queries whose top gallery hit is correct. mAP averages the precision at every correct hit, so it rewards finding all appearances rather than one.

reid_metrics.py
import numpy as np

rng = np.random.default_rng(0)
N_ID, PER_CAM, D = 12, 4, 16

identity = rng.normal(0, 1, (N_ID, D))          # what each person truly looks like
cam_bias = rng.normal(0, 1, (2, D)) * 1.2       # each camera adds its own colour cast

feats, pids, cams = [], [], []
for pid in range(N_ID):
    for cam in (0, 1):
        for _ in range(PER_CAM):
            feats.append(identity[pid] + cam_bias[cam] + rng.normal(0, 1.0, D))
            pids.append(pid)
            cams.append(cam)
feats, pids, cams = np.array(feats), np.array(pids), np.array(cams)
print("gallery:", feats.shape, "|", N_ID, "identities seen from 2 cameras")


def l2(x):
    return x / np.linalg.norm(x, axis=1, keepdims=True)


def evaluate(x, keep):
    """Rank-1 and mAP, Market-1501 protocol: same-camera gallery images are ignored."""
    x, p, c = l2(x[keep]), pids[keep], cams[keep]
    sim = x @ x.T
    r1, aps = [], []
    for q in range(len(x)):
        valid = ~((p == p[q]) & (c == c[q]))            # also drops the query itself
        order = np.argsort(-sim[q][valid], kind="mergesort")
        hit = (p[valid][order] == p[q]).astype(int)
        r1.append(hit[0])
        prec = np.cumsum(hit) / np.arange(1, len(hit) + 1)
        aps.append((prec * hit).sum() / hit.sum())      # average precision for this query
    return np.mean(r1), np.mean(aps)


everything = np.ones(len(feats), bool)
r1, mAP = evaluate(feats, everything)
print(f"\nraw features            : Rank-1 {r1:.3f}   mAP {mAP:.3f}")

centred = feats.copy()
for c in (0, 1):
    centred[cams == c] -= feats[cams == c].mean(0)      # remove each camera's own mean
r1c, mAPc = evaluate(centred, everything)
print(f"after camera-mean removal: Rank-1 {r1c:.3f}   mAP {mAPc:.3f}")

q = 0
sim = l2(centred) @ l2(centred).T
valid = ~((pids == pids[q]) & (cams == cams[q]))
order = np.argsort(-sim[q][valid], kind="mergesort")[:6]
print(f"\nquery is identity {pids[q]} on camera {cams[q]}. Top 6 gallery hits:")
for rank, i in enumerate(order, 1):
    verdict = "correct" if pids[valid][i] == pids[q] else "WRONG  "
    print(f"  rank {rank}: identity {pids[valid][i]:2d} cam {cams[valid][i]}  "
          f"cosine {sim[q][valid][i]:+.3f}  {verdict}")
Output
gallery: (96, 16) | 12 identities seen from 2 cameras

raw features            : Rank-1 0.125   mAP 0.175
after camera-mean removal: Rank-1 0.802   mAP 0.670

query is identity 0 on camera 0. Top 6 gallery hits:
  rank 1: identity  0 cam 1  cosine +0.636  correct
  rank 2: identity  1 cam 1  cosine +0.459  WRONG  
  rank 3: identity  0 cam 1  cosine +0.455  correct
  rank 4: identity  1 cam 1  cosine +0.440  WRONG  
  rank 5: identity  7 cam 1  cosine +0.420  WRONG  
  rank 6: identity  7 cam 1  cosine +0.418  WRONG  

Three things in this output

Rank-1 goes from 0.125 to 0.802 by subtracting a mean. The camera bias is larger than the identity signal, so raw cosine similarity ranks by camera. Removing each camera's mean vector strips most of it out. Six lines of code, a six-fold improvement, and it is the first thing to try on any cross-camera system.

Rank-1 (0.802) is far above mAP (0.670), and that is normal. Rank-1 asks whether the single best hit is right. mAP asks whether all four other appearances of that person rank highly. Published re-identification results always show this gap, and mAP is the harder and more informative number.

Look at the qualitative list. The correct match is rank 1 with cosine 0.636. The second correct one is only rank 3, beaten by a wrong identity at 0.459. Identity 1 appears twice in the top four; it happens to look similar. This is what a 0.670 mAP looks like from the inside, and it is why a tracker using these embeddings still needs motion to break ties.

The Market-1501 protocol excludes same-camera gallery images. The valid mask does this. Without it, the easiest matches — same person, same camera, seconds apart — dominate the score and the number becomes meaningless. Every re-identification benchmark has this rule and every reimplementation must honour it.

Training an embedding, and the trap that comes with it

The standard loss is batch-hard triplet (Hermans et al., 2017): for each anchor, take the hardest positive and the hardest negative in the batch.

triplet_reid.py
import numpy as np
import torch
import torch.nn as nn

rng = np.random.default_rng(0)
N_ID, PER_CAM, D = 12, 4, 16
identity = rng.normal(0, 1, (N_ID, D))
cam_bias = rng.normal(0, 1, (2, D)) * 1.2
feats, pids, cams = [], [], []
for pid in range(N_ID):
    for cam in (0, 1):
        for _ in range(PER_CAM):
            feats.append(identity[pid] + cam_bias[cam] + rng.normal(0, 1.0, D))
            pids.append(pid)
            cams.append(cam)
feats, pids, cams = np.array(feats), np.array(pids), np.array(cams)

TRAIN_IDS, TEST_IDS = set(range(8)), set(range(8, 12))
is_train = np.array([p in TRAIN_IDS for p in pids])


def l2(x):
    return x / np.linalg.norm(x, axis=1, keepdims=True)


def evaluate(x, keep):
    """Rank-1 and mAP over one identity split, Market-1501 style."""
    x, p, c = l2(x[keep]), pids[keep], cams[keep]
    sim = x @ x.T
    r1, aps = [], []
    for q in range(len(x)):
        valid = ~((p == p[q]) & (c == c[q]))
        order = np.argsort(-sim[q][valid], kind="mergesort")
        hit = (p[valid][order] == p[q]).astype(int)
        r1.append(hit[0])
        aps.append(((np.cumsum(hit) / np.arange(1, len(hit) + 1)) * hit).sum() / hit.sum())
    return np.mean(r1), np.mean(aps)


test = ~is_train
print(f"{'features':32s}{'Rank-1':>9}{'mAP':>8}   (4 unseen identities)")
print(f"{'raw':32s}{evaluate(feats, test)[0]:9.3f}{evaluate(feats, test)[1]:8.3f}")

centred = feats.copy()
for c in (0, 1):
    centred[cams == c] -= feats[cams == c].mean(0)
print(f"{'camera-mean removed':32s}{evaluate(centred, test)[0]:9.3f}"
      f"{evaluate(centred, test)[1]:8.3f}")

torch.manual_seed(0)
X = torch.tensor(feats, dtype=torch.float32)
Xtr, Ytr = X[is_train], torch.tensor(pids[is_train])
net = nn.Sequential(nn.Linear(D, 32), nn.ReLU(), nn.Linear(32, 8))
opt = torch.optim.Adam(net.parameters(), lr=0.01)

for step in range(400):
    e = nn.functional.normalize(net(Xtr), dim=1)
    dist = torch.cdist(e, e)
    same = Ytr[:, None] == Ytr[None, :]
    hardest_pos = (dist * same).max(1).values          # furthest same-person pair
    hardest_neg = (dist + same * 1e6).min(1).values    # closest different-person pair
    loss = torch.relu(hardest_pos - hardest_neg + 0.3).mean()
    opt.zero_grad()
    loss.backward()
    opt.step()

with torch.no_grad():
    emb = nn.functional.normalize(net(X), dim=1).numpy()
print(f"{'batch-hard triplet, unseen ids':32s}{evaluate(emb, test)[0]:9.3f}"
      f"{evaluate(emb, test)[1]:8.3f}")
print(f"\nthe same model scored on the identities it trained on:")
print(f"{'batch-hard triplet, seen ids':32s}{evaluate(emb, is_train)[0]:9.3f}"
      f"{evaluate(emb, is_train)[1]:8.3f}   <- meaningless")
Output
features                           Rank-1     mAP   (4 unseen identities)
raw                                 0.156   0.324
camera-mean removed                 0.875   0.781
batch-hard triplet, unseen ids      0.438   0.444

the same model scored on the identities it trained on:
batch-hard triplet, seen ids        1.000   1.000   <- meaningless

The learned model lost, and that is the lesson

The triplet model scores 1.000 on the identities it trained on and 0.444 on unseen ones. Perfect on the training identities, worse than subtracting a mean on new ones.

It memorised eight people. That is what a triplet loss will do when you give it eight people and enough capacity, and no training loss curve will warn you. Re-identification is an open-set problem: the people at deployment were never in the training set, so training-identity scores carry no information at all.

Every re-identification benchmark enforces disjoint train and test identities for exactly this reason. Market-1501 trains on 751 identities and tests on 750 different ones. If your evaluation split shares identities with training, your number is fiction.

A trivial baseline beat a learned model. Camera-mean removal reaches 0.781 mAP against the triplet model's 0.444. With 8 training identities that outcome is expected; real re-identification models train on hundreds to thousands. Run the trivial baseline first and report it, because a learned model that cannot beat it is not earning its complexity.

Exact numbers depend on the PyTorch version's random number generator. The ordering — memorisation on seen identities, a large drop on unseen ones — is robust across seeds.

In a tracker

DeepSORT-style trackers keep a rolling gallery of the last k embeddings per track and use the minimum cosine distance across that gallery as the appearance cost. Combined with a motion gate, as covered in SORT, DeepSORT and ByteTrack:

python
import numpy as np

def appearance_cost(track_gallery, det_embeddings):
    """Smallest cosine distance from each detection to anything in the track's history."""
    g = np.array(track_gallery)                    # (k, D), already L2-normalised
    d = np.array(det_embeddings)                   # (n, D), already L2-normalised
    return 1.0 - (d @ g.T).max(axis=1)             # (n,), one cost per detection

gallery = np.eye(3)[[0, 1]]                        # two remembered looks
dets = np.array([[1., 0., 0.], [0., 0., 1.]])
print(np.round(appearance_cost(gallery, dets), 3))
Output
[0. 1.]

The first detection matches something in the gallery exactly, cost 0. The second matches nothing, cost 1. Taking the minimum over the gallery rather than the mean is deliberate: a person seen from the front and the back has two different looks, and averaging them produces a vector describing neither.

Common mistakes

Skipping L2 normalisation. Cosine similarity on unnormalised vectors is dominated by magnitude, which usually encodes crop size or brightness rather than identity. Normalise, always.

Sharing identities between train and test. See the output above.

Ignoring the same-camera exclusion rule. Your mAP will be inflated and incomparable with any published number.

Using a single stored embedding per track. People turn around. Keep a small gallery, take the minimum distance, and cap the gallery size.

Updating the gallery with every matched detection. An occluded or misassigned box poisons the gallery and the track drifts onto the wrong person. Update only on high-confidence, unoccluded matches.

Assuming an embedding trained on one dataset transfers. A model at 0.83 mAP on Market-1501 can drop to single-digit mAP on MSMT17. Domain shift in re-identification is severe and routinely under-reported.

Try it yourself

Increase N_ID to 200 and put 150 identities in the training split. Re-run and check whether the triplet model now beats camera-mean removal on unseen identities. That experiment is the whole argument for why re-identification datasets are measured in hundreds of identities.

What to learn next

Researcher — Mathematics and papers.

The task

Person re-identification is open-set retrieval. Given a query image or tracklet, rank a gallery so that images of the same identity appear first. Training and test identities are disjoint by construction, so the learned function must be a general similarity metric, not a classifier over known people.

Two evaluation regimes:

  • Image-based: single crops. Market-1501, MSMT17, CUHK03.
  • Video-based / tracklet: sequences per identity, allowing temporal pooling. MARS, DukeMTMC-VideoReID, LS-VID.

Video re-identification is the setting relevant to tracking, and temporal pooling gives a consistent gain: averaging embeddings across a tracklet suppresses per-frame noise from blur, partial occlusion and awkward poses.

Losses

Triplet with margin $m$:

$$ \mathcal{L} = \left[ \, d(f_a, f_p) - d(f_a, f_n) + m \, \right]_{+} $$

Where $f_a, f_p, f_n$ are anchor, positive and negative embeddings, $d$ a distance (usually Euclidean on L2-normalised vectors), and $[\cdot]_+$ the hinge.

The difficulty is sampling. Most random triplets satisfy the margin immediately and contribute zero gradient. Batch-hard mining (Hermans et al., 2017) constructs batches of $P$ identities with $K$ images each, and for each anchor selects the hardest positive and hardest negative within the batch:

$$ \mathcal{L}{BH} = \sum{i=1}^{P}\sum_{a=1}^{K} \left[ m + \max_{p} d(f_a^i, f_p^i) - \min_{j \ne i,\, n} d(f_a^i, f_n^j) \right]_{+} $$

Batch construction is the hyperparameter that matters. $P{=}16$, $K{=}4$ is a common choice. Mining globally hardest triplets across the whole dataset destabilises training, because the hardest examples are frequently label noise.

Softmax with label smoothing over training identities is a strong alternative and combines well with triplet loss. The "bag of tricks" baseline (Luo et al., 2019) — warmup learning rate, random erasing augmentation, label smoothing, a last-stride change from 2 to 1, and a batch-norm neck between the feature and the classifier — reaches competitive accuracy with a plain ResNet-50 and is the sanity baseline every new method should be compared against.

Circle loss, ArcFace and CosFace reformulate the objective with angular margins on a hypersphere, and generally match or beat triplet loss with less sampling sensitivity.

Metrics

CMC Rank-k is the probability that a correct match appears in the top $k$. Rank-1 is standard.

mAP averages the average precision over queries:

$$ \text{AP}(q) = \frac{1}{|G_q|}\sum_{k=1}^{|G|} P(k) \cdot \mathbb{1}[\text{rank } k \text{ is correct}] $$

Where $G_q$ is the set of correct gallery items for query $q$ and $P(k)$ is precision at rank $k$.

The distinction matters because CMC only credits the first correct hit. A method that retrieves one appearance and misses five others scores identically at Rank-1 to one that retrieves all six. mAP is the number to report, and the Rank-1-minus-mAP gap is itself diagnostic.

The Market-1501 protocol excludes gallery images sharing both identity and camera with the query, and marks a set of "junk" detections (bad crops, distractors) that are ignored. Reimplementations that skip these rules produce inflated, incomparable numbers.

Architectures

YearMethodIdea
2018PCBHorizontal part stripes with independent classifiers
2019BoT / "bag of tricks"Strong plain baseline; BN neck; random erasing
2019OSNetOmni-scale feature learning with a lightweight backbone
2021TransReIDViT backbone plus jigsaw patch shuffling and side information embedding
2023CLIP-ReIDVision-language pretraining adapted without concrete text labels

CLIP-ReID reports 63.0% mAP and 84.4% Rank-1 on MSMT17 with a CNN backbone, and 73.4% mAP / 88.7% Rank-1 with a ViT backbone; adding side information embedding and overlapping patches raises this to 75.8% mAP / 89.7% Rank-1. Vision-language pretraining also generalises better across domains than supervised-only models, which is the more useful property in practice.

MSMT17 is the discriminating benchmark. Market-1501 is close to saturated, and a model can score above 83% mAP on Market-1501 while collapsing to single-digit mAP on MSMT17 — a gap that shows up repeatedly in cross-dataset evaluations and is the honest picture of domain generalisation in this field.

Post-processing

Re-ranking with k-reciprocal encoding (Zhong et al., 2017) is worth knowing because it is nearly free and gives a large mAP gain. Two images are $k$-reciprocal neighbours if each appears in the other's top-$k$ list. The method expands the query with its reciprocal neighbours and combines Jaccard distance over neighbour sets with the original distance. Gains of 5 to 10 mAP points are typical.

The costs are real: it is transductive, needing the whole gallery at once, and its complexity is quadratic in gallery size. That rules it out for streaming trackers and makes it standard for offline benchmark evaluation, which is one reason published mAP figures can exceed what an online tracker achieves with the same embeddings.

The ethical record

DukeMTMC and DukeMTMC-reID were withdrawn by their creators in 2019. The recordings deviated from the university's institutional review board guidelines in two respects: they were made outdoors, and the data was released without protections. The dataset was also being used for surveillance research. A Duke professor publicly apologised.

The dataset continued to be used after withdrawal; a survey of the literature found over a hundred papers using it post-takedown. This is the field's own record, and it is the reason DukeMTMC results should not appear in new work.

Three things a researcher should carry forward:

  • Consent is not implied by public space. People filmed in a public square did not agree to become a benchmark.
  • Withdrawal must propagate. Citing a retracted dataset because a competitor did is how a harm persists.
  • Demographic performance gaps are documented and under-measured. Re-identification models trained on one population perform worse on others, and most papers report a single aggregate number.

Privacy in ML and bias in datasets cover the general framework. For this task specifically: retention limits, no cross-site identity linkage, and a documented purpose are the minimum, and there are deployments where the correct answer is not to build it.

References

  • Zheng et al., Scalable Person Re-identification: A Benchmark (Market-1501), ICCV 2015.
  • Hermans, Beyer and Leibe, In Defense of the Triplet Loss for Person Re-Identification, 2017 — arxiv.org/abs/1703.07737
  • Zhong et al., Re-ranking Person Re-identification with k-reciprocal Encoding, CVPR 2017 — arxiv.org/abs/1701.08398
  • Luo et al., Bag of Tricks and a Strong Baseline for Deep Person Re-identification, CVPR Workshops 2019 — arxiv.org/abs/1903.07071
  • Zhou et al., Omni-Scale Feature Learning for Person Re-Identification (OSNet), ICCV 2019 — arxiv.org/abs/1905.00953
  • Li, Sun and Li, CLIP-ReID, AAAI 2023 — arxiv.org/abs/2211.13977
  • Peng et al., Mitigating Dataset Harms Requires Stewardship: Lessons from 1000 Papers, NeurIPS 2021 — arxiv.org/abs/2108.02922

What to learn next