Evaluating Vision Models

Tracking metrics: MOTA and HOTA

Tracking has two jobs, finding things and keeping their identity straight, and MOTA, IDF1 and HOTA weigh those two jobs very differently.

On this page 10
  1. The short answer
  2. The analogy you have lived
  3. What a tracker outputs
  4. MOTA: the counting metric
  5. IDF1: the identity metric
  6. HOTA: the one that balances them
  7. Where you have seen this
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Tracking metrics ask two separate questions: did you find everyone, and did you keep their identities straight?

The analogy you have lived

At a family wedding, someone hands you a job. Keep an eye on four children running around the hall.

At the end of the evening you can fail in two very different ways.

You might have lost sight of one child for twenty minutes. That is a finding failure. You did not know where they were.

Or you might have watched all four the whole time, but mixed up the twins halfway through. Every child was in view. Your record of who was who is wrong. That is an identity failure.

Both are failures. They are not the same failure, and they do not have the same fix.

What a tracker outputs

A detector answers "what is in this frame". A tracker answers a longer question. What is in this frame, and which of them did I see before?

So its output has an extra column: an id number. That number should stay with one person while they are on screen.

   frame 1   [person id=1 at left ]  [person id=2 at right]
   frame 2   [person id=1 at left ]  [person id=2 at right]
   frame 3   [person id=1 at left ]  ... id=2 vanished ...
   frame 4   [person id=1 at left ]  [person id=3 at right]   <- same person,
                                                                 new number

That last line is an identity switch. The person never left. The tracker forgot who they were.

MOTA: the counting metric

MOTA takes the count of everything that went wrong and subtracts it from a perfect score.

Missed people, invented people and identity switches all get added up. That total is divided by how many were there. One number, easy to compute, easy to explain.

Its weakness is hidden in that sentence. An identity switch is counted once, at the moment it happens. Swap two people's ids halfway through a five-minute video and MOTA loses two points out of hundreds. The video is wrong from that moment onward, and the score barely notices.

MOTA is mostly a detection metric wearing a tracking metric's clothes.

IDF1: the identity metric

IDF1 goes the other way. It matches each tracker id to one real person, once, for the whole video. Then it asks: across all frames, how often did that pairing hold?

A swap halfway through means each id was right for only half the video. IDF1 halves. That is the honest penalty.

Its weakness is the mirror image of MOTA's. A tracker can find very few people and never confuse the ones it finds. IDF1 rates it well. It is useless.

HOTA: the one that balances them

HOTA was built because the field kept arguing about the two above.

It measures detection quality and association quality separately, then combines them so that neither can hide the other. A tracker has to be good at both to score well.

That is why HOTA is now the headline number on the main tracking benchmarks. It is also why HOTA is harder to explain — the price of not being fooled.

   HOTA  =  a balance of  [ did you find them ]  and  [ did you keep them straight ]
                              detection                    association

Where you have seen this

  • Shop cameras counting how many people entered, without counting one person five times.
  • Traffic systems measuring how long a car took to cross a junction.
  • Sports broadcasts drawing a trail behind one player.
  • Warehouse robots following the same box down a conveyor.

Notice what each needs. Counting entries needs identity. If one shopper becomes three ids, your footfall report is nonsense.

The honest part

There is no correct answer to how you weigh finding against remembering. It depends entirely on what the tracker feeds into.

A counter cares about identity above everything. A collision-avoidance system cares about detection above everything. Seeing an object with the wrong id beats not seeing it.

Report all three numbers. Then argue about which one your product should be judged on, with your actual product in the room.

Remember this

  • Tracking fails in two ways: not finding things, and mixing up who is who.
  • MOTA barely punishes identity switches. IDF1 punishes them heavily. HOTA balances the two.
  • Which one is right depends on what your tracker feeds into. Report all three.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

The demonstration below is deliberately built so the metrics disagree. That disagreement is the entire lesson.

Two broken trackers, three metrics

tracking_metrics.py
import numpy as np
from scipy.optimize import linear_sum_assignment

# Six frames of a corridor camera. Two people, A and B, walk past each other.
GT = {f: {"A": [10 + 8*f,  10,  30 + 8*f,  60],      # walking left to right
          "B": [70 - 8*f, 70,  90 - 8*f, 120]}      # walking right to left
      for f in range(1, 7)}

def box(f, who):
    """The tracker's box: the true box shifted one pixel, as a real one would be."""
    x1, y1, x2, y2 = GT[f][who]
    return [x1 + 1, y1 + 1, x2 + 1, y2 + 1]

# Tracker 1: never misses anyone, but swaps the two ids halfway through.
SWAPPER = {f: ({1: box(f, "A"), 2: box(f, "B")} if f <= 3
               else {2: box(f, "A"), 1: box(f, "B")}) for f in range(1, 7)}

# Tracker 2: keeps ids perfectly, but loses B after frame 2 and hallucinates
# a person standing in the corner from frame 5.
LOSER = {}
for f in range(1, 7):
    d = {1: box(f, "A")}
    if f <= 2:
        d[2] = box(f, "B")
    if f >= 5:
        d[9] = [150, 20, 170, 70]
    LOSER[f] = d


def iou(a, b):
    ix1, iy1 = max(a[0], b[0]), max(a[1], b[1])
    ix2, iy2 = min(a[2], b[2]), min(a[3], b[3])
    inter = max(0, ix2 - ix1) * max(0, iy2 - iy1)
    ua = (a[2]-a[0])*(a[3]-a[1]) + (b[2]-b[0])*(b[3]-b[1]) - inter
    return inter / ua


def match_frame(gt, tr, alpha):
    """Best pairing inside ONE frame, by IoU, above the threshold alpha."""
    gids, tids = list(gt), list(tr)
    if not gids or not tids:
        return []
    m = np.array([[iou(gt[g], tr[t]) for t in tids] for g in gids])
    rows, cols = linear_sum_assignment(-m)
    return [(gids[r], tids[c]) for r, c in zip(rows, cols) if m[r, c] >= alpha]


def mota(tr, alpha=0.5):
    tp = fp = fn = idsw = 0
    last = {}
    for f in sorted(GT):
        pairs = match_frame(GT[f], tr.get(f, {}), alpha)
        tp += len(pairs)
        fn += len(GT[f]) - len(pairs)
        fp += len(tr.get(f, {})) - len(pairs)
        for g, t in pairs:
            if g in last and last[g] != t:      # this person's id changed
                idsw += 1
            last[g] = t
    n_gt = sum(len(v) for v in GT.values())
    return dict(TP=tp, FP=fp, FN=fn, IDSW=idsw,
                MOTA=1 - (fn + fp + idsw) / n_gt)


def idf1(tr, alpha=0.5):
    """Pair each tracker id with one gt id ONCE, for the whole sequence."""
    gids = sorted({g for f in GT for g in GT[f]})
    tids = sorted({t for f in tr for t in tr[f]})
    ov = np.zeros((len(gids), len(tids)))
    for f in sorted(GT):
        for i, g in enumerate(gids):
            for j, t in enumerate(tids):
                if g in GT[f] and t in tr.get(f, {}) and iou(GT[f][g], tr[f][t]) >= alpha:
                    ov[i, j] += 1
    rows, cols = linear_sum_assignment(-ov)
    idtp = ov[rows, cols].sum()
    n_gt = sum(len(v) for v in GT.values())
    n_tr = sum(len(v) for v in tr.values())
    return 2 * idtp / (n_gt + n_tr)


def hota(tr, alpha):
    pairs = [(f, g, t) for f in sorted(GT)
             for g, t in match_frame(GT[f], tr.get(f, {}), alpha)]
    n_gt = sum(len(v) for v in GT.values())
    n_tr = sum(len(v) for v in tr.values())
    if not pairs:
        return 0.0, 0.0, 0.0
    det_a = len(pairs) / (n_gt + n_tr - len(pairs))     # TP / (TP + FN + FP)

    gt_count, tr_count = {}, {}
    for f in GT:
        for g in GT[f]:
            gt_count[g] = gt_count.get(g, 0) + 1
    for f in tr:
        for t in tr[f]:
            tr_count[t] = tr_count.get(t, 0) + 1

    scores = []
    for _, g, t in pairs:
        tpa = sum(1 for _, g2, t2 in pairs if g2 == g and t2 == t)
        fna = gt_count[g] - tpa      # frames of this person carrying another id
        fpa = tr_count[t] - tpa      # frames of this id spent on someone else
        scores.append(tpa / (tpa + fna + fpa))
    ass_a = float(np.mean(scores))
    return det_a, ass_a, float(np.sqrt(det_a * ass_a))


for name, tr in (("swapper", SWAPPER), ("loser", LOSER)):
    c = mota(tr)
    d, a, h = hota(tr, 0.5)
    sweep = np.mean([hota(tr, x)[2] for x in np.arange(0.05, 1.0, 0.05)])
    print(f"--- {name} ---")
    print(f"TP {c['TP']}  FP {c['FP']}  FN {c['FN']}  ID switches {c['IDSW']}")
    print(f"MOTA {c['MOTA']:.4f}   IDF1 {idf1(tr):.4f}")
    print(f"DetA {d:.4f}  AssA {a:.4f}  HOTA@0.5 {h:.4f}  HOTA(swept) {sweep:.4f}\n")
Output
--- swapper ---
TP 12  FP 0  FN 0  ID switches 2
MOTA 0.8333   IDF1 0.5000
DetA 1.0000  AssA 0.3333  HOTA@0.5 0.5774  HOTA(swept) 0.5166

--- loser ---
TP 8  FP 2  FN 4  ID switches 0
MOTA 0.5000   IDF1 0.7273
DetA 0.5714  AssA 0.8333  HOTA@0.5 0.6901  HOTA(swept) 0.6174

The whole point is in those six numbers

MOTA ranks the swapper above the loser: 0.8333 against 0.5000. The swapper found every single person in every single frame. Its only sin was two id switches, and MOTA charges two points out of twelve for them.

IDF1 ranks them the other way: 0.5000 against 0.7273. IDF1 pairs each id with one person for the whole clip. After the swap, id 1 matches person A for three frames out of six and person B for the other three. Whichever pairing you choose, half the clip is wrong.

HOTA agrees with IDF1 here, and shows its reasoning. The swapper has a perfect DetA of 1.0000 and a wrecked AssA of 0.3333. The loser has a poor DetA of 0.5714 and a healthy AssA of 0.8333. Neither component can hide behind the other, and the square root of their product is the score.

A one-pixel box offset costs HOTA about 0.06. Compare HOTA@0.5 with HOTA(swept): 0.5774 versus 0.5166 for the swapper. Every box in this demonstration is off by one pixel, which is invisible at IoU 0.5 and fatal above 0.9. HOTA sweeps the localisation threshold the way COCO sweeps IoU, so sloppy boxes cost you even when the identities are perfect.

Neither MOTA nor IDF1 sweeps anything. They are computed at IoU 0.5 and stay there, which means a tracker can loosen its boxes considerably without either metric reacting.

Which number should you look at

Ask what consumes the tracker's output.

A footfall counter cares only about identity. One shopper becoming three ids inflates your number by 200 percent. Optimise IDF1 or AssA.

A collision-avoidance system cares only about detection. It would far rather see an object with a scrambled id than not see it. Watch DetA and recall.

A dwell-time or queue-length measurement needs both, which is what HOTA was designed for.

Report all three. A paper that reports only MOTA is telling you about its detector.

Use the reference implementation

bash
pip install motmetrics       # CLEAR MOT, IDF1
# HOTA: the official code is github.com/JonathonLuiten/TrackEval

The forty lines above are a teaching implementation. Real evaluation has to handle crowd regions, distractor classes, sequences where a person leaves and re-enters, and the MOT-challenge file formats. TrackEval is the reference that benchmark numbers are produced with, and py-motmetrics includes a parity test against it.

Common mistakes

Reporting MOTA improvements from a better detector as tracking progress. Swap in a stronger detector and MOTA rises without a single line of the tracker changing. This is why the field moved to HOTA.

Evaluating on ground-truth detections. It measures the association logic alone. Useful as an ablation, dishonest as a headline.

Ignoring the identity of an object after it is occluded and returns. Benchmarks differ on whether a re-entry must keep its old id. MOT17 says yes. Some in-house evaluations quietly say no, which makes the tracker look much better than it is.

Comparing across detection thresholds. Lowering the tracker's confidence threshold raises recall, lowers precision, and moves MOTA and HOTA in opposite directions. Fix the threshold, or report the curve.

Try it yourself

Give the swapper a third id switch at frame 5 and watch MOTA lose one more point while HOTA's AssA falls much further. Then make the loser miss person B for only one frame instead of four, and see whether it overtakes the swapper on all three metrics or only some.

What to learn next

Researcher — Mathematics and papers.

CLEAR MOT

Bernardin and Stiefelhagen (2008) define, over all frames $t$:

$$ \mathrm{MOTA} = 1 - \frac{\sum_t \left(\mathrm{FN}_t + \mathrm{FP}_t + \mathrm{IDSW}_t\right)}{\sum_t g_t} $$

$g_t$ is the number of ground-truth objects in frame $t$. MOTA is unbounded below: a tracker emitting many false positives scores arbitrarily negative. Localisation quality is reported separately as

$$ \mathrm{MOTP} = \frac{\sum_{t,i} d_{t,i}}{\sum_t c_t} $$

with $d_{t,i}$ the overlap (or distance) of matched pair $i$ and $c_t$ the number of matches. MOTP is not combined into MOTA, so a tracker can win on MOTA with systematically poor boxes.

The structural criticism is arithmetic. $\mathrm{IDSW}$ is $O(\text{number of switches})$ while $\mathrm{FN}$ is $O(\text{number of frames} \times \text{objects})$. On long sequences the identity term is numerically negligible.

Identity metrics

Ristani et al. (2016) compute a single global bipartite matching between ground-truth trajectories and predicted trajectories, minimising total mismatched frames, then:

$$ \mathrm{IDF1} = \frac{2\,\mathrm{IDTP}}{2\,\mathrm{IDTP} + \mathrm{IDFP} + \mathrm{IDFN}} $$

Because the matching is global and per-trajectory, a single mid-sequence swap costs proportionally to the length of the affected segments. This is the intended behaviour and the reason IDF1 became the second standard number.

Its bias runs the other way: a tracker producing few, long, clean trajectories and missing many objects can score respectably, since unmatched ground truth contributes to IDFN once per frame but the surviving pairings dominate the ratio in short sequences.

HOTA

Luiten et al. (2021), HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking, IJCV. For a localisation threshold $\alpha$, let $\mathcal{M}_\alpha$ be the set of matched detection pairs (Hungarian, maximising association score). Then

$$ \mathrm{DetA}_\alpha = \frac{|\mathrm{TP}|}{|\mathrm{TP}| + |\mathrm{FN}| + |\mathrm{FP}|} $$

$$ \mathcal{A}(c) = \frac{|\mathrm{TPA}(c)|}{|\mathrm{TPA}(c)| + |\mathrm{FNA}(c)| + |\mathrm{FPA}(c)|}, \qquad \mathrm{AssA}\alpha = \frac{1}{|\mathrm{TP}|}\sum{c \in \mathrm{TP}} \mathcal{A}(c) $$

$\mathrm{TPA}(c)$ counts matched detections sharing both the ground-truth id and the predicted id of pair $c$; $\mathrm{FNA}(c)$ counts detections of that ground-truth id given a different predicted id or none; $\mathrm{FPA}(c)$ is the mirror image. Finally

$$ \mathrm{HOTA}\alpha = \sqrt{\mathrm{DetA}\alpha \cdot \mathrm{AssA}_\alpha}, \qquad \mathrm{HOTA} = \int_0^1 \mathrm{HOTA}\alpha \, d\alpha \approx \frac{1}{19}\sum{\alpha \in {0.05, \dots, 0.95}} \mathrm{HOTA}_\alpha $$

Three properties follow directly from that construction. The geometric mean forbids compensation: zero on either component is zero overall. Association is scored globally over trajectories while matching is done per detection, which is precisely the middle ground between MOTA's local view and IDF1's trajectory-level view. And the $\alpha$ sweep folds localisation quality into the headline number, where MOTP had been an ignored side channel.

Where the three disagree, systematically

The runnable example above is the canonical case, and it generalises:

FailureMOTAIDF1HOTA
Better detector, same trackerlarge gainmodest gainmoderate gain (DetA)
Mid-sequence id swap~2 points~50 percent of affected spanAssA collapses
Fragmentation without swapssmall penaltymoderate penaltyAssA penalty
Systematically loose boxesnone above IoU 0.5none above IoU 0.5penalised by the $\alpha$ sweep
High-confidence false positiveslarge penaltymoderateDetA penalty

Practical evaluation notes

Distractor handling. MOT17 marks certain ground-truth boxes as distractors (reflections, mannequins, static persons). Detections matching them must be removed before scoring, not counted as false positives. Implementations that skip this step report inflated FP counts.

The --USE_PARALLEL and threshold flags in TrackEval change results. Fix them in your scripts and record them alongside the numbers.

Sequence-level averaging. MOTChallenge computes metrics over concatenated sequences, not as a mean of per-sequence scores. The two differ substantially when sequence lengths vary, and mixing them is a common source of irreproducible numbers.

Single-object tracking uses a different family entirely — success plots, precision plots, area under the overlap curve (OTB, VOT, LaSOT). Do not carry MOT metrics across.

Papers

  • Bernardin and Stiefelhagen, Evaluating Multiple Object Tracking Performance: The CLEAR MOT Metrics, EURASIP JIVP 2008
  • Ristani et al., Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking, ECCVW 2016 — arxiv.org/abs/1609.01775
  • Luiten et al., HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking, IJCV 2021 — arxiv.org/abs/2009.07736
  • Dendorfer et al., MOT20: A Benchmark for Multi Object Tracking in Crowded Scenes, 2020 — arxiv.org/abs/2003.09003
  • Zhang et al., ByteTrack: Multi-Object Tracking by Associating Every Detection Box, ECCV 2022 — arxiv.org/abs/2110.06864

What to learn next