Video Understanding and Tracking

Video anomaly detection

You cannot collect examples of every unusual event, so these systems learn what normal looks like and score everything by how badly it fits.

On this page 7
  1. Why it exists
  2. How it works
  3. The three settings
  4. Where you have already seen this
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Video anomaly detection learns what a scene normally looks like, then flags moments that do not fit.

Think about the corridor outside your flat. You have walked it a thousand times and you never look closely.

One evening something is different. A stranger is standing still at the far end. You noticed instantly, and nobody ever taught you what "a stranger loitering" looks like.

You noticed because it did not match the corridor you know. That is the whole method.

Why it exists

The straightforward approach would be to collect examples of every bad event and train a classifier. That approach fails, for two reasons that are worth stating plainly.

You cannot collect them. Fires, fights, thefts, machine failures and falls are rare. You will not gather thousands of labelled examples of each.

You cannot list them. Even with unlimited data, tomorrow brings a kind of unusual event nobody wrote down. A classifier trained on a fixed list is blind to everything outside it.

So the task is inverted. Instead of learning what is bad, learn what is normal. Normal footage is free and endless.

How it works

   TRAINING          hours of ordinary footage, no labels needed
                                  │
                                  ▼
                     build a model of "usual"
                                  │
   RUNNING           new frame ──►│
                                  ▼
                     how well does it fit "usual"?
                                  │
                     ┌────────────┴────────────┐
                  fits well               fits badly
                     │                         │
                  ignore                  raise a score

The model never sees an anomaly during training. It only learns the shape of ordinary footage.

At runtime every frame gets a score for how poorly it fits. High score means unusual. Unusual is not the same as dangerous. That gap is the source of most of the pain in this field.

The three settings

Fully unsupervised: train only on normal footage. What the picture above shows.

Weakly supervised: you know which videos contain something unusual somewhere, but not when. Far more realistic. A person can label a video in seconds. Marking exact timings means watching all of it.

Fully supervised: exact start and end times for every anomaly. Rare, expensive, and usually not worth the labelling budget.

Where you have already seen this

  • Bank fraud alerts on a card transaction that does not match your habits.
  • Factory cameras spotting a machine behaving oddly before it breaks.
  • Traffic systems flagging a vehicle going the wrong way.
  • Hospital ward monitors detecting a patient fall.

What is honestly hard here

Unusual and important are different things, and the model has no way to tell them apart.

A delivery van parked where vans never park is unusual. A cat on the wall is unusual. A camera nudged by wind makes every frame unusual. None of them needs a human called.

This is why deployed systems drown in false alarms, and why the operators stop looking. Say that out loud before promising anyone a detector.

The second hard part is that a scene changes. Rain, festival lights, a new shop opening. Your model of normal was built last month and the world moved on. These systems need retraining on a schedule, not once.

Remember this

  • Learn normal from ordinary footage, then score how badly new footage fits.
  • Weakly supervised is the realistic setting: video-level labels, no timings.
  • Unusual does not mean important, and false alarms are the reason deployments fail.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy torch

Run against numpy 1.26.4 and torch 2.5.1 on CPU. Both examples use tiny synthetic feature streams so you can run them in seconds and see exactly what each metric rewards.

A complete normality model, in thirty lines

Real systems extract features from each frame with a 3D CNN or a video transformer. We skip that and work directly on 8-dimensional features, because everything interesting happens after feature extraction.

normality_model.py
import numpy as np

rng = np.random.default_rng(0)
D = 8                                    # 8 numbers describing each frame
SCALE = np.diag([3, 3, 1, 1, .3, .3, .3, .3])


def normal(n):
    """People walking: slow, similar-looking motion in a few directions."""
    return rng.normal(0, 1, (n, D)) @ SCALE


train = normal(600)                      # the training set contains NO anomalies at all
test = normal(200)
labels = np.zeros(200, dtype=int)

test[120:140] += np.array([0, 0, 0, 0, 0.85, 0.85, 0, 0])   # a cyclist: the anomaly
labels[120:140] = 1
test[40:50] += np.array([0, 0, 0, 0, 0.9, 0.9, 0, 0])       # someone running: still normal

mu = train.mean(0)
U, s, Vt = np.linalg.svd(train - mu, full_matrices=False)
basis = Vt[:4]                           # the 4 directions normal frames actually use


def anomaly_score(x):
    """How badly does the 'normal' subspace fail to explain this frame?"""
    coords = (x - mu) @ basis.T          # project onto the normal subspace
    rebuilt = coords @ basis + mu        # rebuild from that projection alone
    return np.linalg.norm(x - rebuilt, axis=1)


scores = anomaly_score(test)
print(f"mean score on normal frames   : {scores[labels == 0].mean():.3f}")
print(f"mean score on anomalous frames: {scores[labels == 1].mean():.3f}")


def roc_auc(y, s):
    """AUC is the chance a random anomaly outranks a random normal frame."""
    order = np.argsort(s, kind="mergesort")
    ranks = np.empty(len(s), float)
    ranks[order] = np.arange(1, len(s) + 1)
    for v in np.unique(s):               # average ranks within ties
        tie = s == v
        if tie.sum() > 1:
            ranks[tie] = ranks[tie].mean()
    n_pos, n_neg = y.sum(), (1 - y).sum()
    return (ranks[y == 1].sum() - n_pos * (n_pos + 1) / 2) / (n_pos * n_neg)


print(f"frame-level ROC AUC: {roc_auc(labels, scores):.3f}")

print("\npicking a threshold is a separate problem, and a harder one:")
for thr in (0.8, 1.0, 1.2, 1.5):
    pred = scores > thr
    caught = int((pred & (labels == 1)).sum())
    false_alarms = int((pred & (labels == 0)).sum())
    print(f"  threshold {thr:.1f} -> caught {caught:2d}/20 anomalous frames, "
          f"{false_alarms:2d} false alarms")
Output
mean score on normal frames   : 0.597
mean score on anomalous frames: 1.405
frame-level ROC AUC: 0.971

picking a threshold is a separate problem, and a harder one:
  threshold 0.8 -> caught 20/20 anomalous frames, 26 false alarms
  threshold 1.0 -> caught 19/20 anomalous frames, 12 false alarms
  threshold 1.2 -> caught 17/20 anomalous frames,  9 false alarms
  threshold 1.5 -> caught  6/20 anomalous frames,  4 false alarms

This output contains the entire field's problem

AUC is 0.971 and the system is not deployable. Read that sentence again, because it is the thing to take away from this page.

AUC measures ranking: how often an anomalous frame scores above a normal one. It never picks a threshold, so it never has to face the false alarm count. A published AUC of 0.97 sounds close to solved.

Now look at the threshold table. To catch all 20 anomalous frames you must accept 26 false alarms — more false alarms than true detections, on a test set that is 90% normal. To cut false alarms to 4, you catch 6 of 20.

The 10 "running" frames are not labelled as anomalies, but they score high, because running is genuinely unusual for this scene. The model is behaving correctly and the operator still gets called out for nothing.

The gap between a good AUC and a usable alarm is the deployment problem. Report both. A paper reporting only AUC has told you the model ranks well and nothing about whether anyone can use it.

Why implement AUC by hand? The rank formula makes clear what the metric is: mean rank of positives, corrected for the count of positives, divided by all positive-negative pairs. This implementation returns 0.971, identical to sklearn.metrics.roc_auc_score on the same inputs. Tie handling is the part people get wrong when they write their own — hence the averaging loop.

Weak supervision: learning from video-level labels

Nobody will mark exact anomaly timings across a thousand hours. They will tell you which videos contain something. Multiple-instance learning turns that into a training signal.

Treat each video as a bag of segments. A normal video's bag contains no anomalous segments. An anomalous video's bag contains at least one, and you are not told which.

mil_ranking.py
import torch

torch.manual_seed(0)
SEGMENTS, D = 32, 8          # each video is cut into 32 segments, 8 features each


def make_bag(anomalous):
    x = torch.randn(SEGMENTS, D) * 0.5
    if anomalous:                        # only 3 of the 32 segments actually contain it
        x[13:16, 4:6] += 2.0
    return x


scorer = torch.nn.Sequential(torch.nn.Linear(D, 16), torch.nn.ReLU(),
                             torch.nn.Linear(16, 1), torch.nn.Sigmoid())
opt = torch.optim.Adam(scorer.parameters(), lr=0.01)

for step in range(301):
    pos = scorer(make_bag(True)).squeeze(1)      # video labelled "contains an anomaly"
    neg = scorer(make_bag(False)).squeeze(1)     # video labelled "normal"
    # Only the highest-scoring segment of each bag is compared. That is the whole trick.
    ranking = torch.relu(1.0 - pos.max() + neg.max())
    smooth = ((pos[1:] - pos[:-1]) ** 2).sum()   # scores should not flicker frame to frame
    sparse = pos.sum()                           # anomalies are rare inside the video
    loss = ranking + 8e-4 * smooth + 8e-4 * sparse
    opt.zero_grad()
    loss.backward()
    opt.step()
    if step % 100 == 0:
        print(f"step {step:3d}  loss {loss.item():.3f}  "
              f"max score: anomalous bag {pos.max().item():.3f}, "
              f"normal bag {neg.max().item():.3f}")

with torch.no_grad():
    s = scorer(make_bag(True)).squeeze(1)
print("\nper-segment scores on a held-out anomalous video (truth: segments 13-15):")
print(" ".join(f"{v:.2f}" for v in s))
print("highest-scoring segment:", int(s.argmax()))
Output
step   0  loss 1.026  max score: anomalous bag 0.491, normal bag 0.505
step 100  loss 0.061  max score: anomalous bag 0.982, normal bag 0.038
step 200  loss 0.008  max score: anomalous bag 1.000, normal bag 0.005
step 300  loss 0.021  max score: anomalous bag 0.994, normal bag 0.012

per-segment scores on a held-out anomalous video (truth: segments 13-15):
0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.12 0.93 0.44 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
highest-scoring segment: 14

The model located segments 13 to 15 and was never told where to look. Every training signal it received was one bit per video: "this one has something". That is the appeal of multiple-instance learning, and this output is it working.

The three loss terms each do a specific job:

  • Ranking compares only the maximum of each bag. Every other segment is unconstrained, which is what lets an anomalous bag contain mostly normal segments.
  • Smoothness penalises the squared difference between neighbouring scores. Without it, scores flicker frame to frame and the output is unreadable.
  • Sparsity pushes the total score down, encoding "anomalies are rare inside a video". Without it the model can raise every score in an anomalous bag and still satisfy the ranking term.

Exact numbers depend on the PyTorch version's random number generator, so yours may differ slightly. The pattern — near-zero everywhere, a sharp peak at the planted segments — is robust across seeds.

Common mistakes

Putting anomalies in the training set of an unsupervised model. The model learns them as normal and stops flagging them. Audit your "normal only" footage; it usually is not.

Reporting AUC alone. See above. Report the false alarm rate at an operating point a human would accept.

Using accuracy. With 1% anomalies, predicting "normal" always scores 99%. Imbalanced data covers why, and this task is one of the most imbalanced there is.

Ignoring camera drift. A camera nudged by wind changes every feature at once and the whole recording scores as anomalous. Detect and reject global shifts before scoring.

Evaluating with a threshold tuned on the test set. Tune it on a validation split. This is a specific and common form of the leakage described in train-test split.

Assuming per-frame independence. Anomalies last for seconds. Smooth the score over a short window before thresholding, and count events rather than frames when reporting to an operator.

Try it yourself

Add a slow drift to the test features, a small constant creeping in over the 200 frames, as a camera slipping would produce. Watch the AUC. Then subtract a running mean from the features before scoring and watch it recover.

What to learn next

Researcher — Mathematics and papers.

Problem settings

One-class / unsupervised. Train on normal-only footage; score by how poorly a new frame fits. Benchmarks: UCSD Ped1/Ped2, CUHK Avenue, ShanghaiTech. Methods: frame reconstruction autoencoders, future-frame prediction (Liu et al., 2018), memory-augmented autoencoders such as MemAE (Gong et al., 2019).

The recurring failure of pure reconstruction models is that autoencoders generalise too well. A network trained on normal frames often reconstructs anomalous frames adequately, so the reconstruction error separates poorly. MemAE addresses this by forcing reconstruction to use a discrete memory of normal prototypes, so unusual input must be rebuilt from normal parts and the error grows.

Weakly supervised. Video-level labels only. Sultani et al. (CVPR 2018) introduced both the deep MIL ranking formulation and the UCF-Crime dataset: 1,900 untrimmed real surveillance videos, 128 hours, 13 anomaly types. This is now the dominant setting.

Open-set and multimodal. XD-Violence (Wu et al., ECCV 2020) adds audio: 4,754 untrimmed videos, over 217 hours, drawn from films and online video, evaluated by average precision rather than AUC.

The MIL ranking objective

For an anomalous bag $\mathcal{B}_a$ and a normal bag $\mathcal{B}_n$ with per-segment scores $f(\cdot)$:

$$ \ell = \max\left(0,\; 1 - \max_{i \in \mathcal{B}_a} f(v_i) + \max_{j \in \mathcal{B}_n} f(v_j)\right)

  • \lambda_1 \sum_{i}^{n-1} \left(f(v_i) - f(v_{i+1})\right)^2
  • \lambda_2 \sum_i^{n} f(v_i) $$

Where the first term is hinge ranking on bag maxima, the second is temporal smoothness across consecutive segments, and the third is sparsity. $\lambda_1, \lambda_2$ are typically around $8 \times 10^{-5}$ in the original work.

The known weakness is the $\max$ operator. Early in training the highest-scoring segment of an anomalous bag is chosen at random, so the gradient signal is noisy and the model can lock onto a wrong segment. Successors address this directly:

  • RTFM (Tian et al., ICCV 2021) ranks by feature magnitude over the top-$k$ segments instead of the single maximum, giving a denser and more stable signal.
  • MIST (Feng et al., CVPR 2021) generates pseudo-labels and refines the encoder in two stages.
  • VadCLIP and related work align segment features to text prompts using CLIP, adding open-vocabulary anomaly descriptions.

Metrics, and what they hide

Frame-level AUC is the field's default on UCF-Crime and ShanghaiTech, computed over all frames of all test videos concatenated.

Three problems with it, all worth knowing before reading a leaderboard:

  1. It never picks an operating point. A model with excellent AUC can still produce an unusable false alarm rate, as the developer section shows numerically.
  2. Frame-level concatenation lets long normal videos dominate. A model good at scoring long boring sequences low can inflate AUC without localising any anomaly well.
  3. It is insensitive to temporal fragmentation. A score that flickers on and off across an event scores the same as a clean contiguous detection, and is far worse operationally.

XD-Violence uses average precision instead, which is more sensitive in the low-false-positive regime and is the more informative metric for imbalanced detection generally. Reporting AP alongside AUC is becoming standard, and reporting a false alarm rate at a fixed recall is still rare and still the most useful number.

Recent reported results give a sense of the current ceiling: on UCF-Crime, methods around 2025 report frame-level AUC near 90%; on XD-Violence, average precision near 88 to 89%. Both benchmarks remain far from saturated, and the gap between them and deployment is larger than either number suggests.

The domain-shift problem

Every anomaly detector is trained on one camera's notion of normal. Move the camera, change the season, add a festival, and the distribution shifts. There is no consensus solution, and the practical mitigations are unglamorous:

  • Normalise features per camera and per time-of-day bucket.
  • Detect global shifts (camera moves, exposure changes) and suppress scoring during them.
  • Retrain or recalibrate on a fixed schedule, treating the model as perishable.
  • Track the alarm rate itself as a monitored metric — a step change in alarm rate usually means the scene changed, not that crime rose. See monitoring and drift.

The ethical position, stated plainly

This technology is surveillance infrastructure. UCF-Crime is real CCTV footage of real people who did not consent to being in a machine learning benchmark.

Three consequences a researcher should hold:

  • Base rates make false positives dominate. At a genuinely low anomaly rate, even a highly accurate detector produces mostly false alarms among its positives. When a false alarm means a person is stopped or questioned, that arithmetic is the ethical centre of the system.
  • "Anomalous" encodes whoever's normal was recorded. A detector trained on one neighbourhood's footage flags behaviour common in another. This is not a bug to be tuned away; it is what the method does.
  • Automation bias is real. Operators trust flagged clips more than unflagged ones, so errors propagate rather than being caught.

Responsible deployment and privacy in ML cover the framework. The specific recommendation here: treat the output as a ranking aid for a human reviewer, never as a trigger for automated action.

References

What to learn next