Outlier and Anomaly Detection

Anomaly detection by reconstruction error

Train an autoencoder to compress and rebuild normal data, and it rebuilds normal points well but anomalies badly — the size of the rebuilding mistake becomes the anomaly score.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Teach a network to compress and redraw only normal data, and it becomes bad at redrawing anything else.

How badly it redraws a point is that point's anomaly score.

An artist has sketched cows her whole life. Show her a cow for two seconds and she redraws it almost perfectly — she knows the rules of cows, so a glance is enough. Show her a camel for two seconds and her redrawing comes out lumpy and wrong. Her skill is specialised, and the size of her mistake reveals what she has never seen.

Reconstruction-based anomaly detection trains exactly such a specialist.

Why it exists

Fences and scores from earlier lessons work on rows of a few dozen numbers. But what is "distance" between two ECG heartbeats, or two photos of a weld? Anomalies in rich, high-dimensional data live in the relationships between many values — patterns, not positions.

An autoencoder is a network squeezed in the middle: it must compress its input down to a few numbers, then rebuild the original from that summary. The squeeze forbids memorising. To rebuild well through a tiny middle, the network is forced to learn the rules of the data. Which patterns go together. What follows from what.

Train it purely on normal examples and those rules are normal's rules. A rule-breaking input gets rebuilt according to rules it does not follow — badly.

How it works

input            squeeze         rebuilt          verdict

normal   ->  [tiny summary]  ->  nearly identical  ->  small error: ok
anomaly  ->  [tiny summary]  ->  visibly wrong     ->  BIG error: flag!

The final score is one number: how far the rebuilt version is from the original. Score every incoming point, flag the worst rebuilders. No anomaly examples were needed at any stage — this is the novelty mode in deep-learning clothes.

A real example you have seen

Factory camera inspection. Thousands of photos of good products exist; defects are rare and endlessly varied. Systems train a rebuilder on good photos only. They then flag any item the rebuilder redraws poorly. Scratches, dents and misprints it has never seen are all caught by the same mechanism.

Remember this

  • The autoencoder learns to compress and rebuild normal data through a squeeze.
  • Anomalies break the learned rules, so they rebuild badly.
  • The rebuilding mistake is the anomaly score. No anomaly examples needed.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Outputs verified with torch 2.5.1 on CPU. Neural training is only bit-exact on matching versions and hardware — your last decimals may differ, and the gap between the two final numbers is the part that must reproduce.

A two-sensor machine with one hidden rule

Flow and pressure move together — pressure sits near twice the flow. The autoencoder must discover that rule to survive its one-number squeeze:

reconstruction_anomaly.py
import torch
from torch import nn

torch.manual_seed(0)
flow = torch.rand(400, 1) * 10
pressure = 2 * flow + 0.3 * torch.randn(400, 1)   # the hidden rule: pressure = 2 x flow
normal = torch.hstack([flow, pressure])

model = nn.Sequential(nn.Linear(2, 1), nn.Linear(1, 2))   # bottleneck of width 1
opt = torch.optim.Adam(model.parameters(), lr=0.05)
for epoch in range(300):
    opt.zero_grad()
    loss = ((model(normal) - normal) ** 2).mean()
    loss.backward()
    opt.step()
print(f"final training loss: {loss.item():.4f}")

def reconstruction_error(reading):
    x = torch.tensor([reading])
    return ((model(x) - x) ** 2).mean().item()

print(f"obeys the rule  [4.0, 8.1]: {reconstruction_error([4.0, 8.1]):.4f}")
print(f"breaks the rule [4.0, 1.0]: {reconstruction_error([4.0, 1.0]):.4f}")
Output
final training loss: 0.0106
obeys the rule  [4.0, 8.1]: 0.0008
breaks the rule [4.0, 1.0]: 5.1765

The walkthrough

Both test readings have individually normal values. Flow 4.0 is mid-range; pressure 1.0 occurs in the data (at low flow). Per-column rules see nothing. The pair (4.0, 1.0) violates the pressure-is-double-flow rule — and its reconstruction error is 6,000 times the obedient reading's. Correlation-breakers are precisely where reconstruction beats per-column rules and axis-cut methods like isolation forest.

Why it works, mechanically. The width-1 bottleneck forces two numbers into one. The cheapest accurate summary of this data is "the flow value", with pressure rebuilt as roughly double it. The network finds that compression during training. Handed (4.0, 1.0), it encodes something flow-like and decodes a pressure near 8 — far from the actual 1.0. The error is the rule violation.

The bottleneck size is the capacity dial. Width 2 here would let the network learn the identity function and rebuild everything perfectly — anomalies included, scores all zero. Real projects tune the squeeze so normal data reconstructs well and held-out checks stay sane. This lesson's autoencoder background covers the architecture spectrum.

Thresholding comes after scoring. Standard practice: compute errors on held-out normal data, set the alarm at a high percentile of those errors (say the 99th), and monitor the false-alarm rate it implies.

Common mistakes

Training on contaminated data. Anomalies in the training set teach the rebuilder their patterns, and those anomalies go quiet forever after. The clean-data requirement is stricter here than for shallow methods, because the network has capacity to learn contamination well.

Making the model too powerful. More width, more depth, longer training — reconstruction of everything improves, including anomalies, and score separation shrinks. Here, worse modelling is better detecting. Genuinely counterintuitive; watch the gap, not the loss.

Scoring in raw units across mixed-scale features. A feature living in thousands dominates the squared error. Standardise features before training, or the score measures one column.

Comparing error values across retrained models. Like LOF scores, reconstruction errors are relative to one trained artefact. Recalibrate the threshold after every retraining.

Try it yourself

Change the bottleneck to nn.Linear(2, 2) then nn.Linear(2, 2) — no squeeze — retrain, and rescore both readings. Then restore the squeeze but train for only 20 epochs. One experiment kills the detector by capacity, the other by undertraining; check which damages the gap more.

What to learn next

Researcher — Mathematics and papers.

The objective and its geometry

Train encoder $f_\theta: \mathbb{R}^d \to \mathbb{R}^m$ and decoder $g_\phi: \mathbb{R}^m \to \mathbb{R}^d$ with $m < d$ on normal data:

$$ \min_{\theta, \phi} \; \mathbb{E}{x \sim P{\text{normal}}} \left[ \lVert x - g_\phi(f_\theta(x)) \rVert^2 \right] $$

Where:

  • $m$ — bottleneck width, the capacity constraint.
  • $g \circ f$ — the reconstruction map; its image is an $m$-dimensional manifold in $\mathbb{R}^d$.
  • The anomaly score is $s(x) = \lVert x - g_\phi(f_\theta(x)) \rVert^2$.

Geometrically, training fits a manifold to the normal data's support, and $s(x)$ approximates squared distance to that manifold. With linear layers and squared loss, the optimum spans the top-$m$ principal subspace (Baldi and Hornik, 1989) — the developer example is PCA-to-one-component in costume, and PCA reconstruction error is the correct shallow baseline to report.

Failure modes with names

  • Anomaly generalisation: autoencoders reconstruct some unseen anomalies well, especially those lying near the learned manifold (e.g. simpler than normal data). Gong et al. (2019, MemAE) counter with memory modules restricting the decoder to prototype combinations.
  • The likelihood trap: deep generative models can assign higher likelihood to out-of-distribution data than in-distribution (Nalisnick et al., 2019) — a caution against assuming VAE likelihoods fix reconstruction's flaws. Typicality tests and likelihood-ratio methods (Ren et al., 2019) are partial repairs.
  • Score saturation under capacity: as $m \to d$ or training lengthens, $s(x) \to 0$ everywhere. Regularised variants constrain the map: contractive penalties, denoising objectives (Vincent et al., 2008), or sparsity.

Variants in production use

  • Denoising AE: corrupt inputs, reconstruct originals — the score becomes robustness to perturbation, often more stable on tabular data.
  • VAE (Kingma and Welling, 2014): probabilistic encoder plus reconstruction likelihood; score by reconstruction probability (An and Cho, 2015).
  • AnoGAN and successors (Schlegl et al., 2017): GAN-based reconstruction for imaging; f-AnoGAN made it fast enough for clinics.
  • Deep SVDD (Ruff et al., 2018) drops reconstruction entirely — contract normal representations around a centre — bridging to the one-class SVM lineage.
  • For time-indexed signals, sequence autoencoders with LSTM or attention backbones score reconstruction per window — the deep end of time-series anomaly detection.

Ruff et al. (2021), A unifying review of deep and shallow anomaly detection, ties these threads together; ADBench (Han et al., 2022) tempers enthusiasm — on plain tabular data, shallow detectors match deep ones far more often than the paper count suggests. Reconstruction methods earn their complexity on images, sequences and signals.

What to learn next