Evaluating an outlier detector
With anomalies at 1-in-100 or rarer, accuracy and even ROC-AUC flatter a detector — precision-based views and precision-at-k reflect what an alert queue actually delivers.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Judging an anomaly detector with ordinary accuracy is like praising a guard who never leaves his chair.
With thieves this rare, saying "all clear" forever scores 99.9%.
A guard watches a godown gate that sees one thief per thousand visitors. He waves everyone through, eyes closed. His accuracy: 99.9%. A rival guard catches half the thieves at the cost of bothering a few honest visitors — and scores lower on accuracy. The lazy guard wins the metric while failing the job.
Rare targets break ordinary scoring. Anomaly detection is nothing but rare targets.
Why this matters
Every serious detector outputs a suspicion score per point, and someone must act on the most suspicious ones. Real operations have a budget: a fraud analyst can investigate perhaps twenty alerts a day. The honest question is therefore not "how accurate is the detector?" but "of the twenty alerts I can afford to check, how many are real?"
That question has a name: precision at k — among the k highest-scoring points, the fraction that are true anomalies. It is the metric that matches how the tool gets used.
How it works
all 1,000 events, sorted by suspicion score
most suspicious -> [!] [!] [ ] [!] [ ] ... [ ] <- top 20: the day's alerts
^ ^
real false
anomaly alarm 8 real in top 20 = precision 0.40Two curve-based scores summarise the whole ranking, and they can disagree wildly. One family of scores gives the detector credit for correctly ignoring normal events. With 999 normal events per thief, ignoring normals is far too easy. The score then comes out flattering. The other family only gives credit for what the alerts themselves contain. On rare targets, always trust the second family, and read the disagreement itself as a warning.
A real example you have seen
UPI fraud teams live on exactly this arithmetic. Millions of daily transactions, a fixed analyst pool, and dashboards tracking "hit rate in reviewed alerts" — precision at k under its industry name. A model upgrade ships when the top of the queue gets denser with real fraud, whatever the accuracy number says.
Remember this
- With rare targets, accuracy is a lie — "always normal" scores near-perfect.
- Precision at k measures the alert queue someone actually works through.
- A score can flatter by crediting the easy job of ignoring normals. Check both views.
What to learn next
- ROC vs precision-recall curves — the full anatomy of the two disagreeing curves.
- Choosing a threshold from costs — from ranked scores to a defensible alarm line.
- Imbalanced data — the training-side toolkit for rare classes.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU; last-digit drift elsewhere is normal.
Three verdicts on one detector
An isolation forest hunts 20 planted frauds among 980 normal events — 2% contamination, and this evaluation needs the labels only at scoring time:
import numpy as np
from sklearn.ensemble import IsolationForest
from sklearn.metrics import average_precision_score, roc_auc_score
rng = np.random.default_rng(5)
normal = rng.normal(0, 1, size=(980, 2))
fraud = rng.uniform(1.2, 3.0, size=(20, 2)) * rng.choice([-1, 1], size=(20, 2))
X = np.vstack([normal, fraud])
y = np.array([0] * 980 + [1] * 20) # 2% positives, like real fraud data
scores = -IsolationForest(random_state=0).fit(X).score_samples(X)
print("ROC-AUC:", round(roc_auc_score(y, scores), 3))
print("PR-AUC: ", round(average_precision_score(y, scores), 3))
k = 20 # alerts one analyst can check per day
top = np.argsort(scores)[-k:]
print(f"precision in the top {k} alerts:", round(y[top].mean(), 2))ROC-AUC: 0.961 PR-AUC: 0.459 precision in the top 20 alerts: 0.4
The walkthrough
Three numbers, three stories. ROC-AUC of 0.961 sounds like an A grade. PR-AUC of 0.459 sounds mediocre. Precision-at-20 of 0.4 is the operational truth: an analyst working this queue finds 8 real frauds and wastes 12 investigations. All three describe the same ranking.
Why ROC-AUC flatters. It asks how often a random fraud outscores a random normal event — and gets credit for all 980 normals ranked low. Rank 500 normal events above a fraud and ROC-AUC barely moves; the alert queue, meanwhile, is garbage. PR-AUC and precision-at-k feel that pain immediately, because they look only at what the top of the ranking contains. The full anatomy of the two curves gets its own lesson.
Baselines differ, too. Random guessing scores 0.5 on ROC-AUC but only the positive rate — 0.02 here — on PR-AUC. So 0.459 is twenty-three times better than chance, while 0.961 is barely better than double it. Judge each number against its own floor.
The negation on score_samples converts the library's higher-is-normal convention into higher-is-suspicious, so all three metrics read the ranking in the same direction. Silent sign errors here produce impressively terrible dashboards.
Common mistakes
Evaluating on data the detector trained on. Unsupervised is not exempt from train/test discipline: the forest above scored its own training data, fine for a demo, flattering in production. Score held-out events.
Trusting a handful of labelled anomalies too much. With 20 positives, PR-AUC has huge variance — one lucky rank swap moves it visibly. Report uncertainty (bootstrap the labelled set) before declaring a winner between two detectors.
Tuning the contamination parameter against the test labels. The moment labels steer hyperparameters, the "unsupervised" evaluation is quietly supervised, and production performance will disappoint. Keep a untouched final test set — the same leakage discipline as always.
Labelling only the alerts your old system produced. Historical labels come from what past detectors flagged; frauds nobody caught are labelled "normal". New detectors get punished precisely where they improve. Where possible, label a random sample too, and read results with this bias in mind.
Try it yourself
Compute precision at k for k = 5, 10, 50, 100 and plot the four values. Then rerun the whole script with rng = np.random.default_rng(6) and watch how much the PR numbers wobble with only 20 positives — that wobble is the error bar your report needs.
What to learn next
- ROC vs precision-recall curves — the full anatomy of the two disagreeing curves.
- Choosing a threshold from costs — from ranked scores to a defensible alarm line.
- Imbalanced data — the training-side toolkit for rare classes.
Researcher — Mathematics and papers.
Metric definitions and their behaviour under rarity
For scored data with positive prevalence $\pi$, ROC-AUC equals $P(s^+ > s^-)$ for random positive and negative scores — the Mann–Whitney statistic — and is prevalence-invariant. Average precision is:
$$ AP = \sum_k \left( R(k) - R(k-1) \right) P(k) $$
Where:
- $P(k), R(k)$ — precision and recall at rank cut-off $k$.
- $AP$ — the area under the precision-recall curve by step-wise integration; its random baseline is $\pi$.
Prevalence-invariance is ROC's flaw here, not its virtue: as $\pi \to 0$, false positives per true positive scale as $(1-\pi)\,\mathrm{FPR} / (\pi\,\mathrm{TPR})$, so operationally catastrophic rankings retain high ROC-AUC. Davis and Goadrich (2006) prove the curves are linked (one dominates in ROC space iff it dominates in PR space) yet AUC orderings between models can disagree — comparing two detectors on ROC-AUC and deploying on precision is how teams ship regressions.
Precision-at-k connects to both: $P@k = \mathbb{E}[y \mid \text{rank} \le k]$, and under a budget interpretation is the value delivered per investigation. Its estimator has binomial variance $P(1-P)/k$ — wide at operational k, hence the bootstrap advice.
Evaluation without labels
Fully unsupervised assessment is an active subfield:
- Mass-volume and excess-mass curves (Clémençon and Jakubowicz, 2013; Goix, 2016): score a detector by how small a volume its high-score regions occupy at given data mass — density-level-set quality without labels. Requires volume estimates, so dimension bites.
- EMMV (Goix, 2016) operationalises both criteria for model selection among unsupervised detectors.
- Internal consistency approaches: agreement between detectors (Marques et al., 2020, IREOS) as a proxy for quality — with the circularity caveat that implies.
None replaces even a small labelled audit sample; they triage which configurations deserve one.
Benchmark hygiene
Campos et al. (2016) dissected the classic LOF-family comparisons and found conclusions driven by dataset preparation choices — downsampled classes as "anomalies", duplicate handling, normalisation — as much as by algorithms. ADBench (Han et al., 2022; 57 datasets, 30 algorithms) standardises this and reports three sobering regularities: no unsupervised method dominates, a little supervision reorders everything, and preprocessing rivals algorithm choice. Emmott et al. (2013) remains the reference for constructing honest synthetic benchmarks from real classification data.
For deployment, the natural end-point of this lesson is decision-theoretic: convert scores plus cost estimates into thresholds via cost-based threshold choice, and report expected cost — the metric the business already believes in.
What to learn next
- ROC vs precision-recall curves — the full anatomy of the two disagreeing curves.
- Choosing a threshold from costs — from ranked scores to a defensible alarm line.
- Imbalanced data — the training-side toolkit for rare classes.