Outlier and Anomaly Detection

Outliers vs novelties

Outlier detection hunts the strangers already inside your data; novelty detection learns from clean data and judges each new arrival — the same tools behave differently in each mode.

Read these first

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Outlier detection asks "who in this data does not belong?"

Novelty detection asks "does this new arrival match the data I trust?" Same instinct, different timing, different setup.

Picture a hostel warden. One evening she scans the mess hall and spots a face that matches no resident — a stranger already inside. That is outlier detection: examine the crowd you have, find who sticks out.

The same warden also stands at the gate each morning. She knows every resident's face from months of duty, and she judges each new person as they arrive. That is novelty detection: learn what normal looks like from clean examples, then screen newcomers one by one.

Why this matters

The two jobs sound interchangeable and are not. In outlier detection, your training data is already contaminated — the strangers are mixed in, and the method must resist being fooled by them. In novelty detection, your training data is trusted clean, and the method's whole job is to draw a boundary around it for future arrivals.

Mixing the modes up produces real failures. Train a "normal behaviour" model on data that quietly contains fraud, and the model learns fraud as normal. Or run a contaminated-data method on a live stream one event at a time, and it has no crowd to compare against.

How it works

outlier detection             novelty detection

 [x x x x X x x]              clean: [x x x x x] -> learn a fence
        ^
   find the X                 later:  o  -> inside fence?  ok
   inside the pile                     O  -> outside fence? alarm!

Ask two questions about any anomaly task before choosing tools. First: do I already have the suspicious rows in hand, or do they arrive later? Second: can I trust some portion of my data as clean? The answers place you in one mode or the other.

There is a third character in this story: the anomaly itself, meaning any point that deviates from expected behaviour. Outliers and novelties are both anomalies — the words describe when and where you meet them, not different kinds of strangeness.

A real example you have seen

Your bank does both. Cleaning historical transaction data for analysis — removing corrupted rows and one-off errors — is outlier detection over data at rest. The instant SMS asking "was this you?" seconds after an odd payment is novelty detection on a live stream, screened against your learned normal pattern.

Remember this

  • Outlier mode: contaminated data in hand; find the strangers inside it.
  • Novelty mode: trusted clean data; fence it, then screen newcomers.
  • Decide your mode before choosing a tool — most tools behave differently in each.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

One tool, two modes

scikit-learn encodes the distinction directly: LocalOutlierFactor has a novelty switch that changes what the object is for.

two_modes.py
import numpy as np
from sklearn.neighbors import LocalOutlierFactor

rng = np.random.default_rng(0)
clean = rng.normal(0, 1, size=(100, 2))            # a normal week of sensor readings
dirty = np.vstack([clean, [[6.0, 6.0]]])           # one bad reading already inside

# Outlier detection: judge the data you already have.
flags = LocalOutlierFactor().fit_predict(dirty)
print("flagged inside existing data:", int((flags == -1).sum()))

# Novelty detection: learn from clean data, judge NEW points later.
detector = LocalOutlierFactor(novelty=True).fit(clean)
print("new point [0, 0]:", detector.predict([[0.0, 0.0]])[0])
print("new point [6, 6]:", detector.predict([[6.0, 6.0]])[0])
Output
flagged inside existing data: 5
new point [0, 0]: 1
new point [6, 6]: -1

The walkthrough

The convention throughout scikit-learn: -1 means anomaly, 1 means normal. Not 0 and 1. Every anomaly estimator in the library speaks this dialect.

Outlier mode has no predict for new data — on purpose. With novelty=False (the default), the object only offers fit_predict on the data it was given. Call .predict() on it and you get an AttributeError. The API physically prevents you from misusing an already-contaminated fit as a gate for newcomers.

Why did 5 points get flagged when we planted 1? Every detector needs an assumed contamination rate — the fraction of the data presumed anomalous. LOF's default (contamination="auto") uses a fixed score cut-off, and four ordinary points from the cloud's thin edge fell past it. Pass contamination=0.01 and only the sharpest 1% get flagged. This rate is the single most consequential setting in the whole field.

In novelty mode, the model never saw [6, 6] during fit, and correctly rejects it at prediction time. It equally correctly admits [0, 0]. The fence was learned from clean alone.

Common mistakes

Training a novelty detector on unverified data. If strangers hide in the "clean" set, the fence is drawn around them too. Fraud that resembles past undetected fraud sails through. Verify your normal data by domain checks before trusting the mode.

Scoring a novelty detector on its own training data. It will flag roughly contamination percent of them regardless of truth — that is a property of the fitting, not evidence about accuracy. Evaluation needs held-out data, ideally with a few known anomalies: see evaluating an outlier detector.

Assuming anomalies are rare in every application. Defect rates in a struggling factory line can hit 20%. Methods and thresholds assume rarity; check the assumption against reality first.

Forgetting time. In streams, "normal" drifts — festival-season spending is not January spending. A fence learned once and never refreshed slowly turns every customer into an anomaly. For time-indexed data, start with anomaly detection for time series.

Try it yourself

Fit the novelty detector on dirty instead of clean — poisoned training on purpose — and re-test both new points. Watch what happens to the verdict on [6, 6], then explain it in one sentence.

What to learn next

Researcher — Mathematics and papers.

The standard taxonomy

Chandola, Banerjee and Kumar (2009), Anomaly detection: a survey, remains the field's organising reference. Their axes:

  • Supervision: supervised (labelled anomalies — effectively imbalanced classification); semi-supervised (labelled normal only — novelty detection); unsupervised (no labels — outlier detection under a rarity assumption).
  • Anomaly type: point anomalies; contextual anomalies (normal value, wrong context — 30°C in Delhi in January); collective anomalies (individually normal points, anomalous as a group — a slow port scan).
  • Output: binary flags versus ranked scores. Scores dominate in practice, since thresholds encode operational budgets: see the cost-based threshold lesson.

Formal problem statements

Novelty detection is density-level-set estimation: given samples from clean distribution $P$, estimate a region $C_\alpha$ with $P(C_\alpha) \ge 1 - \alpha$ and minimal volume, flagging points outside it. The minimum-volume-set view (Scott and Nowak, 2006) makes the trade-off explicit: $\alpha$ is the tolerated false-alarm rate on normal data.

Where:

  • $P$ — the distribution of normal behaviour, known only through samples.
  • $\alpha$ — the fraction of genuine normal points you accept flagging.
  • $C_\alpha$ — the learned "fence"; its complement is the alarm region.

Unsupervised outlier detection replaces the clean-sample assumption with a mixture $(1-\epsilon) P + \epsilon Q$, contamination rate $\epsilon$ and anomaly distribution $Q$ both unknown. Identifiability requires assumptions — typically rarity ($\epsilon$ small) plus separation ($Q$ places mass where $P$ does not). Robust statistics contributes the breakdown-point lens: an estimator useful here must tolerate an $\epsilon$ fraction of arbitrary points, previewing the robust rules lesson.

Connections worth knowing

  • PU learning (positive-unlabelled): novelty detection with an extra unlabelled pool; Elkan and Noto (2008) give the standard reduction.
  • One-class classification: the discriminative approach to the semi-supervised setting — one-class SVM and Deep SVDD (Ruff et al., 2018).
  • Concept drift: in streams, the fence must adapt; Gama et al. (2014) survey drift-aware evaluation, which reframes "novelty" as "not yet normal".
  • Ruff et al. (2021), A unifying review of deep and shallow anomaly detection, connects the classical and deep families under the density-level-set view — the best single modern read.

Empirically, the mode distinction dominates method choice: Han et al. (2022, ADBench) find that access to even a small verified-clean training set changes method rankings more than any algorithmic difference.

What to learn next

What to learn next

These follow on from what you just read.

  • Outlier and Anomaly Detection

    Z-scores, IQR fences and MAD

    The classic z-score rule can be blinded by the very outliers it hunts, because outliers drag the mean and inflate the spread — median-based rules stay honest.

  • Outlier and Anomaly Detection

    Isolation forest

    Isolation forest flags the points that random yes/no splits separate from the crowd in only a few cuts — anomalies are easy to isolate, and that ease is the score.

  • Outlier and Anomaly Detection

    Local outlier factor

    LOF compares each point's crowding against its neighbours' crowding, so it catches points that are odd for their own neighbourhood even when they look globally ordinary.