Isolation forest
Isolation forest flags the points that random yes/no splits separate from the crowd in only a few cuts — anomalies are easy to isolate, and that ease is the score.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An isolation forest asks how few random questions it takes to single a point out.
Strangers get singled out fast, and that speed is the anomaly score.
Play "guess who" with a crowded wedding photo. To pin down one particular aunty among two hundred guests takes many questions — left half? wearing red? front row? But the one guest standing alone by the gate takes a single question. Odd points are easy to describe, precisely because nothing else is like them.
Isolation forest turns that ease into a number.
Why this matters
Older rules ask "how far is this point from typical?" That needs a definition of typical and a scale for far. It also breaks in ugly ways when outliers corrupt those definitions. And distance-based methods slow to a crawl on big data, comparing everything with everything.
Isolation forest sidesteps the whole framing. It never models "normal" at all. It splits the data randomly, again and again, and watches which points fall out alone almost immediately. No distances, no distribution assumptions, and fast enough for millions of rows.
How it works
Build a tree of random cuts: pick a random column, cut at a random value, repeat on each half. Track how many cuts it takes before each point sits alone.
all 302 orders
/ \
value < 2100? value >= 2100
/ \ \
... many ALONE after 1 cut
more cuts <- the ₹4,000 order
needed to
isolate any
normal orderOne random tree proves little, so build a few hundred on different random samples. That is a forest, like the random forests used for prediction. The difference: these trees are grown blind, with no labels to guide the cuts. Average each point's isolation depth across all trees. Consistently-shallow points are the anomalies.
The averaging is what makes it reliable. Any single tree's cuts are luck. But a point that every random tree isolates in two or three cuts is genuinely sitting apart from the crowd.
A real example you have seen
Payment gateways screen millions of daily transactions. A ₹2 lakh order at 3 a.m. from a fresh device separates from the pile in a couple of random cuts on any tree. Isolation-style scoring is popular in exactly these high-volume screens because it costs so little per transaction.
Remember this
- Anomalies are easy to isolate; normal points hide deep inside the crowd.
- The score is the average number of random cuts to single a point out.
- No distances, no "typical" to corrupt — and it scales to millions of rows.
What to learn next
- Local outlier factor — the local-density view isolation trees lack.
- Evaluating an outlier detector — deciding whether these scores are any good.
- Random forest — the supervised cousin, if trees themselves are new.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU; expect last-digit drift elsewhere.
Scoring orders, and taming the flag count
import numpy as np
from sklearn.ensemble import IsolationForest
rng = np.random.default_rng(42)
orders = rng.normal([500, 3], [120, 1], size=(300, 2)) # order value, delivery days
X = np.vstack([orders, [[510, 3.2]], [[4000, 14.0]]]) # one typical, one strange
iso = IsolationForest(random_state=0).fit(X)
scores = iso.score_samples(X)
print("typical order score:", round(scores[-2], 3))
print("strange order score:", round(scores[-1], 3))
print("flagged by default settings:", int((iso.predict(X) == -1).sum()), "of", len(X))
# Tell it how dirty you believe the data is, and the flags become sane.
iso1 = IsolationForest(contamination=0.01, random_state=0).fit(X)
print("flagged with contamination=0.01:", int((iso1.predict(X) == -1).sum()))typical order score: -0.381 strange order score: -0.892 flagged by default settings: 37 of 302 flagged with contamination=0.01: 4
The walkthrough
score_samples is the real output; predict is a threshold on it. Scores are negated so that lower means more anomalous: the planted strange order sits at -0.892, deep below the typical order's -0.381. Rank by score and you have a triage list — usually more useful than binary flags.
37 of 302 flagged by default. The default contamination="auto" applies a fixed offset from the original paper, and on tight data it flags generously. It has no idea only two rows in this table deserve attention.
contamination=0.01 moves the threshold, not the scores. Refit with a stated belief — "about 1% of my data is bad" — and the flags drop to 4. Choosing this number honestly is half the job in production; the evaluation lesson shows what to do when nobody knows it.
Two more knobs worth knowing. n_estimators (default 100) buys score stability; max_samples (default 256) is the per-tree sample size, and the paper's insight is that small samples help — they keep normal points crowded, so anomalies stand out more.
Common mistakes
Reading the sign backwards. In score_samples, more negative means more anomalous. The related decision_function shifts scores so negatives are flags. Mixing the two silently inverts your triage list.
Expecting isolation to catch inliers-with-wrong-combinations. A point inside the overall cloud but violating a correlation — normal value, normal day count, impossible pair — isolates slowly on axis-aligned cuts. Reconstruction-based detection handles correlation-breakers better.
Ignoring categorical columns. Random numeric cuts need numbers. One-hot columns work poorly (a cut on a 0/1 column isolates whole categories, not odd rows). Prefer numeric encodings with real order, or detectors designed for mixed data.
Refitting on every new batch and comparing scores across fits. Scores are relative to the fitted forest. Fit once on reference data, then score new arrivals with the same object — the novelty-style workflow from the first lesson.
Try it yourself
Add a correlation-breaking order at [200, 12] — cheap but two-week delivery, both values individually common. Check its score against the two planted points. Then rerun with max_samples=16 and watch what small samples do to score separation.
What to learn next
- Local outlier factor — the local-density view isolation trees lack.
- Evaluating an outlier detector — deciding whether these scores are any good.
- Random forest — the supervised cousin, if trees themselves are new.
Researcher — Mathematics and papers.
The scoring formula
Liu, Ting and Zhou (2008), Isolation forest, define the anomaly score for point $x$ over trees grown on subsamples of size $\psi$:
$$ s(x, \psi) = 2^{-\, \mathbb{E}[h(x)] / c(\psi)} $$
Where:
- $h(x)$ — the path length: number of splits before $x$ is isolated in one tree, plus an adjustment for unresolved leaves.
- $\mathbb{E}[h(x)]$ — the average path length across the forest.
- $c(\psi) = 2H(\psi - 1) - 2(\psi - 1)/\psi$ — the expected path length in a binary search tree of $\psi$ points, with $H$ the harmonic number; it normalises depth so scores compare across subsample sizes.
$s \to 1$ signals a clear anomaly (paths far shorter than random expectation), $s \approx 0.5$ signals an ordinary point. scikit-learn reports score_samples $= -s$, hence "lower is stranger", and decision_function $= -s + \text{offset}$ with the offset set by contamination.
Why subsampling helps
Two failure modes of density-style detectors — swamping (normal points near anomalies get flagged) and masking (anomaly clusters look dense) — are reduced by small $\psi$: subsampling thins anomaly clusters apart and separates them from normal mass. The paper fixes $\psi = 256$ and $t = 100$ trees as defaults, with training cost $O(t\,\psi \log \psi)$ and scoring cost $O(t \log \psi)$ per point — independent of $n$, the property that makes the method viable at stream scale.
Known weaknesses and refinements
- Axis-parallel artefacts: random single-feature cuts produce rectangular score bands ("ghost" regions of spuriously low scores along axes through sparse areas). Extended Isolation Forest (Hariri, Kind, Brunner, 2019) cuts with random hyperplanes, removing the artefacts at slight cost.
- SCiForest (Liu et al., 2010) selects splits to maximise separation rather than randomly, targeting clustered anomalies.
- Local structure blindness: a point sparse relative to its local cluster but inside the global envelope scores as normal; this is the regime where LOF dominates — Breunig et al.'s local ratio is exactly what isolation depth lacks.
- Interpretability: per-feature depth contributions (e.g. the SHAP-style DIFFI method, Carletti et al., 2020) recover which columns drove a short path — often a deployment requirement.
In ADBench's 57-dataset comparison (Han et al., 2022), isolation forest sits in the strongest tier of unsupervised detectors while being among the cheapest — the reason it remains the default first tool despite its age.
What to learn next
- Local outlier factor — the local-density view isolation trees lack.
- Evaluating an outlier detector — deciding whether these scores are any good.
- Random forest — the supervised cousin, if trees themselves are new.