Outlier and Anomaly Detection

Local outlier factor

LOF compares each point's crowding against its neighbours' crowding, so it catches points that are odd for their own neighbourhood even when they look globally ordinary.

Read these first

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The local outlier factor compares how crowded a point's spot is against how crowded its neighbours' spots are.

Strange here means strange for your neighbourhood, not strange for the whole map.

A house 500 metres from its nearest neighbour means nothing by itself. In a packed Hyderabad colony, that house is bizarre — everyone else lives wall to wall. In Punjab farmland, it is completely ordinary — every farmhouse sits that far apart.

Same distance, opposite meanings. The neighbourhood defines normal.

Why this matters

Global rules use one yardstick for the entire dataset. Real data almost never has one honest yardstick. It contains regions of different natural density. Think of a tight cluster of salaried customers beside a diffuse cloud of business accounts.

Any single yardstick must fail one region. Set it for the tight cluster and it floods the diffuse region with false alarms. Set it for the diffuse region and it goes blind inside the tight cluster — where a modestly-displaced point may be the fraud you wanted.

LOF, the local outlier factor, drops the single yardstick. Every point is judged by its own neighbourhood's standard.

How it works

dense colony            spread-out farms

  xxx                    x     x
  xxxx   <- o                x       x
  xxx                    x       x
      odd here!          o  <- ordinary here

For each point, LOF asks three questions. How crowded is the space right around me? How crowded is the space around each of my nearest neighbours? And what is the ratio between the two?

A point exactly as crowded as its neighbours scores about 1 — a local citizen. A point living in noticeably emptier space than its neighbours scores well above 1 — a local stranger. The score is a comparison, so it works in the colony and the farmland at the same time, with no shared yardstick anywhere.

A real example you have seen

Card-fraud screening again, one level deeper. A ₹15,000 purchase is unremarkable across all customers. Against your pattern of ₹200 tea-stall payments, it is a screaming outlier. Fraud teams call this peer-group analysis: everyone is judged against similar behaviour, which is LOF's exact logic.

Remember this

  • Strangeness is local: the same gap is odd in a colony, normal in farmland.
  • LOF is a ratio of crowdings — yours versus your neighbours'.
  • Around 1 means local citizen; well above 1 means local stranger.

What to learn next

  • One-class SVM — drawing one boundary around the normal, instead of scoring point by point.
  • HDBSCAN — the same local-density thinking, used to find clusters instead of strangers.
  • Evaluating an outlier detector — how to tell 4.74 is genuinely good.

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

A point global rules cannot catch

The odd point below sits closer to the data than most of the sparse cluster's members. Any global distance rule that flags it must flood the sparse cluster first. LOF catches it alone:

lof_local.py
import numpy as np
from sklearn.neighbors import LocalOutlierFactor

rng = np.random.default_rng(7)
tight = rng.normal(0, 0.3, size=(60, 2))       # dense housing colony
loose = rng.normal(8, 2.0, size=(60, 2))       # spread-out farmland houses
odd = np.array([[0.0, 1.8]])                   # near the colony, but not in it
X = np.vstack([tight, loose, odd])

lof = LocalOutlierFactor(n_neighbors=20)
flags = lof.fit_predict(X)
factor = -lof.negative_outlier_factor_          # LOF score, higher = stranger

print("LOF of the odd point:        ", round(factor[-1], 2))
print("median LOF, dense colony:    ", round(np.median(factor[:60]), 2))
print("median LOF, spread-out area: ", round(np.median(factor[60:120]), 2))
print("odd point flagged:", flags[-1] == -1)
Output
LOF of the odd point:         4.74
median LOF, dense colony:     1.04
median LOF, spread-out area:  1.08
odd point flagged: True

The walkthrough

Both clusters sit near 1.0 — that is the magic. The colony's houses are 0.3 apart; the farmland's are 2.0 apart. A global method sees two different worlds. LOF's ratio construction scores both as "normal for their own kind": 1.04 and 1.08.

The odd point scores 4.74 — its neighbourhood (the colony) is roughly five times more crowded than the space it lives in. Note its raw distance from the colony is only about 1.8, smaller than typical farmland gaps. Distance said ordinary; the ratio said stranger.

negative_outlier_factor_ is negated LOF, following the library-wide "lower means stranger" convention. Negate it back, as above, when you want the paper's scale.

n_neighbors=20 defines "the neighbourhood" and is the parameter that matters. Too small, and single random gaps look like anomalies. Larger than a real cluster's size, and that whole cluster gets judged against a different region's density. The default of 20 is a reasonable start; sweep it and prefer regions where results are stable.

Common mistakes

Comparing LOF scores across datasets or refits. The score is a local ratio, not a probability. A 1.8 in one dataset can be more alarming than a 3.0 in another. Rank within one fit; never port thresholds between fits.

Using fit_predict when you meant to screen new arrivals. Default LOF judges only its training data. For gate-duty on new points you need novelty=True, fitted on clean data — the two modes again.

Skipping feature scaling. Neighbourhoods come from distances, and one wide-ranged column decides them. Standardise first, always.

Running LOF on high-dimensional data unchanged. Past a few dozen dimensions, nearest-neighbour distances concentrate and everyone's neighbourhood looks equally crowded. Reduce dimensions first, or switch to methods that survive dimensionality better, like isolation forest.

Try it yourself

Move the odd point to [4.0, 4.0] — the empty no-man's-land between the clusters — and rerun. Predict first whether its LOF rises or falls, remembering the score compares against whichever points happen to be its 20 nearest neighbours.

What to learn next

  • One-class SVM — drawing one boundary around the normal, instead of scoring point by point.
  • HDBSCAN — the same local-density thinking, used to find clusters instead of strangers.
  • Evaluating an outlier detector — how to tell 4.74 is genuinely good.

Researcher — Mathematics and papers.

The construction

Breunig, Kriegel, Ng and Sander (2000), LOF: identifying density-based local outliers. Three layered definitions, with $k$ the neighbour count:

Reachability distance smooths distance measurements inside dense regions:

$$ \text{reach-dist}_k(p, o) = \max{ \text{k-dist}(o),\; d(p, o) } $$

Where $\text{k-dist}(o)$ is the distance from $o$ to its own $k$-th neighbour, and $d(p,o)$ the base metric. Points inside $o$'s $k$-neighbourhood all count as equally far — a variance-reduction device.

Local reachability density — inverse mean reachability from $p$ to its neighbour set $N_k(p)$:

$$ \text{lrd}_k(p) = \left( \frac{1}{|N_k(p)|} \sum_{o \in N_k(p)} \text{reach-dist}_k(p, o) \right)^{-1} $$

LOF — the mean ratio of neighbours' densities to $p$'s own:

$$ \text{LOF}_k(p) = \frac{1}{|N_k(p)|} \sum_{o \in N_k(p)} \frac{\text{lrd}_k(o)}{\text{lrd}_k(p)} $$

$\text{LOF} \approx 1$ for points whose density matches their neighbours'; the paper proves bounds placing inliers near 1 regardless of their cluster's absolute density — the formal version of "both clusters near 1.0" in the developer block.

Complexity and parameters

Naive computation needs all-pairs distances, $O(n^2 d)$. With tree indexes (KD/ball tree) the neighbour queries drop toward $O(n \log n)$ in low dimensions; index efficiency decays as $d$ grows, and past $d \approx 20$ the concentration-of-measure problem bites the semantics, not only the speed (Beyer et al., 1999). The original paper recommends scanning a range of $k$ and taking each point's maximum LOF — a detail most library defaults quietly drop.

Descendants

  • LoOP (Kriegel et al., 2009) recasts the ratio as a probability via a Gaussian assumption on distance distributions — comparable across datasets, addressing LOF's main usability flaw.
  • COF (Tang et al., 2002) replaces spherical neighbourhoods with chaining structure, for anomalies relative to curved manifolds.
  • INFLO (Jin et al., 2006) symmetrises neighbourhoods (reverse neighbours) to fix border effects between adjacent clusters of different density.
  • ODIN / kNN-distance is the degenerate baseline: score by $\text{k-dist}$ alone. Campos et al. (2016), the standard LOF-family benchmark, finds plain kNN-distance embarrassingly competitive — always report it as a baseline.
  • GLOSH from the HDBSCAN hierarchy provides a multi-scale local score without fixing $k$, inheriting density-tree consistency.

Aggarwal, Outlier Analysis (2nd ed., 2017), chapter 4, is the definitive survey of this family and its evaluation pitfalls.

What to learn next

  • One-class SVM — drawing one boundary around the normal, instead of scoring point by point.
  • HDBSCAN — the same local-density thinking, used to find clusters instead of strangers.
  • Evaluating an outlier detector — how to tell 4.74 is genuinely good.