Outlier and Anomaly Detection

Z-scores, IQR fences and MAD

The classic z-score rule can be blinded by the very outliers it hunts, because outliers drag the mean and inflate the spread — median-based rules stay honest.

Read these first

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The oldest outlier rules measure how far a value sits from "typical".

But if "typical" is worked out with the outlier included, the outlier helps hide itself.

A billionaire walks into a tea shop with nine regular customers. The "average customer wealth" in that room rockets into the crores. Measured against that absurd average, the billionaire suddenly looks... not so unusual. And the nine regulars now all look oddly poor.

One stranger bent the ruler that was supposed to measure him.

Why this matters

The classic rule of thumb says: compute the average, compute the typical spread around it, and flag anything sitting several spreads away. The bent-ruler problem is that the outlier itself inflates both the average and the spread. A monstrous value can end up looking borderline-acceptable on a ruler it bent. With two or three outliers helping each other, all of them can slip under the bar together.

This failure has a name: masking — outliers hiding each other by corrupting the measurements used to find them.

How it works

The cure is to build the ruler from measurements a stranger cannot bend. Two do the job:

The median is the middle value when everything is sorted. A billionaire can walk in and the middle customer barely changes. Values based on position in the sorted order ignore how extreme the extremes are.

values:   32  35  36  38  40  41  43  45  47  900
                          ^^
mean   = 125.7   <- dragged far above 9 of 10 values
median = 40.5    <- barely moved

From the median come two robust rules. The IQR fence takes the range covered by the middle half of the data. It flags anything far outside that range. These are the whiskers in every box plot. The MAD, the median distance of values from the median, is a spread measurement immune to bending; distances measured in MADs replace the classic rule safely.

A real example you have seen

Your electricity bill. Billing systems flag a month that leaps far outside your usual range before assuming the meter reading is real. If they measured "usual" with a mean that one faulty 10,000-unit reading had corrupted, that faulty reading would look almost normal. Median-based checks are standard in exactly such systems.

Remember this

  • Outliers bend the mean and the spread, then hide behind the damage — that is masking.
  • Median-based rules (IQR fence, MAD) cannot be bent by a few wild values.
  • Reach for the classic z-score rule only when the data is verified clean.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Outputs verified with numpy 1.26.4.

Watching the z-score fail, and the robust rules succeed

robust_rules.py
import numpy as np

# Nine ordinary salaries and one director, in thousands per month.
salaries = np.array([32, 35, 36, 38, 40, 41, 43, 45, 47, 900.0])

z = (salaries - salaries.mean()) / salaries.std()
print("z-score of the 900:", round(z[-1], 4))

# Add a second big salary. Watch both hide each other.
padded = np.append(salaries, 850.0)
z2 = (padded - padded.mean()) / padded.std()
print("z-scores after adding an 850:", round(z2[-2], 2), "and", round(z2[-1], 2))

q1, q3 = np.percentile(padded, [25, 75])
fence_high = q3 + 1.5 * (q3 - q1)
print("IQR fence catches:", padded[padded > fence_high])

med = np.median(padded)
mad = np.median(np.abs(padded - med))
robust_z = 0.6745 * (padded - med) / mad
print("robust z-scores of the two:", np.round(robust_z[-2:], 1))
Output
z-score of the 900: 2.9996
z-scores after adding an 850: 2.2 and 2.04
IQR fence catches: [900. 850.]
robust z-scores of the two: [115.9 109.1]

The walkthrough

A salary 20 times everyone else's scores z = 2.9996. The textbook threshold is "flag beyond 3". This value fails the test. In fact with 10 data points, no single value can push far past z = 3 — the outlier inflates the standard deviation as fast as it inflates its own distance. The rule contains its own blind spot.

Two outliers mask each other down to 2.2 and 2.04. Each inflates the spread that measures the other. A z-threshold of 3 now misses both by a mile. More outliers make the classic rule worse, exactly when you need it most.

The IQR fence catches both without blinking. q1 and q3 are the values a quarter and three-quarters of the way up the sorted order; wild values cannot move them. The multiplier 1.5 is Tukey's convention — raise it to 3.0 for a "far out" fence that flags less.

The robust z-scores read 115.9 and 109.1. Same idea as a z-score, with median replacing mean and MAD replacing standard deviation. The 0.6745 factor rescales MAD so the score reads in familiar standard-deviation-like units for bell-shaped data. Threshold at 3.5, per the Iglewicz–Hoaglin convention.

Common mistakes

Deleting flagged rows automatically. A flag is a question, not a verdict. That ₹9,00,000 salary might be the CEO — real, and analytically important. Investigate, clean deliberately, and log what you removed.

Applying any of these rules to skewed data. Monthly incomes, order values and page views have long natural tails. On skewed data, both z and MAD rules flag legitimate tail values relentlessly. Work on a log scale first, or fence the two sides asymmetrically.

Computing z-scores column by column and calling multivariate data covered. A 55 kg weight is normal; a 190 cm height is normal; the combination is remarkable. Per-column rules never see combinations — that gap is what isolation forest and friends exist for.

MAD = 0 on chunky data. When over half the values are identical (lots of zeros, say), the MAD is zero and the robust z divides by it. Fall back to the IQR fence or handle the zero-inflation explicitly.

Try it yourself

Generate 1,000 values with rng.lognormal(mean=10, sigma=1) — a realistic right-skewed income shape. Count how many the three rules each flag. Then repeat on np.log of the values and compare counts. The lesson writes itself.

What to learn next

Researcher — Mathematics and papers.

Breakdown points

The breakdown point of an estimator is the smallest fraction of arbitrarily bad data that can drive it to arbitrary values. The mean and standard deviation break at $1/n$ — one point suffices, as the developer block demonstrates. The median and MAD achieve the maximum possible breakdown of $1/2$ (Hampel, 1971; the framework text is Huber and Ronchetti, Robust Statistics, 2nd ed., 2009).

The classical-rule failure is quantifiable: with $n$ points, the maximum attainable z-score is $(n-1)/\sqrt{n}$, so for $n = 10$ no point can exceed $z = 2.846$ under the textbook one-outlier configuration — a "flag at 3" rule is provably unable to fire on small clean-plus-one samples. (The 2.9996 above edges past that bound only because the nine regulars also deviate from the mean.)

The MAD constants

For a sample $X$ with median $m$, $\mathrm{MAD} = \mathrm{median}_i\, |x_i - m|$. Under a normal distribution, $\mathrm{MAD} = \sigma \, \Phi^{-1}(3/4) \approx 0.6745\,\sigma$.

Where:

  • $\sigma$ — the true standard deviation.
  • $\Phi^{-1}$ — the standard normal quantile function; $\Phi^{-1}(3/4) \approx 0.6745$.

Hence the two equivalent conventions: multiply MAD by $1.4826$ to get a consistent $\sigma$ estimate, or (as in the code) multiply deviations by $0.6745/\mathrm{MAD}$. The modified z-score threshold of 3.5 is from Iglewicz and Hoaglin (1993), How to detect and handle outliers.

MAD's efficiency at the normal is poor (about 37%), motivating the estimators $S_n$ and $Q_n$ of Rousseeuw and Croux (1993) — same 50% breakdown, far better efficiency, $O(n \log n)$ algorithms.

Beyond one dimension

The multivariate analogue of the z-score is the Mahalanobis distance $d^2(x) = (x - \mu)^\top \Sigma^{-1} (x - \mu)$, with $\mu$ the mean vector and $\Sigma$ the covariance matrix — and it inherits the same fragility, since outliers corrupt $\hat{\Sigma}$. The robust replacement estimates location and scatter by the Minimum Covariance Determinant (Rousseeuw, 1984): find the half-sample whose covariance has the smallest determinant, computable via FastMCD (Rousseeuw and Van Driessen, 1999) — in scikit-learn as EllipticEnvelope. Distances then follow a known chi-squared reference under normality, giving principled cut-offs.

Swamping is masking's mirror: outliers dragging the estimates so far that normal points get flagged. Sequential-deletion identifiers (Hampel identifier, Rosner's generalised ESD test, 1983) formalise iterative detection with controlled error rates; Rosner's test requires an upper bound on the outlier count, which is its honest admission of the problem's circularity.

What to learn next

What to learn next

These follow on from what you just read.

  • Outlier and Anomaly Detection

    Isolation forest

    Isolation forest flags the points that random yes/no splits separate from the crowd in only a few cuts — anomalies are easy to isolate, and that ease is the score.

  • Outlier and Anomaly Detection

    Local outlier factor

    LOF compares each point's crowding against its neighbours' crowding, so it catches points that are odd for their own neighbourhood even when they look globally ordinary.

  • Outlier and Anomaly Detection

    One-class SVM

    A one-class SVM learns a single flexible fence around your normal data, with a dial that states upfront how many training points you are willing to leave outside.