Linear Models and Regularisation

Robust regression

A handful of corrupted points can drag an ordinary regression line anywhere — robust methods cap or discard their influence and recover the real trend.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Robust regression fits the trend of the majority, refusing to let a few wild points drag the line around.

Ten friends rate a restaurant between 8 and 9, and one furious ex-employee gives it 0. The average crashes to about 7.7 — below every genuine rating. One voice, shouting, outweighed nine speaking normally. Any method built on averages inherits this weakness: the loudest point wins.

Ordinary linear regression is built on averages. Robust regression is the family of fixes.

Why it exists

Regression fits its line by punishing mistakes. Ordinary least squares punishes a big miss far more harshly than a small one. Twice the miss does not cost twice as much. It costs out of all proportion. So one wildly wrong point shouts down a whole crowd of honest ones. The line bends toward the shouter, sometimes until the real trend vanishes entirely.

And corrupted points are not rare. Sensors glitch. Someone types 950 instead of 95. A scraper reads the wrong column for one row. Real datasets ship with a percentage of garbage as a fact of life.

Two classic defences:

  • Cap the shouting. Punish small misses normally, but stop raising the penalty once a miss gets large. A wild point still speaks, at normal volume. This is the Huber approach.
  • Vote the shouters out. Repeatedly fit the line to small random samples of points, and keep the line that the largest crowd of points agrees with. Points that agree with nothing are declared outliers and ignored. This is RANSAC.

How it works

        ●  ●   ← glitches                    ●  ●  (ignored)
   │                                     │
   │        ___────                      │            ──
   │  ──────       OLS line              │       ──── 
   │ ∙ ∙ ∙∙ ∙ ∙∙ ∙  dragged up           │ ∙ ∙∙─∙──∙∙ ∙  robust line
   │                toward glitches      │      follows the crowd
   └──────────────                       └──────────────

A real example you have seen

Every "average income" statistic that feels wrong to you is this lesson — one billionaire in the sample and the mean stops describing anyone. It is why reports quote median income. Robust regression brings that same median instinct to fitted lines. GPS apps use it to discard impossible location jumps. Camera panorama-stitching depends on it too, running RANSAC under the hood.

Remember this

  • Squared punishment lets one wild point outshout hundreds of honest ones.
  • Huber caps the shouting; RANSAC votes the shouters out.
  • If your data can contain glitches — and it can — an OLS-only fit is a gamble.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

Four glitches versus three fitters

robust_fit.py
import numpy as np
from sklearn.linear_model import LinearRegression, HuberRegressor, RANSACRegressor

rng = np.random.default_rng(2)
hours = rng.uniform(0, 10, 40)                  # machine running hours
temp = 5 + 2 * hours + rng.normal(0, 1, 40)     # sensor temperature in deg C
temp[:4] = [95, 90, 88, 99]                     # four glitched readings

X = hours.reshape(-1, 1)
ols = LinearRegression().fit(X, temp)
hub = HuberRegressor().fit(X, temp)
ran = RANSACRegressor(random_state=0).fit(X, temp)

print("truth : slope 2.00, intercept 5.00")
print(f"OLS   : slope {ols.coef_[0]:.2f}, intercept {ols.intercept_:.2f}")
print(f"Huber : slope {hub.coef_[0]:.2f}, intercept {hub.intercept_:.2f}")
print(f"RANSAC: slope {ran.estimator_.coef_[0]:.2f}, intercept {ran.estimator_.intercept_:.2f}")
print("points RANSAC threw out:", int((~ran.inlier_mask_).sum()))
Output
truth : slope 2.00, intercept 5.00
OLS   : slope 0.00, intercept 23.28
Huber : slope 1.96, intercept 5.50
RANSAC: slope 2.03, intercept 4.92
points RANSAC threw out: 4

The walkthrough

OLS reported slope 0.00. Not slightly wrong — the trend is gone. Four glitches among forty points (10% contamination) were enough to flatten a strong, clean relationship into "temperature does not depend on hours". If you only remember one output block from this section, make it this one.

Huber landed at 1.96. Its loss is quadratic for small residuals and linear beyond a threshold (epsilon, default 1.35 in scaled units). The glitches still participate, at capped volume — enough participation to matter little.

RANSAC landed at 2.03 and named the culprits. inlier_mask_ flags which points the consensus model accepted; exactly our 4 planted glitches were rejected. That mask is a diagnostic gift: print the rejected rows and you often find the cause — one bad sensor, one bad day, one bad join in your pipeline.

Which one when? Moderate noise plus occasional wild values: Huber, deterministic and smooth. Gross corruption, mixed data sources, or a need to identify outliers: RANSAC. Vertical-only outliers with many features: also consider TheilSenRegressor (median-of-slopes, deterministic, slower).

Common mistakes

Deleting outliers by eyeballing residuals from an OLS fit. The glitches bent the line toward themselves — OLS residuals are computed from a corrupted line, so genuine points can look like outliers and glitches can look fine. This failure mode is called masking. Fit robustly first; judge residuals second.

Treating every outlier as garbage. Sometimes the "glitch" is the discovery — the fraud, the failing machine. Robust fits separate the crowd from the exceptions; investigating the exceptions is your job, not the model's.

Forgetting RANSAC is randomised. Different seeds can select different consensus sets. Fix random_state, and on small data check stability across seeds before shipping conclusions.

Scaling-blind Huber. HuberRegressor's epsilon is defined on standardised residuals internally, but features still deserve scaling for the optimiser's sake, as ever.

Try it yourself

Increase contamination: replace 12 of the 40 readings with glitches. Which fitter breaks first? Then move the glitches horizontally instead — put them at hours = 30 with plausible temperatures — and watch what that does to each method. Leverage points are a different beast from vertical outliers.

What to learn next

Researcher — Mathematics and papers.

M-estimation

Replace OLS's squared loss with $\rho$ applied to scaled residuals:

$$ \hat\beta = \arg\min_\beta \sum_{i=1}^{n} \rho!\left( \frac{y_i - x_i^\top \beta}{\hat\sigma} \right) $$

Where:

  • $\rho$ — a symmetric loss; $\hat\sigma$ — a robust scale estimate (MAD or jointly estimated, as in HuberRegressor).
  • The Huber loss (Huber, 1964, Robust estimation of a location parameter): quadratic for $|u| \leq \epsilon$, linear beyond — convex, hence a unique optimum, solved by IRLS.
  • Tukey's biweight flattens completely beyond a cutoff — wild points get zero gradient — but non-convexity demands good initialisation.

The influence function $\psi = \rho'$ formalises "shouting": for OLS $\psi(u) = u$ is unbounded (influence grows without limit); Huber clips it at $\pm\epsilon$; redescending $\psi$ (Tukey) returns to zero. Bounded influence is the defining property of a robust estimator (Hampel et al., 1986).

Breakdown points

The breakdown point is the smallest contamination fraction that can carry the estimate to infinity. OLS: $0$ (one point suffices — the slope-0.00 demo). Huber M-estimators: $0$ against arbitrary leverage outliers, though excellent against vertical ones. High-breakdown estimators reach the theoretical maximum near $1/2$:

  • LMS/LTS — least median/trimmed squares (Rousseeuw, 1984): minimise the median (or trimmed sum) of squared residuals; combinatorial, approximated by subsampling (FAST-LTS).
  • S- and MM-estimators (Yohai, 1987): high breakdown and high Gaussian efficiency — the default in R's MASS::rlm(method="MM") and the recommended modern compromise.

RANSAC (Fischler and Bolles, 1981, Random sample consensus) reaches high breakdown by hypothesise-and-verify: sample minimal subsets ($p$ points for $p$ parameters), fit, count inliers within a tolerance, keep the best consensus. Success probability with inlier fraction $w$ after $k$ rounds is $1 - (1 - w^p)^k$, giving the standard formula for choosing $k$. It remains ubiquitous in computer vision for fundamental-matrix and homography estimation.

Efficiency trade-off and context

Robustness is bought with efficiency at the Gaussian ideal: Huber with $\epsilon = 1.345$ retains 95% efficiency versus OLS on clean normal data — the standard tuning. Quantile regression at $\tau = 0.5$ (least absolute deviations) is the historical ancestor (Boscovich, 18th c.) with breakdown 0 against leverage but full resistance to vertical outliers.

In modern ML, Huber loss reappears as the default regression loss in gradient boosting (loss="huber"), as Smooth L1 in object-detection box regression (Fast R-CNN; Girshick, 2015), and robust aggregation ideas resurface in federated learning against poisoned clients.

What to learn next