Linear Models and Regularisation

Ridge regression

Ridge regression tames wild coefficients by charging a penalty on their size, trading a little bias for a lot of stability.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Ridge regression is linear regression with a penalty on big coefficients, which keeps the model calm when the data tries to make it wild.

Think of a set of old-fashioned weighing scales that wobble. Put the same rice on them twice and the needle settles in two different places. A shopkeeper fixes this with a gentle damper — a small resistance that stops the needle swinging wildly. The reading becomes slightly conservative, but it becomes repeatable.

Ridge is that damper for linear regression. The damper's strength is a knob called alpha.

Why it exists

Plain linear regression has a fragile spot: features that carry overlapping information. A flat's size in square feet and its number of rooms tell nearly the same story. The model cannot decide how to split credit between them — so tiny accidents in the data decide instead.

The result is coefficients that swing violently. Collect the data again and "size" might flip from strongly positive to negative, with "rooms" absorbing the difference. Predictions may stay adequate, but the model's explanation is garbage, and on new data the swings eventually hurt predictions too.

Ridge adds one rule to the fitting: big coefficients cost extra. To claim a large effect, a feature must earn it with strong, consistent evidence. Overlapping features stop fighting and share the credit.

How it works

plain regression:  find coefficients that fit the data best

ridge regression:  find coefficients that fit well AND stay small

fit twice on similar data:
   plain:  size: -1.3, rooms: +12.3   →  size: +8.4, rooms: +2.5   (chaos)
   ridge:  size: +3.9, rooms: +3.8   →  size: +4.1, rooms: +4.2   (steady)

A real example you have seen

Any house-price site publishes numbers like "one extra room adds about 4 lakhs". Behind such claims sits a model whose coefficients must not flip sign every time a new month of listings arrives. Stabilised regression is how those numbers stay printable.

Remember this

  • Overlapping features make plain regression's coefficients swing wildly.
  • Ridge charges a penalty for big coefficients, so features share credit and stay stable.
  • The knob alpha sets the penalty: zero is plain regression, huge is a flat model.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

Two halves of the same data, two stories

ridge_stability.py
import numpy as np
from sklearn.linear_model import LinearRegression, Ridge
from sklearn.preprocessing import StandardScaler

rng = np.random.default_rng(7)
size = rng.uniform(400, 1600, 30)                 # flat size in sq ft
rooms = size / 350 + rng.normal(0, 0.1, 30)       # rooms is nearly size in disguise
X = StandardScaler().fit_transform(np.column_stack([size, rooms]))
price = 20 + 0.03 * size + rng.normal(0, 3, 30)   # price in lakhs; only size matters

half1, half2 = slice(0, 15), slice(15, 30)
for name, model in [("OLS  ", LinearRegression()), ("Ridge", Ridge(alpha=10.0))]:
    c1 = model.fit(X[half1], price[half1]).coef_.round(1)
    c2 = model.fit(X[half2], price[half2]).coef_.round(1)
    print(f"{name} first half {c1}  second half {c2}")
Output
OLS   first half [-1.3 12.3]  second half [8.4 2.5]
Ridge first half [3.9 3.8]  second half [4.1 4.2]

The walkthrough

Read the OLS row twice. On the first half, ordinary least squares (OLS — plain linear regression) says size lowers price and rooms carry everything. On the second half, the story reverses. Same city, same month, same true relationship — the coefficients are noise wearing a suit.

The ridge row is boring, and boring is the point. Roughly [4, 4] both times: the two overlapping features split the credit, and resampling barely moves them. Note ridge does not choose between the twins — for that behaviour see lasso.

Why StandardScaler first, always. The penalty charges coefficients by size. A feature measured in big units (square feet) needs a tiny coefficient, so it is barely charged, while a small-unit feature pays full price for the same influence. Scaling puts every feature on one price list. Unscaled ridge is a subtle, common bug.

Choosing alpha. alpha=10 here is illustrative. In practice sweep a log scale with built-in cross-validation:

python
from sklearn.linear_model import RidgeCV
best = RidgeCV(alphas=np.logspace(-3, 3, 13)).fit(X, price)
print("chosen alpha:", best.alpha_)
Output
chosen alpha: 1.0

RidgeCV uses efficient leave-one-out cross-validation by default — cheap even on large data.

Common mistakes

Skipping the scaler. Covered above; wrap it in a pipeline — make_pipeline(StandardScaler(), Ridge(alpha=10)) — so it cannot be forgotten.

Judging ridge by training error. Ridge deliberately fits training data slightly worse than OLS. Its win shows on held-out data and in coefficient stability. Compare with cross-validation, never on the training set.

Expecting zeros. Ridge shrinks coefficients toward zero but never exactly to zero. Twenty junk features stay in the model with small coefficients. If you need automatic feature removal, that is lasso's job.

Penalising the intercept. The baseline level of the target is not model complexity. scikit-learn leaves the intercept unpenalised — if you implement ridge yourself, remember to do the same.

Try it yourself

Sweep alpha through 0.01, 1, 100, and 10000, printing both halves' coefficients each time. Watch chaos fade into stability and then into everything-crushed-toward-zero. Where does the damping stop helping and start erasing signal?

What to learn next

Researcher — Mathematics and papers.

Objective and closed form

Ridge minimises the penalised least-squares objective:

$$ \hat{\beta}^{ridge} = \arg\min_{\beta} \; \lVert y - X\beta \rVert_2^2 + \lambda \lVert \beta \rVert_2^2 $$

Where:

  • $X \in \mathbb{R}^{n \times d}$ — the (centred, scaled) design matrix; $y \in \mathbb{R}^n$ — targets.
  • $\beta$ — coefficients; the intercept is excluded from the penalty.
  • $\lambda \geq 0$ — the penalty strength (scikit-learn's alpha).
  • $\lVert \beta \rVert_2^2 = \sum_j \beta_j^2$ — the squared L2 norm.

The solution is closed-form:

$$ \hat{\beta}^{ridge} = (X^\top X + \lambda I)^{-1} X^\top y $$

Adding $\lambda I$ lifts every eigenvalue of $X^\top X$ by $\lambda$, so the inverse exists even when features are exactly collinear or $d > n$ — the algebraic reason ridge cannot blow up where OLS does.

The spectral view

With singular value decomposition $X = U \Sigma V^\top$ (singular values $\sigma_j$), predictions decompose as:

$$ X\hat{\beta}^{ridge} = \sum_j u_j \, \frac{\sigma_j^2}{\sigma_j^2 + \lambda} \, u_j^\top y $$

OLS uses factor 1 for every direction; ridge shrinks each direction by $\sigma_j^2 / (\sigma_j^2 + \lambda)$. Strong directions (large $\sigma_j$) pass almost untouched; weak, noise-dominated directions are crushed. Ridge is a smooth low-pass filter on the data's spectrum — precisely targeted at the unstable directions multicollinearity creates.

Bias–variance accounting

Ridge estimates are biased: $\mathbb{E}[\hat{\beta}^{ridge}] \neq \beta$ for $\lambda > 0$. But variance falls faster than squared bias rises for small $\lambda$: Hoerl and Kennard (1970), Ridge regression: biased estimation for nonorthogonal problems, proved there always exists $\lambda > 0$ with lower total mean-squared error than OLS. A free lunch, paid for by accepting bias.

Bayesian reading: ridge is the posterior mode under prior $\beta \sim \mathcal{N}(0, \tau^2 I)$ with $\lambda = \sigma^2 / \tau^2$ — the full posterior treatment is Bayesian linear regression.

Cost and context

Direct solve: $O(nd^2 + d^3)$; the SVD route prices all $\lambda$ values at once, which is how RidgeCV makes leave-one-out CV nearly free (closed-form LOO residuals). For huge sparse problems, conjugate-gradient and SAGA solvers apply.

Ridge (Tikhonov regularisation; Tikhonov, 1943, for ill-posed integral equations) is the template for weight decay in deep learning — the same $\ell_2$ penalty, SGD-shaped (Krogh and Hertz, 1992; decoupled form Loshchilov and Hutter, 2019, AdamW). Kernelised, it becomes kernel ridge regression, the mean function of Gaussian process regression.

What to learn next

What to learn next

These follow on from what you just read.

  • Linear Models and Regularisation

    Lasso regression

    Lasso's penalty pushes weak coefficients to exactly zero, so the model selects its own features while it fits.

  • Linear Models and Regularisation

    Elastic net

    Elastic net mixes ridge's credit-sharing with lasso's feature-dropping, fixing lasso's coin-flip behaviour on correlated features.

  • Linear Models and Regularisation

    Multicollinearity and VIF

    When features tell the same story, coefficients turn unstable and unreadable — VIF is the number that catches it before you publish nonsense.