Linear Models and Regularisation

Lasso regression

Lasso's penalty pushes weak coefficients to exactly zero, so the model selects its own features while it fits.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Lasso regression fits a linear model that throws away useless features by setting their coefficients to exactly zero.

Think of packing one small suitcase for a long trip. Every item must argue for its place: the raincoat earns it, the third pair of shoes does not. Items that cannot pay their way stay home — not "packed but smaller", but left out entirely.

Ridge regression is a different packer: it shrinks everything a little and still packs all of it. Lasso leaves things behind.

Why it exists

Real tables carry dozens or hundreds of candidate features, and usually only a handful matter. A model that keeps all of them has three problems. It reads badly — nobody can explain a formula with ninety tiny terms. It costs more — every feature must be collected, cleaned and monitored forever. And it overfits — each junk feature is one more way to memorise noise.

Ridge's penalty cannot fix the first two: it makes junk coefficients small, never zero. Lasso changes the shape of the penalty, and that one change makes weak coefficients snap exactly to zero. The model does its own feature selection while fitting, and the strength knob alpha sets how ruthless the packing is.

How it works

10 candidate features, only 3 truly matter

ridge:  [ 2.9  0.02 -0.03 -2.1  0.02 -0.0 -0.02  0.5  0.03 -0.02 ]
         everything survives, junk gets tiny

lasso:  [ 2.9  0    0    -2.0  0     0    0     0.4  0    0     ]
         junk is GONE — the model kept features 1, 4 and 8

A real example you have seen

A hospital predicting patient risk can measure two hundred things per patient. Each test costs money and time. A lasso-style model that keeps twelve of them is not a worse model with fewer features. It is a cheaper form and a shorter consent list. It is a formula a doctor can read aloud.

Remember this

  • Lasso sets weak coefficients to exactly zero — features are removed, not shrunk.
  • The surviving features are lasso's built-in feature selection.
  • Bigger alpha, smaller suitcase: more features left behind.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

Ten candidates, three real

lasso_select.py
import numpy as np
from sklearn.linear_model import Lasso, Ridge

rng = np.random.default_rng(0)
X = rng.normal(size=(80, 10))                     # 10 candidate features
y = 3 * X[:, 0] - 2 * X[:, 3] + 0.5 * X[:, 7] + rng.normal(0, 0.5, 80)

lasso = Lasso(alpha=0.1).fit(X, y)
ridge = Ridge(alpha=1.0).fit(X, y)

print("lasso:", lasso.coef_.round(2))
print("ridge:", ridge.coef_.round(2))
print("features lasso kept:", np.flatnonzero(lasso.coef_).tolist())
Output
lasso: [ 2.87  0.    0.   -2.01  0.    0.    0.    0.36  0.   -0.  ]
ridge: [ 2.92  0.02 -0.03 -2.07  0.02 -0.   -0.02  0.48  0.03 -0.02]
features lasso kept: [0, 3, 7]

The truth planted in y used features 0, 3 and 7. Lasso recovered exactly that set from noisy data. Ridge kept all ten.

The walkthrough

Why does the penalty's shape matter so much? Ridge charges the square of each coefficient: near zero, the charge fades to almost nothing, so there is never a reason to finish the job. Lasso charges the absolute value: the price per unit stays full all the way down. Any feature whose contribution cannot cover that flat rate gets pushed to exactly zero. Same idea, different tariff — completely different behaviour.

Watch alpha control the ruthlessness:

python
for a in (0.01, 0.1, 0.5, 1.0):
    m = Lasso(alpha=a).fit(X, y)
    print(f"alpha={a}: kept {np.count_nonzero(m.coef_)} of 10, coef[0]={m.coef_[0]:.2f}")
Output
alpha=0.01: kept 8 of 10, coef[0]=2.95
alpha=0.1: kept 3 of 10, coef[0]=2.87
alpha=0.5: kept 2 of 10, coef[0]=2.51
alpha=1.0: kept 2 of 10, coef[0]=2.04

Two things happen at once as alpha grows. Features drop out — 8, then 3, then 2. And the survivors shrink: the true coefficient is 3.0, but at alpha=1.0 lasso reports 2.04. Lasso taxes the innocent along with the guilty; that shrinkage bias is the price of the selection.

Choosing alpha for real: LassoCV(cv=5).fit(X, y) sweeps a path of alphas efficiently and picks by cross-validation. Scale features first — the flat tariff, like ridge's, assumes every feature is priced in the same units.

Common mistakes

Forgetting to scale. One big-unit feature needs a tiny coefficient, dodges the tariff, and survives while honest features die. make_pipeline(StandardScaler(), Lasso(...)).

Reading the kept set as ground truth. With correlated features, lasso keeps one representative and zeroes its twins — which twin survives can flip with a new sample of data. Stability matters before you print "these are the 12 factors that matter". See elastic net.

Using survivor coefficients for effect sizes. They are shrunk, as the sweep showed. A common remedy is the relaxed lasso: select with lasso, then refit plain regression on the survivors.

Ignoring convergence warnings. ConvergenceWarning: Objective did not converge means the coordinate-descent solver ran out of iterations — often unscaled data or a near-zero alpha. Scale, raise max_iter, or raise alpha.

Try it yourself

Add a feature that is a noisy copy of feature 0 — X[:, 5] = X[:, 0] + rng.normal(0, 0.05, 80) — and refit at alpha=0.1 a few times with different rng seeds. Watch which twin survives. That instability is the opening argument of the elastic net lesson.

What to learn next

Researcher — Mathematics and papers.

Objective

$$ \hat{\beta}^{lasso} = \arg\min_{\beta} \; \frac{1}{2n} \lVert y - X\beta \rVert_2^2 + \alpha \lVert \beta \rVert_1 $$

Where:

  • $\lVert \beta \rVert_1 = \sum_j |\beta_j|$ — the L1 norm; the intercept is unpenalised.
  • $\alpha$ — penalty strength; $n$ — sample count (scikit-learn's parameterisation divides the loss by $n$, so published $\lambda$ values differ by that factor).

No closed form exists — $\lVert \cdot \rVert_1$ is non-differentiable at zero, and exactly that kink produces sparsity. Geometrically: the L1 constraint region is a diamond whose corners sit on the axes; the loss contours typically first touch a corner, where some coordinates are zero. The L2 ball is round and has no corners, hence ridge's no-zeros behaviour.

Optimisation

For a single coefficient with others fixed, the solution is soft-thresholding:

$$ \beta_j \leftarrow \operatorname{sign}(z_j)\,\max(|z_j| - \alpha, 0) $$

with $z_j$ the univariate least-squares update for coordinate $j$ (standardised features). Coordinate descent cycles this rule to convergence (Friedman, Hastie, Tibshirani, 2010 — the glmnet algorithm scikit-learn mirrors), at $O(nd)$ per full cycle, exploiting warm starts along an alpha path. The LARS algorithm (Efron et al., 2004) traces the exact piecewise-linear solution path in $O(nd\,\min(n,d))$.

Theory

  • Origin: Tibshirani (1996), Regression shrinkage and selection via the lasso; the signal-processing twin is basis pursuit (Chen, Donoho, Saunders, 1998).
  • Support recovery requires the irrepresentable condition — junk features must not be too correlated with true ones (Zhao and Yu, 2006; Wainwright, 2009 gives sharp sample-size thresholds $n \gtrsim s \log d$ for $s$ true features).
  • Prediction consistency holds far more generally, with oracle-style bounds under restricted-eigenvalue conditions (Bickel, Ritov, Tsybakov, 2009).
  • The estimator is biased; the adaptive lasso (Zou, 2006) reweights penalties to gain oracle properties, and debiased/desparsified variants (Zhang and Zhang, 2014; van de Geer et al., 2014) recover valid confidence intervals — the entry point to post-selection inference.

Standing

Lasso remains the default tool for interpretable sparse regression in genomics ($d \gg n$ by design), econometrics and risk modelling. Its penalty generalises: group lasso (Yuan and Lin, 2006) zeroes whole feature blocks; graphical lasso sparsifies precision matrices; L1 penalties on network weights drive pruning research in deep learning.

What to learn next

What to learn next

These follow on from what you just read.

  • Linear Models and Regularisation

    Elastic net

    Elastic net mixes ridge's credit-sharing with lasso's feature-dropping, fixing lasso's coin-flip behaviour on correlated features.

  • Linear Models and Regularisation

    Multicollinearity and VIF

    When features tell the same story, coefficients turn unstable and unreadable — VIF is the number that catches it before you publish nonsense.

  • Linear Models and Regularisation

    Generalised linear models

    GLMs are one recipe with three swappable parts — a linear score, a link that bends it to the target's range, and a noise family — uniting linear, logistic and Poisson regression.