Linear Models and Regularisation
Elastic net
Elastic net mixes ridge's credit-sharing with lasso's feature-dropping, fixing lasso's coin-flip behaviour on correlated features.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Elastic net blends ridge and lasso: it drops useless features like lasso, and shares credit between similar ones like ridge.
Two sisters run a tiffin service together, cooking side by side, splitting every task. Ask a customer to name "the cook" behind a great meal. They pick one sister, whoever answered the door. The other's work vanishes from the story. That is lasso with near-identical features: it names one and zeroes its twin. Ask instead "who contributed how much?" and the honest answer credits both — that is ridge's instinct.
Elastic net keeps both instincts on one leash.
Why it exists
Lasso has a real weakness. Two features can carry almost the same information. Monthly income and annual income. Humidity and dew point. Lasso arbitrarily keeps one and kills the other. Refit on next month's data and the survivor may swap. Any "these are the features that matter" report built on that is a coin flip wearing a lab coat.
Ridge never has this problem: twins share credit smoothly. But ridge never removes anything, so junk features linger forever.
Elastic net applies both penalties at once. The lasso part clears out genuine junk. The ridge part makes correlated survivors share — twins enter or leave the model together, in what is called the grouping effect. Two knobs control it: overall strength, and the mix between the two penalties.
How it works
features: x1 and x2 are near-twins, x3 is independent, truth uses all three
lasso : [ 3.9 0.0 1.0 ] one twin erased — which one is luck
elastic net : [ 2.0 2.0 1.0 ] twins share, junk (if any) still diesA real example you have seen
Genetics is the classic case: thousands of genes, many moving in lockstep. A lasso paper that names "the gene" for a trait, when thirty correlated neighbours would each pass, misleads everyone downstream. Elastic net was invented in exactly this setting — reports that keep the whole correlated group are more honest and more repeatable.
Remember this
- Lasso on correlated twins picks one arbitrarily — a coin flip in your feature report.
- Elastic net = lasso's dropping + ridge's sharing, on one leash.
- Twins enter or leave together — the grouping effect.
What to learn next
- Multicollinearity and VIF — measuring the twin problem before choosing a cure.
- Lasso regression — the selection machinery this lesson repaired.
- Feature engineering — avoiding accidental twins at the source.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU.
The twin experiment
import numpy as np
from sklearn.linear_model import Lasso, ElasticNet
rng = np.random.default_rng(2)
x1 = rng.normal(size=100)
x2 = x1 + rng.normal(0, 0.01, 100) # near-perfect copy of x1
x3 = rng.normal(size=100) # genuinely independent
X = np.column_stack([x1, x2, x3])
y = 2 * x1 + 2 * x2 + 1 * x3 + rng.normal(0, 0.3, 100) # truth: twins share
lasso = Lasso(alpha=0.05).fit(X, y)
enet = ElasticNet(alpha=0.05, l1_ratio=0.5).fit(X, y)
print("lasso :", lasso.coef_.round(2))
print("elastic net:", enet.coef_.round(2))lasso : [3.95 0. 0.97] elastic net: [1.97 1.95 0.97]
The truth is [2, 2, 1]. Lasso reported [3.95, 0, 0.97] — one twin swallowed the other's credit whole. Elastic net reported [1.97, 1.95, 0.97] — nearly the truth.
The walkthrough
Both models predict about equally well here. x1 and x2 are interchangeable, so 3.95 × x1 and 2 × x1 + 2 × x2 produce almost identical predictions. The damage is to interpretation: lasso's report says "x2 is irrelevant", which is false, and which twin gets erased flips with the random seed. If coefficients feed a decision — which sensor to keep, which test to run — that flip matters.
The two knobs. alpha is total penalty strength, as before. l1_ratio is the mix: 1.0 is pure lasso, 0.0 is pure ridge, 0.5 charges half of each. The grouping effect strengthens as l1_ratio falls.
Tune both together. The knobs interact, and ElasticNetCV sweeps them jointly at path speed:
from sklearn.linear_model import ElasticNetCV
cv = ElasticNetCV(l1_ratio=[0.2, 0.5, 0.8, 0.95], cv=5).fit(X, y)
print("best l1_ratio:", cv.l1_ratio_, " best alpha:", round(cv.alpha_, 4))best l1_ratio: 0.8 best alpha: 0.0054
A caution hides in that result. Cross-validation scores prediction only, and interchangeable twins predict identically however credit is split — so its l1_ratio choice here is nearly arbitrary. When the coefficients themselves will be read and acted on, that is a reason to prefer a lower l1_ratio than CV picked. The metric cannot see interpretability; you must.
Scaling, again, always. Both penalties price coefficients by magnitude; unscaled features rig the pricing. Pipeline a StandardScaler in front.
Common mistakes
Treating l1_ratio=0.5 as a law. It is a default, not a recommendation. Text-like data with many independent weak features often wants lasso-leaning mixes (0.9+); dense correlated tabular data wants ridge-leaning ones. Let ElasticNetCV decide.
Reading grouped coefficients as independent effects. Elastic net splits credit between twins roughly equally — that split reflects the penalty, not separated causal effects. No linear model can untangle what the data itself cannot distinguish; see multicollinearity.
Using ElasticNet(l1_ratio=0) for pure ridge. It works but the coordinate-descent solver is a poor fit and warns; use Ridge, whose solvers are built for it.
Forgetting that alpha scales differ across libraries. scikit-learn's (alpha, l1_ratio) maps to glmnet's (lambda, alpha) differently (and glmnet standardises internally). Porting hyperparameters between R and Python without translating is a classic silent bug.
Try it yourself
Rerun the twin experiment with seeds 0 through 4, recording which twin lasso keeps each time. Then loosen the twins — noise 0.5 instead of 0.01 — and find the correlation level at which lasso stops flipping. np.corrcoef(x1, x2) tells you where you are.
What to learn next
- Multicollinearity and VIF — measuring the twin problem before choosing a cure.
- Lasso regression — the selection machinery this lesson repaired.
- Feature engineering — avoiding accidental twins at the source.
Researcher — Mathematics and papers.
Objective
$$ \hat{\beta} = \arg\min_{\beta} \; \frac{1}{2n} \lVert y - X\beta \rVert_2^2
- \alpha \left( \rho \lVert \beta \rVert_1 + \frac{1 - \rho}{2} \lVert \beta \rVert_2^2 \right) $$
Where:
- $\alpha$ — overall penalty strength; $\rho \in [0, 1]$ — scikit-learn's
l1_ratio. - $\lVert \beta \rVert_1$, $\lVert \beta \rVert_2^2$ — the lasso and ridge penalties respectively.
- $\rho = 1$ recovers lasso; $\rho = 0$ recovers ridge (up to solver choice).
The penalty's unit ball interpolates between the L1 diamond and the L2 sphere: corners survive (hence exact zeros) but the faces bulge (hence sharing). The objective is strictly convex for $\rho < 1$, so the solution is unique even with exactly duplicated columns — where lasso's optimum is non-unique, the algebraic root of its coin-flip behaviour.
The grouping effect, quantified
Zou and Hastie (2005), Regularization and variable selection via the elastic net, prove that for standardised features with sample correlation $r_{jk}$:
$$ |\hat{\beta}_j - \hat{\beta}_k| \;\leq\; \frac{\lVert y \rVert_1}{n\,\alpha(1-\rho)} \sqrt{2(1 - r_{jk})} $$
As $r_{jk} \to 1$, coefficients of the twins are forced together — the grouping effect as a theorem rather than a tendency. The same paper introduces the rescaling $(1 + \alpha(1-\rho))\hat{\beta}$ (the "corrected" elastic net) to undo part of the double shrinkage from stacking two penalties.
Computation
Coordinate descent handles the combined penalty with a modified soft-threshold update:
$$ \beta_j \leftarrow \frac{\operatorname{sign}(z_j)\max(|z_j| - \alpha\rho,\, 0)}{1 + \alpha(1-\rho)} $$
with $z_j$ the partial residual correlation for coordinate $j$. Cost per sweep is $O(nd)$, with warm-started regularisation paths and active-set tricks as in glmnet (Friedman, Hastie, Tibshirani, 2010). Strong rules (Tibshirani et al., 2012) discard most features per path step before solving.
Selection consistency and practice
Lasso's support recovery needs the irrepresentable condition, which correlated designs violate; the elastic net relaxes this, and the adaptive elastic net (Zou and Zhang, 2009) attains oracle properties in high dimension. Empirically, $d \gg n$ fields with correlated blocks — genomics, neuroimaging, chemometrics, text with engineered n-gram families — are where elastic net beats both parents. When features are near-orthogonal, plain lasso typically matches it with one fewer hyperparameter; when nothing should be dropped, ridge does. The blend earns its second knob exactly when the correlation structure is real.
What to learn next
- Multicollinearity and VIF — measuring the twin problem before choosing a cure.
- Lasso regression — the selection machinery this lesson repaired.
- Feature engineering — avoiding accidental twins at the source.