Linear Models and Regularisation
Multicollinearity and VIF
When features tell the same story, coefficients turn unstable and unreadable — VIF is the number that catches it before you publish nonsense.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Multicollinearity means your features overlap so much that the model cannot tell who deserves the credit — and VIF is the number that measures it.
Two friends always visit your shop together, and sales rise every time they come. Which friend brings the luck? From your records alone, unanswerable — they have never come separately. Any answer you force out of the data is a guess dressed as a finding.
Features can be inseparable friends too. A flat's size and its room count arrive together in every row. Multicollinearity is this inseparability, and it quietly poisons a model's explanations.
Why it exists — or rather, why it hurts
The model's predictions usually survive: whichever friend gets credit, "the pair visited" still predicts sales. What breaks is the story. Coefficients swing wildly between datasets, signs flip, and a feature everyone knows matters can show up as "insignificant" because its twin absorbed the evidence. The ridge lesson shows exactly this chaos in action.
So you want a smoke detector: a number, per feature, saying how redundant it is. The variance inflation factor (VIF) is that number. The idea is disarmingly direct — try to predict each feature from the other features. Predictable feature, redundant feature. VIF repackages "how predictable" into "how many times more wobbly this coefficient is than it should be".
How it works
for each feature:
hide it → predict it using the OTHER features
predicted well? → it's redundant → VIF is high
VIF ≈ 1 independent — no wobble added
VIF ≈ 5 noticeably entangled
VIF > 10 coefficient wobble is 10x — stop trusting this numberA real example you have seen
Nutrition headlines. "Tea drinkers live longer" — but tea drinking travels with income, and income with healthcare. Entangled inputs, unstable credit. Statisticians running such studies check VIF-style diagnostics before naming a culprit; headlines usually do not.
Remember this
- Multicollinearity rarely hurts predictions; it wrecks explanations.
- VIF asks: can the other features predict this one? High answer, high redundancy.
- Rules of thumb: VIF near 1 is clean, above ~10 means the coefficient is untrustworthy.
What to learn next
- Regression diagnostics — the rest of the pre-flight checklist for linear models.
- Ridge regression — the standard treatment for the disease measured here.
- Elastic net — selection that respects correlated groups.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU.
VIF from first principles
VIF needs no special library — it is one regression per feature:
import numpy as np
from sklearn.linear_model import LinearRegression
rng = np.random.default_rng(7)
size = rng.uniform(400, 1600, 30) # sq ft
rooms = size / 350 + rng.normal(0, 0.1, 30) # almost a copy of size
metro = rng.uniform(0.5, 8, 30) # km to metro, unrelated
X = np.column_stack([size, rooms, metro])
names = ["size", "rooms", "metro"]
for i, name in enumerate(names):
others = np.delete(X, i, axis=1)
r2 = LinearRegression().fit(others, X[:, i]).score(others, X[:, i])
vif = 1 / (1 - r2)
print(f"{name:>5}: R^2 against the others = {r2:.3f} -> VIF = {vif:7.1f}")size: R^2 against the others = 0.992 -> VIF = 131.2 rooms: R^2 against the others = 0.992 -> VIF = 131.1 metro: R^2 against the others = 0.011 -> VIF = 1.0
The walkthrough
The loop is the definition. For each feature, regress it on the rest and read the R² — the fraction of it the others already explain. rooms is 99.2% predictable from size and metro; it brings almost nothing of its own.
VIF = 1 / (1 - R²) converts that into a wobble multiplier. R² of 0.992 gives VIF 131: this coefficient's uncertainty is 131 times what it would be with independent features. Whatever the fitted value says about rooms, its error bars are enormous — sign flips between samples are expected, not surprising.
metro at VIF 1.0 is the healthy baseline. Independent features earn VIF near 1, and their coefficients mean what they appear to mean.
What to actually do about VIF 131 — one honest menu:
- Drop or merge. Keep
size, droprooms— or engineer one combined feature. Cheapest fix, costs nothing if the pair is truly redundant. - Regularise. Ridge or elastic net stabilise the credit split without dropping anything.
- Accept it. If you only need predictions and the deployment data resembles training data, high VIF is survivable. Never quote individual coefficients, though.
If you prefer a library, statsmodels ships the identical computation as variance_inflation_factor — remember it expects a constant column added.
Common mistakes
Screening with pairwise correlation instead of VIF. Three features can be pairwise mildly correlated yet jointly redundant — a sum of two others has moderate pairwise correlations but infinite VIF. VIF sees group redundancy; the correlation matrix does not.
The dummy-variable trap. One-hot encoding all categories and keeping an intercept makes the encoding sum to the constant column — perfect multicollinearity by construction, R² of exactly 1, VIF infinite. Drop one category (OneHotEncoder(drop="first")) for readable coefficients.
Dropping a variable your question is about. If the study's purpose is the effect of rooms, deleting it for VIF hygiene answers a different question. High VIF may honestly mean this dataset cannot answer your question — more data or a designed experiment can, deletion cannot.
Reading a twin's "insignificance" as irrelevance. Wide error bars from redundancy make real effects look insignificant. "High VIF and insignificant" means "cannot tell", not "does not matter".
Try it yourself
Add a fourth feature total = size + 50 * rooms and rerun. Predict metro's VIF before running; then explain the other three values you see. Then reduce the rooms noise from 0.1 to 0.01 and watch VIF race toward infinity.
What to learn next
- Regression diagnostics — the rest of the pre-flight checklist for linear models.
- Ridge regression — the standard treatment for the disease measured here.
- Elastic net — selection that respects correlated groups.
Researcher — Mathematics and papers.
What is being inflated
For OLS with design matrix $X$ (standardised columns) and noise variance $\sigma^2$, the variance of coefficient $j$:
$$ \operatorname{Var}(\hat{\beta}_j) = \frac{\sigma^2}{(n-1)\operatorname{Var}(x_j)} \cdot \frac{1}{1 - R_j^2} $$
Where:
- $R_j^2$ — the R² from regressing feature $j$ on all other features.
- $\text{VIF}_j = 1 / (1 - R_j^2)$ — the second factor: the multiplier on the variance a feature would have if orthogonal to the rest.
- $n$ — sample count.
The first factor is the irreducible part; VIF isolates exactly the collinearity contribution. Standard treatment: Belsley, Kuh and Welsch (1980), Regression Diagnostics.
The spectral picture
Collinearity lives in the eigenvalues $\lambda_1 \geq \dots \geq \lambda_d$ of $X^\top X$ (standardised). Near-zero $\lambda_d$ means some direction $v_d$ of feature space is almost unrepresented in the data, and coefficient error along it explodes as $\sigma^2 / \lambda_d$. The condition number $\kappa = \sqrt{\lambda_1 / \lambda_d}$ summarises this; Belsley's guideline flags $\kappa > 30$, and condition indices with variance-decomposition proportions localise which coefficients share each weak direction — sharper than VIF, which cannot say who is entangled with whom.
This is precisely the pathology ridge treats: adding $\lambda I$ lifts the tiny eigenvalues. Principal-component regression instead discards the weak directions outright.
Thresholds, honestly
The VIF > 10 rule traces to convention, not theory; O'Brien (2007), A caution regarding rules of thumb for variance inflation factors, shows large VIFs can be harmless when $n$ is large or $\sigma^2$ small, since total variance is what matters. Treat VIF as a diagnostic to investigate, never an automatic deletion criterion. Note also that VIF is a property of the design, not the response: it is computed from $X$ alone.
Perfect collinearity and modern practice
When $R_j^2 = 1$ exactly, $X^\top X$ is singular; OLS solutions form an affine subspace, and software returns one of them via the pseudoinverse — coefficients then depend on the solver, a trap for the unwary. Structural causes (dummy trap, derived sums, $d > n$) should be fixed structurally.
In causal-inference settings the same issue appears as lack of positivity/overlap; in deep learning, feature collinearity is largely ignored because nobody reads individual first-layer weights — interpretation, not optimisation, is what collinearity kills. For inference with correlated designs, the current toolset is regularisation plus post-selection inference (lasso debiasing; van de Geer et al., 2014) or explicit experimental design.
What to learn next
- Regression diagnostics — the rest of the pre-flight checklist for linear models.
- Ridge regression — the standard treatment for the disease measured here.
- Elastic net — selection that respects correlated groups.