Linear Models and Regularisation

Multicollinearity and VIF

When features tell the same story, coefficients turn unstable and unreadable — VIF is the number that catches it before you publish nonsense.

On this page 5
  1. Why it exists — or rather, why it hurts
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Multicollinearity means your features overlap so much that the model cannot tell who deserves the credit — and VIF is the number that measures it.

Two friends always visit your shop together, and sales rise every time they come. Which friend brings the luck? From your records alone, unanswerable — they have never come separately. Any answer you force out of the data is a guess dressed as a finding.

Features can be inseparable friends too. A flat's size and its room count arrive together in every row. Multicollinearity is this inseparability, and it quietly poisons a model's explanations.

Why it exists — or rather, why it hurts

The model's predictions usually survive: whichever friend gets credit, "the pair visited" still predicts sales. What breaks is the story. Coefficients swing wildly between datasets, signs flip, and a feature everyone knows matters can show up as "insignificant" because its twin absorbed the evidence. The ridge lesson shows exactly this chaos in action.

So you want a smoke detector: a number, per feature, saying how redundant it is. The variance inflation factor (VIF) is that number. The idea is disarmingly direct — try to predict each feature from the other features. Predictable feature, redundant feature. VIF repackages "how predictable" into "how many times more wobbly this coefficient is than it should be".

How it works

for each feature:
    hide it  →  predict it using the OTHER features
    predicted well?  →  it's redundant  →  VIF is high

VIF ≈ 1     independent — no wobble added
VIF ≈ 5     noticeably entangled
VIF > 10    coefficient wobble is 10x — stop trusting this number

A real example you have seen

Nutrition headlines. "Tea drinkers live longer" — but tea drinking travels with income, and income with healthcare. Entangled inputs, unstable credit. Statisticians running such studies check VIF-style diagnostics before naming a culprit; headlines usually do not.

Remember this

  • Multicollinearity rarely hurts predictions; it wrecks explanations.
  • VIF asks: can the other features predict this one? High answer, high redundancy.
  • Rules of thumb: VIF near 1 is clean, above ~10 means the coefficient is untrustworthy.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU.

VIF from first principles

VIF needs no special library — it is one regression per feature:

vif.py
import numpy as np
from sklearn.linear_model import LinearRegression

rng = np.random.default_rng(7)
size = rng.uniform(400, 1600, 30)               # sq ft
rooms = size / 350 + rng.normal(0, 0.1, 30)     # almost a copy of size
metro = rng.uniform(0.5, 8, 30)                 # km to metro, unrelated
X = np.column_stack([size, rooms, metro])
names = ["size", "rooms", "metro"]

for i, name in enumerate(names):
    others = np.delete(X, i, axis=1)
    r2 = LinearRegression().fit(others, X[:, i]).score(others, X[:, i])
    vif = 1 / (1 - r2)
    print(f"{name:>5}: R^2 against the others = {r2:.3f}  ->  VIF = {vif:7.1f}")
Output
 size: R^2 against the others = 0.992  ->  VIF =   131.2
rooms: R^2 against the others = 0.992  ->  VIF =   131.1
metro: R^2 against the others = 0.011  ->  VIF =     1.0

The walkthrough

The loop is the definition. For each feature, regress it on the rest and read the R² — the fraction of it the others already explain. rooms is 99.2% predictable from size and metro; it brings almost nothing of its own.

VIF = 1 / (1 - R²) converts that into a wobble multiplier. R² of 0.992 gives VIF 131: this coefficient's uncertainty is 131 times what it would be with independent features. Whatever the fitted value says about rooms, its error bars are enormous — sign flips between samples are expected, not surprising.

metro at VIF 1.0 is the healthy baseline. Independent features earn VIF near 1, and their coefficients mean what they appear to mean.

What to actually do about VIF 131 — one honest menu:

  • Drop or merge. Keep size, drop rooms — or engineer one combined feature. Cheapest fix, costs nothing if the pair is truly redundant.
  • Regularise. Ridge or elastic net stabilise the credit split without dropping anything.
  • Accept it. If you only need predictions and the deployment data resembles training data, high VIF is survivable. Never quote individual coefficients, though.

If you prefer a library, statsmodels ships the identical computation as variance_inflation_factor — remember it expects a constant column added.

Common mistakes

Screening with pairwise correlation instead of VIF. Three features can be pairwise mildly correlated yet jointly redundant — a sum of two others has moderate pairwise correlations but infinite VIF. VIF sees group redundancy; the correlation matrix does not.

The dummy-variable trap. One-hot encoding all categories and keeping an intercept makes the encoding sum to the constant column — perfect multicollinearity by construction, R² of exactly 1, VIF infinite. Drop one category (OneHotEncoder(drop="first")) for readable coefficients.

Dropping a variable your question is about. If the study's purpose is the effect of rooms, deleting it for VIF hygiene answers a different question. High VIF may honestly mean this dataset cannot answer your question — more data or a designed experiment can, deletion cannot.

Reading a twin's "insignificance" as irrelevance. Wide error bars from redundancy make real effects look insignificant. "High VIF and insignificant" means "cannot tell", not "does not matter".

Try it yourself

Add a fourth feature total = size + 50 * rooms and rerun. Predict metro's VIF before running; then explain the other three values you see. Then reduce the rooms noise from 0.1 to 0.01 and watch VIF race toward infinity.

What to learn next

Researcher — Mathematics and papers.

What is being inflated

For OLS with design matrix $X$ (standardised columns) and noise variance $\sigma^2$, the variance of coefficient $j$:

$$ \operatorname{Var}(\hat{\beta}_j) = \frac{\sigma^2}{(n-1)\operatorname{Var}(x_j)} \cdot \frac{1}{1 - R_j^2} $$

Where:

  • $R_j^2$ — the R² from regressing feature $j$ on all other features.
  • $\text{VIF}_j = 1 / (1 - R_j^2)$ — the second factor: the multiplier on the variance a feature would have if orthogonal to the rest.
  • $n$ — sample count.

The first factor is the irreducible part; VIF isolates exactly the collinearity contribution. Standard treatment: Belsley, Kuh and Welsch (1980), Regression Diagnostics.

The spectral picture

Collinearity lives in the eigenvalues $\lambda_1 \geq \dots \geq \lambda_d$ of $X^\top X$ (standardised). Near-zero $\lambda_d$ means some direction $v_d$ of feature space is almost unrepresented in the data, and coefficient error along it explodes as $\sigma^2 / \lambda_d$. The condition number $\kappa = \sqrt{\lambda_1 / \lambda_d}$ summarises this; Belsley's guideline flags $\kappa > 30$, and condition indices with variance-decomposition proportions localise which coefficients share each weak direction — sharper than VIF, which cannot say who is entangled with whom.

This is precisely the pathology ridge treats: adding $\lambda I$ lifts the tiny eigenvalues. Principal-component regression instead discards the weak directions outright.

Thresholds, honestly

The VIF > 10 rule traces to convention, not theory; O'Brien (2007), A caution regarding rules of thumb for variance inflation factors, shows large VIFs can be harmless when $n$ is large or $\sigma^2$ small, since total variance is what matters. Treat VIF as a diagnostic to investigate, never an automatic deletion criterion. Note also that VIF is a property of the design, not the response: it is computed from $X$ alone.

Perfect collinearity and modern practice

When $R_j^2 = 1$ exactly, $X^\top X$ is singular; OLS solutions form an affine subspace, and software returns one of them via the pseudoinverse — coefficients then depend on the solver, a trap for the unwary. Structural causes (dummy trap, derived sums, $d > n$) should be fixed structurally.

In causal-inference settings the same issue appears as lack of positivity/overlap; in deep learning, feature collinearity is largely ignored because nobody reads individual first-layer weights — interpretation, not optimisation, is what collinearity kills. For inference with correlated designs, the current toolset is regularisation plus post-selection inference (lasso debiasing; van de Geer et al., 2014) or explicit experimental design.

What to learn next