Machine Learning

Linear regression

Linear regression draws the straight line that fits your data best, and uses it to predict a number for inputs it has never seen.

Read these first

On this page 7
  1. Why it exists
  2. What the two numbers mean
  3. How it works
  4. Where you have already seen it
  5. The honest part
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Linear regression finds the straight line that best fits your data, then uses that line to predict a number.

Think about an auto-rickshaw meter. The moment you sit down, the meter already shows a starting amount. Then every kilometre adds roughly the same charge on top. If you know both numbers, you can guess the fare for any trip before you take it.

That is linear regression. A starting amount, plus a steady amount for each unit of something. Two numbers, and you can predict.

Why it exists

Plenty of real relationships work like that meter.

A bigger flat costs more. More rainfall gives more crop. More hours of study give more marks. In each case, going up on one side tends to push the other side up by a fairly steady amount.

You could try to guess these by eye. Draw the dots on graph paper, hold a ruler against them, and slide it around until it looks balanced. People genuinely did this for centuries.

The problem with the ruler is that two people place it differently. Linear regression replaces the eye with a definite rule, so everyone gets the same line from the same dots.

What the two numbers mean

Every straight line is described by two numbers, and both have a plain meaning.

  • The starting value. What the meter reads before you travel any distance. In flat prices, it is roughly what an empty plot would cost.
  • The step. How much the answer rises for each extra unit of input. One more kilometre adds this much fare. One hundred more square feet adds this much price.

The step is the interesting one. It tells you the strength of the relationship, in units you can talk about with a non-technical person.

How it works

STEP 1 — plot what you know

  price
    ^
    |                          o
    |                    o
    |               o
    |         o
    |     o
    |  o
    +-------------------------> area


STEP 2 — try a line, measure every miss

  price
    ^                        /
    |                      /  o     each | is a miss
    |                 o  /
    |               / |
    |         o  /
    |     o  /
    |  o  /
    +-------------------------> area


STEP 3 — keep the line where the total miss is smallest

  price
    ^                       o
    |                    x
    |               o        x = the line's guess
    |          x             o = what really happened
    |     o
    |  x
    +-------------------------> area

The computer does not slide the line around by trial and error. There is a formula that lands on the best line directly, in one shot. That is unusual and lovely — most machine learning has no such shortcut.

Where you have already seen it

  • Property websites showing an estimated price for a flat.
  • Electricity bill predictions in your provider's app.
  • Fitness apps estimating calories burned from distance covered.
  • Delivery apps estimating fare from distance, which is the auto meter again.
  • Salary calculators that estimate pay from years of experience.

The honest part

Linear regression has a real weakness, and it is worth knowing on day one.

It insists on a straight line, even when reality bends. Study hours help your marks — up to a point. Study twenty hours a day and your marks fall, because you are exhausted. A straight line cannot express that. It will predict marks above one hundred and never notice the absurdity.

It cannot see outside the data it was shown. Say every flat in your data is between six hundred and fourteen hundred square feet. The line then knows nothing about a five thousand square foot bungalow. It will still give you a confident number for one. That number is a guess dressed up as an answer.

A pattern is not a cause. Ice cream sales and drowning deaths rise together. Ice cream does not cause drowning — hot weather causes both. Linear regression finds the pattern and stays silent about why. Deciding the why is your job, not the model's.

Remember this

  • Linear regression fits a straight line described by a starting value and a steady step.
  • The step tells you how much the answer moves for each extra unit of input.
  • It fails on bending relationships, and it should not be trusted outside the range of data it saw.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Fitting a line

flat_prices.py
import numpy as np
from sklearn.linear_model import LinearRegression

# Input: flat area in hundreds of square feet
X = np.array([[6.5], [8.0], [9.5], [11.0], [12.5], [14.0]])
# Label: asking price in lakhs of rupees
y = np.array([42.0, 50.0, 61.0, 68.0, 79.0, 87.0])

model = LinearRegression().fit(X, y)

print("slope     :", round(float(model.coef_[0]), 3))
print("intercept :", round(float(model.intercept_), 3))
print("R squared :", round(model.score(X, y), 4))
print("price predicted for 1000 sq ft:", round(float(model.predict([[10.0]])[0]), 2))

# The gap between truth and prediction, one row at a time
residuals = y - model.predict(X)
print("residuals :", np.round(residuals, 2))
Output
slope     : 6.076
intercept : 2.219
R squared : 0.9974
price predicted for 1000 sq ft: 62.98
residuals : [ 0.29 -0.83  1.06 -1.06  0.83 -0.29]

Reading the output

slope is 6.076. Each extra hundred square feet adds about 6.08 lakh to the price. This is the number to quote in a meeting. It means something to a person who has never heard of regression.

intercept is 2.219. This is where the line crosses zero area. A zero-square-foot flat costing 2.22 lakh is nonsense, and that is fine. The intercept is a positioning number for the line, not a real-world price. Treating it as one is a common misreading.

R squared is 0.9974. This is the fraction of the price variation that the line accounts for. On that scale, 1.0 is a perfect fit. A value this high should make you suspicious rather than pleased, because real data is messier. It is high here because these six rows were built to sit near a line.

residuals are the misses, in lakhs. They stay under 1.1, which is small against prices near 60.

Look at their signs though: plus, minus, plus, minus, plus, minus. Residuals from a well-specified model should look like random scatter, with no readable pattern. A perfect alternation is a pattern. Here it is an artefact of six hand-picked rows, but the habit it should teach you is real.

On actual data, always plot residuals against the fitted values. Random cloud means the straight line captured what was there. A smooth arc means the relationship curves. A widening fan means the error grows with the prediction, which breaks the assumptions behind any confidence interval you report.

The same answer without scikit-learn

The formula is a single least-squares solve. Seeing it removes the mystery.

by_hand.py
import numpy as np

area = np.array([6.5, 8.0, 9.5, 11.0, 12.5, 14.0])
price = np.array([42.0, 50.0, 61.0, 68.0, 79.0, 87.0])

# A column of ones lets the same equation solve for the intercept too
A = np.c_[np.ones(len(area)), area]
intercept, slope = np.linalg.lstsq(A, price, rcond=None)[0]

print("intercept:", round(intercept, 3))
print("slope    :", round(slope, 3))
Output
intercept: 2.219
slope    : 6.076

Identical to four decimal places, because it is the identical computation. LinearRegression is a convenience wrapper around scipy.linalg.lstsq. There is no hidden cleverness inside it.

The column of ones is the one non-obvious line. Without it, the line is forced through the origin. Adding a constant column lets the same matrix equation solve for the intercept alongside the slope.

Common mistakes

Predicting far outside the training range. model.predict([[50.0]]) returns about 306, a confident price for a 5000 sq ft flat the model never saw. The model has no mechanism to refuse. Check your input range before trusting an output.

Reporting score() on training data. model.score(X, y) with the same X used in fit() measures memorisation. Split your data first. See train, test and validation splits.

Ignoring outliers. Squared error punishes large misses heavily, so one wrongly-typed price can visibly tilt the whole line. Plot the data before fitting. If outliers are genuine, HuberRegressor downweights them.

Fitting a line to a curve. If residuals form a clear arc rather than scattering randomly, a straight line is the wrong shape. Add polynomial features, and then read overfitting and underfitting before you get carried away.

Comparing coefficients across differently-scaled features. A coefficient on "area in square feet" and one on "number of bedrooms" are not comparable. Standardise features before ranking them by importance.

Try it yourself

Add one badly mistyped row to flat_prices.py. Give area 15.0 a price of 20.0, instead of the roughly 93 it should be:

Add these two lines to flat_prices.py, right after y is defined and before the fit() call:

flat_prices.py (edit)
X = np.vstack([X, [[15.0]]])
y = np.append(y, 20.0)

Rerun it and read the damage:

Output
slope     : 0.989
intercept : 47.334
R squared : 0.0182

The slope fell from 6.076 to 0.989. R squared fell from 0.9974 to 0.0182, meaning the line now explains almost nothing.

One wrong row out of seven destroyed the model. Squared error punishes a large miss enormously, so a single extreme point can outvote every correct one. This is the fastest way to feel why experienced practitioners spend more time cleaning data than choosing algorithms.

What to learn next

Researcher — Mathematics and papers.

Ordinary least squares

The model is linear in its parameters, which is what makes the closed form possible.

y = X * beta + epsilon
  • y — response vector, shape (n, 1)
  • X — design matrix, shape (n, p), first column all ones for the intercept
  • beta — parameter vector, shape (p, 1)
  • epsilon — error term, shape (n, 1), assumed mean zero
  • n — number of observations, p — number of parameters including the intercept

OLS minimises the residual sum of squares:

RSS(beta) = || y - X*beta ||^2 = (y - X*beta)^T (y - X*beta)

Setting the gradient to zero yields the normal equations:

X^T X beta_hat = X^T y

and, when X^T X is invertible:

beta_hat = (X^T X)^{-1} X^T y
  • X^T — transpose of the design matrix
  • (X^T X)^{-1} — inverse of the p x p Gram matrix
  • beta_hat — the estimated parameters

Geometrically, X * beta_hat is the orthogonal projection of y onto the column space of X. The residual vector is orthogonal to every column of X, which is the entire content of the normal equations.

Do not invert the matrix

(X^T X)^{-1} is correct algebra and poor numerics. Forming X^T X squares the condition number of X:

kappa(X^T X) = kappa(X)^2
  • kappa — condition number, the ratio of largest to smallest singular value

Production solvers use QR decomposition or SVD on X directly. scikit-learn's LinearRegression calls scipy.linalg.lstsq, which uses the SVD-based gelsd driver. When X is rank-deficient it returns the minimum-norm solution rather than raising an error.

Normal equations:  O(n p^2 + p^3)
Householder QR:    O(n p^2 - p^3/3)
SVD:               O(n p^2) with a larger constant, most robust
Gradient descent:  O(n p) per iteration, preferred when n or p is very large

Gauss-Markov

The theorem requires four assumptions:

  1. Linearity in the parameters.
  2. Strict exogeneity, E[epsilon | X] = 0.
  3. Homoscedasticity and no autocorrelation, Var(epsilon | X) = sigma^2 * I.
  4. Full column rank of X.

Under these, beta_hat is the Best Linear Unbiased Estimator. It has minimum variance among all linear unbiased estimators.

Note what is not assumed: normality of errors. Normality is needed for exact t and F inference in finite samples, not for the Gauss-Markov result itself.

The variance of the estimator is:

Var(beta_hat) = sigma^2 * (X^T X)^{-1}
  • sigma^2 — error variance, estimated by RSS / (n - p)

Note that "best unbiased" is a weaker claim than "best". Biased estimators can achieve lower mean squared error, which is exactly the opening that regularisation exploits.

Coefficient of determination

R^2 = 1 - RSS / TSS,     TSS = SUM_i ( y_i - y_bar )^2
  • RSS — residual sum of squares
  • TSS — total sum of squares
  • y_bar — mean of the observed response

R^2 never decreases when a predictor is added, including a column of pure noise. Adjusted R^2 penalises p:

R^2_adj = 1 - ( (1 - R^2) * (n - 1) / (n - p) )

On out-of-sample data R^2 can be negative, meaning the model is worse than predicting the training mean. That is a valid and informative result, not a bug.

Regularisation

When p approaches or exceeds n, or predictors are collinear, OLS variance explodes. Adding a penalty trades bias for variance.

Ridge:  beta_hat = argmin  || y - X*beta ||^2 + lambda * || beta ||_2^2
Lasso:  beta_hat = argmin  || y - X*beta ||^2 + lambda * || beta ||_1
  • lambda — regularisation strength, chosen by cross-validation
  • || . ||_2^2 — squared L2 norm, the sum of squared coefficients
  • || . ||_1 — L1 norm, the sum of absolute coefficients

Ridge has the closed form beta_hat = (X^T X + lambda*I)^{-1} X^T y. That matrix is always invertible for lambda > 0.

Lasso has no closed form. Its L1 penalty is non-differentiable at zero, and that is exactly why it drives coefficients to exact zero. The result is genuine variable selection. It is solved by coordinate descent or LARS.

Always standardise features before penalising. The penalty is not scale-invariant, so unstandardised features are penalised in proportion to their arbitrary units.

Diagnostics worth running

CheckDetectsTool
Residuals vs fittedNon-linearity, heteroscedasticityScatter plot
Q-Q plot of residualsNon-normal errorsscipy.stats.probplot
Variance Inflation FactorMulticollinearity; VIF > 10 is a warningstatsmodels
Cook's distanceHigh-influence observationsstatsmodels
Durbin-WatsonAutocorrelated errors in time seriesstatsmodels

For inference — standard errors, p-values, confidence intervals — use statsmodels rather than scikit-learn. scikit-learn is built for prediction and deliberately omits inferential output.

Key references

  • Legendre, A. M. (1805). Nouvelles méthodes pour la détermination des orbites des comètes. First publication of least squares.
  • Gauss, C. F. (1809). Theoria Motus Corporum Coelestium. Claims prior use from 1795 and adds the probabilistic justification.
  • Hoerl, A. & Kennard, R. (1970). Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics 12(1).
  • Tibshirani, R. (1996). Regression Shrinkage and Selection via the Lasso. JRSS-B 58(1).
  • Zou, H. & Hastie, T. (2005). Regularization and Variable Selection via the Elastic Net. JRSS-B 67(2).
  • Golub, G. & Van Loan, C. (2013). Matrix Computations, 4th ed. Johns Hopkins. The reference for the numerical arguments above.

Current state

Linear regression is not superseded, and treating it as a beginner's toy is a mistake. It remains the default in econometrics and epidemiology. It wins wherever a defensible coefficient matters more than a marginal accuracy gain.

Generalised linear models (Nelder & Wedderburn, 1972) extend it to non-Gaussian responses through a link function. That is where logistic and Poisson regression come from. Generalised additive models (Hastie & Tibshirani, 1986) relax linearity per feature while keeping additive interpretability.

It is also the correct baseline for any tabular regression problem. Report it alongside your gradient-boosted model. If the gap is small, the extra complexity is not earning its keep.

What to learn next