Preprocessing and Feature Selection

Feature scaling

Feature scaling squeezes every column onto a similar range so that no column drowns out the others by accident of its units.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Feature scaling squeezes every column of your data onto a similar range, so no column shouts over the others.

Imagine a singing contest with two judges. One scores out of 10, the other scores out of 100,000. You add the two scores to pick the winner. The second judge now decides everything, and not because their opinion is better — their numbers are bigger.

Columns in a dataset are those judges. Income might run from 20,000 to 95,000 rupees. Age runs from 24 to 52. Any model that adds or compares these columns will be deafened by income and will barely hear age.

Scaling fixes this the way you would fix the judges: convert everyone to the same scale first, then compare.

Why it exists

Many models measure distance — how far apart two customers are, treating each row as a point. Distance is dominated by whichever column has the biggest numbers. That dominance comes from the units you happened to measure in, not from importance.

Change income from rupees to lakhs and the same model behaves differently. A model whose answers depend on your choice of units is broken. Scaling removes the units.

How it works

The most common recipe is standardisation: for each column, subtract the average, then divide by the typical spread.

income (rupees):   21000 ..... 95000   ──┐  scale    -1.1 ..... +1.1
age (years):          24 ..... 52      ──┘  ─────→   -1.1 ..... +1.2

After scaling, "high income" and "old" are both written as numbers near +1. "Low income" and "young" both sit near -1. The columns finally speak at the same volume.

A real example you have seen

A maps app choosing your route weighs distance in kilometres against time in minutes against tolls in rupees. Three different units. Before combining them into one "best route" score, they must be pulled onto one scale. Otherwise the toll amount, being the biggest number, would pick every route on its own.

Remember this

  • Big numbers in a column do not mean the column is important. They mean you picked big units.
  • Scaling puts every column on a similar range, usually centred near zero.
  • Models that use distances or add columns together need it. Tree models mostly do not.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Outputs below verified with scikit-learn 1.7.2 and NumPy 1.26.

Watch a model ignore a column

The label here depends on age — the older customers bought health insurance. Income is noise. Watch what an unscaled distance model does with that.

scaling.py
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier

# monthly income in rupees, age in years -- two wildly different scales
X = np.array([[92000, 24], [23000, 52], [88000, 25], [21000, 49],
              [95000, 26], [24000, 51], [90000, 27], [25000, 50]])
y = np.array([0, 1, 0, 1, 0, 1, 0, 1])   # 1 = bought health insurance (the older half)

new_customer = np.array([[91000, 53]])   # high income AND older

knn = KNeighborsClassifier(n_neighbors=3).fit(X, y)
print("unscaled prediction:", knn.predict(new_customer)[0])

scaler = StandardScaler().fit(X)   # learns each column's mean and spread
knn2 = KNeighborsClassifier(n_neighbors=3).fit(scaler.transform(X), y)
print("scaled prediction:  ", knn2.predict(scaler.transform(new_customer))[0])

print(f"income: mean {scaler.mean_[0]:.0f}, std {scaler.scale_[0]:.0f}")
print(f"age:    mean {scaler.mean_[1]:.0f}, std {scaler.scale_[1]:.1f}")
print("new customer, scaled:", scaler.transform(new_customer).round(2))
Output
unscaled prediction: 0
scaled prediction:   1
income: mean 57250, std 34065
age:    mean 38, std 12.5
new customer, scaled: [[0.99 1.2 ]]

The walkthrough

The unscaled model got it wrong. The new customer is 53 — squarely in the group that buys insurance. But the model answered 0. Look at the printed spreads to see why. Income varies by around 34,065; age varies by around 12.5. In a distance calculation, an income gap of 3,000 rupees outweighs an age gap of 25 years. The three nearest neighbours were the three high-income young people. Age never got a vote.

StandardScaler learns, then applies. fit records each column's mean and spread. transform rewrites each value as "how many spreads away from the mean". After that, the new customer reads as [0.99, 1.2] — a bit rich, notably old — and both facts carry similar weight.

Fit on training data only. The scaler is part of the model. If you compute the mean using test rows, information about the test set leaks into training. The pattern is always: fit on train, transform on both.

Other scalers, one line each

ScalerWhat it doesReach for it when
StandardScalercentre on 0, divide by spreadthe default; works almost everywhere
MinMaxScalersquash into [0, 1]you need a bounded range, no big outliers
RobustScalercentre on median, divide by IQRoutliers are wrecking the other two

Common mistakes

Calling fit_transform on the test set. This re-learns the mean and spread from test data. Results look fine and are quietly wrong. Use transform alone on anything the model will be judged on. See train-test split for why this separation is sacred.

Scaling with outliers using MinMaxScaler. One billionaire in the income column squashes everyone else into a tiny sliver near 0. RobustScaler uses the median and the middle 50% spread, so one absurd row cannot distort the rest.

Scaling columns that are already 0 or 1. One-hot columns do not need it, and scaling them can make sparse data dense and huge. Scale the numeric columns, pass the rest through — ColumnTransformer exists for exactly this.

Scaling before a tree model and expecting a change. Decision trees, random forests and XGBoost split on thresholds. "income > 50,000" and "scaled_income > -0.2" are the same split. Scaling neither helps nor hurts them.

Try it yourself

Change the new customer to [[24000, 25]] — poor and young. Predict what both models will say before running. Then break the scaled model on purpose: multiply the income column by 1,000 after scaling and watch age get ignored again.

What to learn next

Researcher — Mathematics and papers.

The transforms

Standardisation maps each feature value x to

z = (x − μ) / σ

Where μ is the feature's training-set mean and σ its training-set standard deviation. Min-max scaling maps x to (x − x_min)/(x_max − x_min); robust scaling to (x − median)/IQR, with IQR the interquartile range (75th minus 25th percentile).

None of these change the shape of a feature's distribution — a skewed feature stays skewed. Reshaping distributions is the job of power transforms.

Why gradient descent cares

For linear models with quadratic loss, the Hessian is H = (1/n) XᵀX (plus regularisation). Wildly different feature scales inflate the condition number κ(H) = λ_max/λ_min — the ratio of the largest to smallest eigenvalue. Gradient descent's convergence rate degrades linearly in κ: the loss surface becomes a long narrow valley, and the step size that is safe along the steep axis crawls along the shallow one. Standardising brings feature variances to 1, which typically collapses κ by orders of magnitude. LeCun et al. (1998), Efficient BackProp, is the classic treatment, including the recommendation to decorrelate inputs when feasible.

Distance-based methods

For k-NN and RBF-kernel SVMs, the metric d(x, x′)² = Σ_j (x_j − x_j′)² weights feature j by its variance. Scaling is therefore not a numerical nicety but part of the model specification: choosing a scaler chooses the metric. There is no "correct" metric — standardisation encodes the assumption that each feature deserves equal a-priori influence.

Invariances worth knowing

  • Tree ensembles are invariant to any strictly monotone per-feature transform — scaling, log, rank. Their splits depend only on feature order.
  • L2 regularisation is not scale-invariant: shrinking a feature's scale inflates its coefficient and its penalty. Regularised linear models (ridge, lasso, logistic with penalty) effectively require standardisation for the penalty to mean the same thing per feature.
  • Batch normalisation (Ioffe and Szegedy, 2015) is this idea moved inside the network: re-standardise at every layer, with learned scale and shift.

What to learn next

What to learn next

These follow on from what you just read.

  • Preprocessing and Feature Selection

    Power transforms

    Power transforms like log and Yeo-Johnson reshape lopsided columns into balanced ones, so a few giant values stop dominating everything the model learns.

  • Preprocessing and Feature Selection

    Binning and discretisation

    Binning chops a continuous number like age or income into a handful of buckets, trading fine detail for robustness and rules a human can read.

  • Preprocessing and Feature Selection

    The hashing trick

    The hashing trick turns words or categories into column numbers with a fixed formula instead of a stored dictionary, so memory stays constant no matter how many new values appear.