Preprocessing and Feature Selection
Feature scaling
Feature scaling squeezes every column onto a similar range so that no column drowns out the others by accident of its units.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Feature scaling squeezes every column of your data onto a similar range, so no column shouts over the others.
Imagine a singing contest with two judges. One scores out of 10, the other scores out of 100,000. You add the two scores to pick the winner. The second judge now decides everything, and not because their opinion is better — their numbers are bigger.
Columns in a dataset are those judges. Income might run from 20,000 to 95,000 rupees. Age runs from 24 to 52. Any model that adds or compares these columns will be deafened by income and will barely hear age.
Scaling fixes this the way you would fix the judges: convert everyone to the same scale first, then compare.
Why it exists
Many models measure distance — how far apart two customers are, treating each row as a point. Distance is dominated by whichever column has the biggest numbers. That dominance comes from the units you happened to measure in, not from importance.
Change income from rupees to lakhs and the same model behaves differently. A model whose answers depend on your choice of units is broken. Scaling removes the units.
How it works
The most common recipe is standardisation: for each column, subtract the average, then divide by the typical spread.
income (rupees): 21000 ..... 95000 ──┐ scale -1.1 ..... +1.1
age (years): 24 ..... 52 ──┘ ─────→ -1.1 ..... +1.2After scaling, "high income" and "old" are both written as numbers near +1. "Low income" and "young" both sit near -1. The columns finally speak at the same volume.
A real example you have seen
A maps app choosing your route weighs distance in kilometres against time in minutes against tolls in rupees. Three different units. Before combining them into one "best route" score, they must be pulled onto one scale. Otherwise the toll amount, being the biggest number, would pick every route on its own.
Remember this
- Big numbers in a column do not mean the column is important. They mean you picked big units.
- Scaling puts every column on a similar range, usually centred near zero.
- Models that use distances or add columns together need it. Tree models mostly do not.
What to learn next
- Power transforms — when a column's problem is its shape, not its size.
- Feature engineering — building the columns worth scaling.
- Train-test split — why the scaler must never see the test set.
Developer — Code and libraries.
Setup
pip install scikit-learn numpyOutputs below verified with scikit-learn 1.7.2 and NumPy 1.26.
Watch a model ignore a column
The label here depends on age — the older customers bought health insurance. Income is noise. Watch what an unscaled distance model does with that.
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier
# monthly income in rupees, age in years -- two wildly different scales
X = np.array([[92000, 24], [23000, 52], [88000, 25], [21000, 49],
[95000, 26], [24000, 51], [90000, 27], [25000, 50]])
y = np.array([0, 1, 0, 1, 0, 1, 0, 1]) # 1 = bought health insurance (the older half)
new_customer = np.array([[91000, 53]]) # high income AND older
knn = KNeighborsClassifier(n_neighbors=3).fit(X, y)
print("unscaled prediction:", knn.predict(new_customer)[0])
scaler = StandardScaler().fit(X) # learns each column's mean and spread
knn2 = KNeighborsClassifier(n_neighbors=3).fit(scaler.transform(X), y)
print("scaled prediction: ", knn2.predict(scaler.transform(new_customer))[0])
print(f"income: mean {scaler.mean_[0]:.0f}, std {scaler.scale_[0]:.0f}")
print(f"age: mean {scaler.mean_[1]:.0f}, std {scaler.scale_[1]:.1f}")
print("new customer, scaled:", scaler.transform(new_customer).round(2))unscaled prediction: 0 scaled prediction: 1 income: mean 57250, std 34065 age: mean 38, std 12.5 new customer, scaled: [[0.99 1.2 ]]
The walkthrough
The unscaled model got it wrong. The new customer is 53 — squarely in the group that buys insurance. But the model answered 0. Look at the printed spreads to see why. Income varies by around 34,065; age varies by around 12.5. In a distance calculation, an income gap of 3,000 rupees outweighs an age gap of 25 years. The three nearest neighbours were the three high-income young people. Age never got a vote.
StandardScaler learns, then applies. fit records each column's mean and spread. transform rewrites each value as "how many spreads away from the mean". After that, the new customer reads as [0.99, 1.2] — a bit rich, notably old — and both facts carry similar weight.
Fit on training data only. The scaler is part of the model. If you compute the mean using test rows, information about the test set leaks into training. The pattern is always: fit on train, transform on both.
Other scalers, one line each
| Scaler | What it does | Reach for it when |
|---|---|---|
StandardScaler | centre on 0, divide by spread | the default; works almost everywhere |
MinMaxScaler | squash into [0, 1] | you need a bounded range, no big outliers |
RobustScaler | centre on median, divide by IQR | outliers are wrecking the other two |
Common mistakes
Calling fit_transform on the test set. This re-learns the mean and spread from test data. Results look fine and are quietly wrong. Use transform alone on anything the model will be judged on. See train-test split for why this separation is sacred.
Scaling with outliers using MinMaxScaler. One billionaire in the income column squashes everyone else into a tiny sliver near 0. RobustScaler uses the median and the middle 50% spread, so one absurd row cannot distort the rest.
Scaling columns that are already 0 or 1. One-hot columns do not need it, and scaling them can make sparse data dense and huge. Scale the numeric columns, pass the rest through — ColumnTransformer exists for exactly this.
Scaling before a tree model and expecting a change. Decision trees, random forests and XGBoost split on thresholds. "income > 50,000" and "scaled_income > -0.2" are the same split. Scaling neither helps nor hurts them.
Try it yourself
Change the new customer to [[24000, 25]] — poor and young. Predict what both models will say before running. Then break the scaled model on purpose: multiply the income column by 1,000 after scaling and watch age get ignored again.
What to learn next
- Power transforms — when a column's problem is its shape, not its size.
- Feature engineering — building the columns worth scaling.
- Train-test split — why the scaler must never see the test set.
Researcher — Mathematics and papers.
The transforms
Standardisation maps each feature value x to
z = (x − μ) / σ
Where μ is the feature's training-set mean and σ its training-set standard deviation. Min-max scaling maps x to (x − x_min)/(x_max − x_min); robust scaling to (x − median)/IQR, with IQR the interquartile range (75th minus 25th percentile).
None of these change the shape of a feature's distribution — a skewed feature stays skewed. Reshaping distributions is the job of power transforms.
Why gradient descent cares
For linear models with quadratic loss, the Hessian is H = (1/n) XᵀX (plus regularisation). Wildly different feature scales inflate the condition number κ(H) = λ_max/λ_min — the ratio of the largest to smallest eigenvalue. Gradient descent's convergence rate degrades linearly in κ: the loss surface becomes a long narrow valley, and the step size that is safe along the steep axis crawls along the shallow one. Standardising brings feature variances to 1, which typically collapses κ by orders of magnitude. LeCun et al. (1998), Efficient BackProp, is the classic treatment, including the recommendation to decorrelate inputs when feasible.
Distance-based methods
For k-NN and RBF-kernel SVMs, the metric d(x, x′)² = Σ_j (x_j − x_j′)² weights feature j by its variance. Scaling is therefore not a numerical nicety but part of the model specification: choosing a scaler chooses the metric. There is no "correct" metric — standardisation encodes the assumption that each feature deserves equal a-priori influence.
Invariances worth knowing
- Tree ensembles are invariant to any strictly monotone per-feature transform — scaling, log, rank. Their splits depend only on feature order.
- L2 regularisation is not scale-invariant: shrinking a feature's scale inflates its coefficient and its penalty. Regularised linear models (ridge, lasso, logistic with penalty) effectively require standardisation for the penalty to mean the same thing per feature.
- Batch normalisation (Ioffe and Szegedy, 2015) is this idea moved inside the network: re-standardise at every layer, with learned scale and shift.
What to learn next
- Power transforms — when a column's problem is its shape, not its size.
- Feature engineering — building the columns worth scaling.
- Train-test split — why the scaler must never see the test set.