AI glossary

Feature scaling

In one sentence Feature scaling puts all input columns on comparable ranges, so no feature dominates purely because its numbers are bigger.

By Updated

Feature scaling rescales each input column to a comparable range, so a feature cannot dominate the model because of its units.

Suppose you compare flats using price in rupees and distance to the metro in kilometres. One column holds numbers like 8,500,000; the other holds numbers like 1.2. Any method that measures distances between examples — or takes gradient steps across them — will be utterly dominated by the rupee column. Not because price matters more, but because its numbers are a million times larger. Scaling converts both to comparable units, the way you convert everything to percentages before comparing.

Two standard recipes:

standardisation :  (value − mean) ÷ standard deviation   → centred near 0, spread ≈ 1
min-max scaling :  (value − min) ÷ (max − min)           → squeezed into [0, 1]

Who needs it: KNN, SVM, k-means, PCA, linear and logistic regression with regularization, and neural networks — anything distance-based or gradient-trained. Who does not: decision-trees and their ensembles, which split on thresholds and never compare across columns.

The critical rule: fit the scaler on training data only, then apply those same numbers to the test set. Computing the mean over all data before splitting leaks test-set information into training — a textbook case of data-leakage.

Where to go next