scikit-learn Properly

Encoders and scalers, and the unknown-category trap

Encoders turn categories into numbers and scalers put numbers on one scale — and both must survive meeting a value they never saw in training.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An encoder turns words like "mumbai" into columns of 0s and 1s, and a scaler brings all numbers onto one comparable scale.

Think of a school attendance register. Each student gets their own column, and each day you tick exactly one box per row. That is one-hot encoding: one column per category, with a single 1 marking which one applies.

Scalers solve the other half. Comparing height in centimetres against weight in kilograms is like comparing distances in footsteps against distances in kilometres. A scaler converts every column to the same kind of unit first.

Why it exists

Models do arithmetic. The word "mumbai" cannot be multiplied by a weight, so categories must become numbers. And when one column ranges 0–1 while another ranges 0–100000, many models let the big column shout down the small one. Encoding and scaling fix both problems before the model ever looks.

How it works

city column        delhi  mumbai  pune
"mumbai"     →       0      1      0
"delhi"      →       1      0      0

new city:
"kolkata"    →       0      0      0    ← the trap: no column exists for it

The trap is that last row. The encoder built its columns from training data. Months later, in production, a customer from a city it never saw walks in. The encoder must either crash or shrug — and you have to choose which, in advance.

A real example you have seen

Every form with a dropdown that ends in "Other". The form's designer knew new answers would appear that the original list missed. "Other" is the graceful shrug, planned ahead of time. An encoder without that plan is a form that crashes on new answers.

Remember this

  • One-hot encoding: one column per category, a single 1 per row.
  • Scaling: all number columns brought to comparable ranges, learned from training data.
  • Decide before deployment what happens when an unseen category arrives.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Tested against scikit-learn 1.7.

The trap, sprung on purpose

unknown_category.py
import numpy as np
from sklearn.preprocessing import OneHotEncoder

train_cities = np.array([["mumbai"], ["delhi"], ["pune"], ["delhi"]])
new_cities = np.array([["mumbai"], ["kolkata"]])   # kolkata was never seen

strict = OneHotEncoder().fit(train_cities)
try:
    strict.transform(new_cities)
except ValueError as err:
    print("strict:", str(err).splitlines()[0])

relaxed = OneHotEncoder(handle_unknown="ignore", sparse_output=False)
relaxed.fit(train_cities)
print("columns:", relaxed.get_feature_names_out())
print(relaxed.transform(new_cities))
Output
strict: Found unknown categories ['kolkata'] in column 0 during transform
columns: ['x0_delhi' 'x0_mumbai' 'x0_pune']
[[0. 1. 0.]
 [0. 0. 0.]]

Three scalers, one contaminated column

scaler_comparison.py
import numpy as np
from sklearn.preprocessing import MinMaxScaler, RobustScaler, StandardScaler

salaries = np.array([[3.], [4.], [5.], [6.], [90.]])   # one founder in the data

print("standard:", StandardScaler().fit_transform(salaries).round(2).ravel())
print("minmax:  ", MinMaxScaler().fit_transform(salaries).round(2).ravel())
print("robust:  ", RobustScaler().fit_transform(salaries).round(2).ravel())
Output
standard: [-0.54 -0.51 -0.49 -0.46  2.  ]
minmax:   [0.   0.01 0.02 0.03 1.  ]
robust:   [-1.  -0.5  0.   0.5 42.5]

The walkthrough

The default is to crash, and that is a reasonable default. An unseen category can mean upstream data corruption, and silence would hide it. handle_unknown="ignore" opts into the shrug: kolkata became [0, 0, 0], an all-zero row meaning "none of the known cities". The model treats it as belonging to no city — a defensible neutral answer.

sparse_output=False is for your eyes. By default the encoder returns a sparse matrix — a memory-saving format that stores only the 1s. Models consume it happily; print does not. Leave the default on in real pipelines with many categories.

Reading the scaler shoot-out. StandardScaler subtracts the mean and divides by spread — but the founder's 90 dragged the mean so far that the four normal salaries got crushed into a 0.08-wide lump. MinMaxScaler is worse here: it maps min to 0 and max to 1, so the outlier owns the whole range and everyone else huddles below 0.03. RobustScaler uses the median and the middle-half spread, which the outlier cannot drag — normal salaries stay nicely spread, and 90 shows up as the absurdity it is (42.5). With outliers, robust scaling is the safe first choice.

Ordinal is for real orders only. OrdinalEncoder maps categories to 0, 1, 2, … in one column. Use it when the order is real — ["small", "medium", "large"]. On cities it invents nonsense arithmetic: pune − delhi = mumbai. Its unknown-category plan is spelled differently: OrdinalEncoder(handle_unknown="use_encoded_value", unknown_value=-1).

Common mistakes

No unknown-category plan before deployment. The strict crash above will happen at 2 a.m., in production, on your first customer from a new city. Choose handle_unknown="ignore" — or the strict default plus monitoring — while you can still think about it.

Fitting the encoder on the full dataset "to see all categories". This is leakage with a convincing excuse. If a category exists only in the test set, honest evaluation should meet it as an unknown — production will.

MinMax on anything with outliers. One extreme value flattens every normal value into a paper-thin band, as the output shows. MinMax shines when bounds are genuinely fixed, like pixel values 0–255.

Scaling tree models out of habit. Decision trees, random forests and XGBoost split on thresholds, so monotonic rescaling changes nothing. Scaling matters for linear models, SVMs, k-NN and neural networks.

Try it yourself

Give OneHotEncoder the arguments min_frequency=2, handle_unknown="infrequent_if_exist" and refit on the training cities. Print the new columns. Where do mumbai and pune go now, and what happens to kolkata? Explain why grouping rare categories into one bucket also softens the unknown-category trap.

What to learn next

Researcher — Mathematics and papers.

Encoding as basis choice

One-hot encoding maps a categorical variable with K levels to the indicator vectors e_1..e_K in R^K. With an intercept, the K indicators are perfectly collinear (they sum to the all-ones vector); the fix is either drop="first" (K−1 dummy coding) or ridge-style regularisation, which selects the minimum-norm solution among the equivalent ones. scikit-learn's penalised defaults are why drop=None is safe there while unpenalised OLS is not.

For high-cardinality K, alternatives trade variance for dimensionality:

  • Feature hashing (Weinberger et al., 2009, Feature hashing for large scale multitask learning): index by hash into d ≪ K buckets; collisions add noise bounded by the inner-product concentration results in the paper.
  • Target encoding (Micci-Barreca, 2001): replace level j with a shrunken estimate λ_j ȳ_j + (1−λ_j) ȳ, where λ_j = n_j/(n_j + m) grows with the level's count n_j. Must be cross-fitted; scikit-learn's TargetEncoder does this internally.
  • Grouped infrequent categories (min_frequency, since 1.1): a bias-variance compromise that also gives unknowns a home bucket.

Why scaling changes optimisation, not information

Standardisation is an affine reparameterisation, so the model class is unchanged — what changes is geometry. For losses minimised by gradient methods, convergence rate depends on the condition number κ of the Hessian; wildly different feature variances inflate κ, and standardising is a diagonal preconditioner. For penalised models the effect is statistical, not only numerical: the penalty ‖w‖² is not invariant to feature rescaling, so an unscaled feature with tiny units receives an effectively lighter penalty. Identical reasoning covers distance-based methods (k-NN, RBF kernels), where unscaled features dominate the metric outright.

Robust location/scale (median, IQR) has bounded influence function, hence RobustScaler's immunity in the demo: the sample mean has breakdown point 0, the median 1/2 (Huber, Robust Statistics, 1981).

The unknown-category event, quantified

Under a Dirichlet-multinomial view, the probability that a fresh sample contains an unseen level relates to the missing-mass problem; the Good-Turing estimator puts E[missing mass] ≈ N_1/n, the count of singleton levels over sample size (Good, 1953). A training column with many singleton categories is a column that will produce unknowns in production — worth computing before choosing the strict default.

Version note: sparse_output replaced the old sparse argument in 1.2; handle_unknown="infrequent_if_exist" requires 1.1+. Code here assumes 1.7.

What to learn next