Preprocessing and Feature Selection

Power transforms

Power transforms like log and Yeo-Johnson reshape lopsided columns into balanced ones, so a few giant values stop dominating everything the model learns.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A power transform reshapes a lopsided column — many small values, a few huge ones — into a balanced one.

Picture drawing every family's income in your city as a bar on a wall. Most bars are knee-high. Then one industrialist's bar smashes through the ceiling and keeps going. To fit that one bar, you shrink the chart until every normal family becomes an unreadable smudge.

This lopsidedness has a name: skew. A skewed column is one where values pile up at one end while a long thin tail stretches out the other way. Money, city populations, website visits, house prices — most real "amount" data looks like this.

Why it exists

Feature scaling changes a column's size but not its shape. A skewed column stays skewed after scaling. The industrialist is still a thousand times the median. A model fitting straight lines or measuring distances still bends everything around that one tail.

The fix is to change what the number means. Instead of asking "how many rupees?", ask "roughly how many times bigger?". On that scale, 1,000 to 10,000 is one step, and 10,000 to 100,000 is one more step. The tail folds in. This is what a logarithm does — and a power transform is a family of such squashing rules, with a knob for how hard to squash.

How it works

raw:      2,000   20,000   36,000   90,000   460,000    <- tail runs far right
             |        |        |        |         |
             v        v        v        v         v     squash big values hardest
reshaped:  7.6      9.9     10.5     11.4      13.0     <- roughly balanced

Small values barely move. Huge values get pulled in hard. Afterwards the column looks like a fair bell shape, which is the shape most models silently hope for.

A real example you have seen

Earthquake strength is reported on the Richter scale — a magnitude 7 is ten times a 6. Sound is measured in decibels. Both are logarithms in disguise, invented for the same reason. The raw numbers span such a huge range that plain differences stop meaning anything. Only a "times bigger" scale makes them comparable.

Remember this

  • Skew means values bunch at one end with a long tail at the other.
  • Scaling changes size; power transforms change shape.
  • Log and its relatives squash big values hard and small values gently.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn scipy numpy

Outputs verified with scikit-learn 1.7.2, SciPy 1.14, NumPy 1.26.

Squashing a thousand fake incomes

power.py
import numpy as np
from scipy.stats import skew
from sklearn.preprocessing import PowerTransformer

rng = np.random.default_rng(42)
# 1,000 fake monthly incomes in rupees -- most small, a few enormous
incomes = rng.lognormal(mean=10.5, sigma=0.8, size=1000)

print(f"min {incomes.min():.0f}   median {np.median(incomes):.0f}   max {incomes.max():.0f}")
print(f"skew of raw incomes: {skew(incomes):.2f}")

logged = np.log1p(incomes)                     # log(1 + x): safe even at zero
print(f"skew after log1p:    {skew(logged):.2f}")

pt = PowerTransformer()                        # Yeo-Johnson by default
squashed = pt.fit_transform(incomes.reshape(-1, 1))
print(f"skew after Yeo-Johnson: {skew(squashed.ravel()):.2f}")
print(f"fitted lambda: {pt.lambdas_[0]:.3f}")  # near 0 means 'basically a log'
Output
min 1961   median 36495   max 461890
skew of raw incomes: 3.11
skew after log1p:    -0.04
skew after Yeo-Johnson: 0.00
fitted lambda: 0.018

The walkthrough

The skew number is your before-and-after meter. Zero means balanced. Above about 1, the tail is long enough to matter. Raw incomes score 3.11 — the max is over 200 times the min. One log1p later, the skew is essentially gone.

log1p instead of log. np.log(0) is negative infinity, and one zero in the column poisons everything downstream. log1p computes log(1 + x), which maps 0 to 0 and behaves like plain log for large values. Make it your reflex for counts and amounts.

PowerTransformer finds the squash strength for you. It searches for lambda — the knob controlling how hard to squash — that makes the output most bell-shaped. Here it found 0.018, close to zero, which is the "use a log" setting. It also standardises its output by default, so you get shape-fixing and scaling in one step.

Yeo-Johnson vs Box-Cox. Box-Cox is the older method and demands strictly positive input. Yeo-Johnson handles zeros and negatives, which is why sklearn made it the default. If your column contains refunds or losses, Box-Cox raises an error; Yeo-Johnson does not.

Common mistakes

Logging a column that contains zeros or negatives. log(0) is -inf, log(-5) is nan, and both flow silently into the model until an error erupts far from the cause. Use log1p for counts, Yeo-Johnson for anything that can go negative.

Forgetting to undo the transform on predictions. Train a regression on log1p(price) and the model predicts log-prices. Report np.expm1(prediction) — the exact inverse — or your "predicted price" of 11.9 will confuse everyone.

Transforming an already-balanced column. Log-transforming symmetric data creates skew in the other direction. Check skew() first; reshape only what needs it.

Fitting the transformer on all rows. PowerTransformer learns lambda from data, so it leaks like any other fitted step. Fit on training rows only, then transform the rest.

Try it yourself

Generate rng.normal(50, 10, 1000) — already bell-shaped — and push it through the same script. Watch the skew before and after, and check what lambda the transformer picks when there is nothing to fix.

What to learn next

Researcher — Mathematics and papers.

The Box-Cox family

For x > 0 and parameter lambda:

y(lam) = (x^lam - 1) / lam, for lam != 0; y(0) = ln x

Where lam is the transform's exponent, chosen by maximising the profile log-likelihood under a Gaussian assumption on the transformed data (Box and Cox, 1964, An analysis of transformations, JRSS B). lam = 1 leaves the shape alone up to an affine shift; lam = 0.5 is a square root; lam = 0 a log; lam = -1 a reciprocal. Writing (x^lam - 1)/lam rather than bare x^lam makes the family continuous in lam at 0.

Yeo-Johnson

Yeo and Johnson (2000, A new family of power transformations to improve normality or symmetry, Biometrika) extend the family to all reals:

  • x >= 0: ((x + 1)^lam - 1) / lam, with ln(x + 1) at lam = 0
  • x < 0: -((-x + 1)^(2 - lam) - 1) / (2 - lam), with -ln(-x + 1) at lam = 2

The negative branch mirrors the positive one with exponent 2 - lam, keeping the function monotone and twice differentiable at 0. This is sklearn's PowerTransformer default.

What it buys, formally

  • Variance stabilisation. For a Poisson-like variable with variance proportional to the mean, the square root (lam near 0.5) makes variance approximately constant; for multiplicative noise (log-normal), the log does. Homoscedastic errors are an assumption of OLS inference, not a nicety.
  • Linearisation. A relationship y = a * x^b becomes linear in logs. The transform can move a problem from "needs a flexible model" to "a linear model suffices".
  • Retransformation bias. A model estimating conditional means on a log target satisfies E[exp(y_hat)] != exp(E[y_hat]) — predictions inverted naively are biased low. Duan's smearing estimator (Duan, 1983) is the standard correction.

Alternatives

QuantileTransformer maps each feature through its empirical CDF to a uniform or normal target — stronger and non-parametric, but it distorts within-tail spacing and needs enough data per feature. Rank-based inverse normal transforms are its statistical cousin. For heavy tails that carry signs, the inverse hyperbolic sine (asinh) behaves like a signed log and needs no parameter search.

What to learn next