Preprocessing and Feature Selection

Binning and discretisation

Binning chops a continuous number like age or income into a handful of buckets, trading fine detail for robustness and rules a human can read.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Binning turns a smooth number into a small set of buckets — age 34 becomes "30 to 40".

You have done this on every form you have ever filled. The form never asks your exact income; it offers tick-boxes: below 3 lakhs, 3 to 10 lakhs, above 10 lakhs. T-shirt makers do not stitch a size for every chest measurement; they make S, M, L, XL. A smooth measurement gets chopped into a few named ranges.

That chopping is discretisation, and each range is a bin.

Why it exists

Detail is not always your friend. Three reasons to throw some away on purpose:

  1. Wild values stop mattering. A crorepati and a lakhpati land in the same "high income" bucket. One absurd row can no longer stretch the whole column, the problem you met in power transforms.
  2. Bent patterns become learnable. Insurance risk is high for the very young, low in the middle, high again for the old. A straight-line model cannot draw that U shape from raw age. Give it three age buckets, and each bucket gets its own vote.
  3. Humans can read the result. "Applicants under 25 default twice as often" survives a meeting. A coefficient of 0.0034 per year does not.

How it works

There are two honest ways to draw the bucket edges:

equal width:    |----|----|----|----|      same size ranges,
                 ####  #                    crowded + empty bins

equal count:    |-|--|---|------|-----|    ranges stretch so each
                 ##  ##  ##  ##  ##         bin holds equal people

Equal width cuts the range into same-sized slices. On lopsided data, almost everyone lands in slice one and the top slices sit empty. Equal count (quantile binning) moves the edges so every bin holds the same number of rows. For skewed data, equal count is nearly always what you want.

A real example you have seen

Your electricity bill is computed on slabs: the first 100 units at one rate, the next 100 at a higher rate, and so on. The board discretised your usage. Delivery apps do the same with distance slabs, and airlines with baggage weight.

Remember this

  • Binning trades detail for robustness and readability.
  • Equal-width bins collapse on skewed data; quantile bins stay balanced.
  • Outliers lose their power: the biggest bin swallows them.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Outputs verified with scikit-learn 1.7.2. The quantile_method argument below silences a FutureWarning new in the 1.7 line; it becomes the default in 1.9.

Watching equal-width binning fail

binning.py
import numpy as np
from sklearn.preprocessing import KBinsDiscretizer

rng = np.random.default_rng(7)
incomes = rng.lognormal(mean=10.5, sigma=0.8, size=1000).reshape(-1, 1)

uniform = KBinsDiscretizer(n_bins=5, encode="ordinal", strategy="uniform")
# quantile_method silences a FutureWarning in scikit-learn 1.7;
# it becomes the default in 1.9
quantile = KBinsDiscretizer(n_bins=5, encode="ordinal", strategy="quantile",
                            quantile_method="averaged_inverted_cdf")

u = uniform.fit_transform(incomes).ravel().astype(int)
q = quantile.fit_transform(incomes).ravel().astype(int)

print("uniform bins  (equal width):", np.bincount(u, minlength=5))
print("quantile bins (equal count):", np.bincount(q, minlength=5))
print("uniform edges: ", uniform.bin_edges_[0].round(0))
print("quantile edges:", quantile.bin_edges_[0].round(0))
Output
uniform bins  (equal width): [764 178  37  16   5]
quantile bins (equal count): [200 200 200 200 200]
uniform edges:  [  2694.  58990. 115285. 171581. 227876. 284172.]
quantile edges: [  2694.  17827.  29056.  39652.  64243. 284172.]

The walkthrough

Read the uniform counts: 764, 178, 37, 16, 5. Three-quarters of all rows fell into bin 0, and the top bin holds five people. Any model consuming this column learns almost nothing — the feature has become "poor or not", plus three nearly-empty categories. This is what equal width does to skewed data.

The quantile edges tell the real story. The first bin spans 2,694 to 17,827 — a narrow slice where people are dense. The last spans 64,243 to 284,172 — a huge slice where people are sparse. The edges adapted so every bucket carries equal evidence.

encode="ordinal" outputs bucket numbers 0 to 4 in one column. The alternative, encode="onehot" (the default), outputs five 0/1 columns. Ordinal keeps the order information — bin 3 is genuinely richer than bin 2 — and suits tree models. One-hot lets a linear model give each bucket a fully independent effect, which is the whole point for U-shaped patterns.

In pandas, the same two strategies are pd.cut (equal width, or your own edges) and pd.qcut (equal count).

Common mistakes

Fitting bin edges on all rows. Edges computed from data are learned parameters. Fit on the training set, apply to the test set — same discipline as any scaler. See train-test split.

Too many bins. Fifty bins over 500 rows means ten rows per bin, and each bin's average is mostly noise. The model memorises accidents — a direct path to overfitting. Start at 4 or 5 bins; add more only with evidence.

Binning a cleanly linear feature. If the target rises steadily with the feature, binning replaces a smooth ramp with a staircase and throws away precision for nothing. Bin to fix a problem — outliers, U shapes, slab-like business rules — not by routine.

Ordinal-encoding truly unordered categories. Bucket numbers imply order. That is right for income slabs and wrong if the buckets are "north, south, east, west". Unordered categories want one-hot instead.

Try it yourself

Rerun the script with n_bins=20. Look at the uniform bin counts and find how many bins hold fewer than three people. Then decide which you would trust in a model: 20 uniform bins or 5 quantile bins.

What to learn next

Researcher — Mathematics and papers.

Discretisation as basis expansion

Binning feature x with edges t_1 < t_2 < ... < t_{k-1}, then one-hot encoding, replaces x with indicator functions 1[t_j <= x < t_{j+1}]. A linear model on these indicators fits a piecewise-constant function of x — a regression histogram. The bias-variance trade-off is controlled by k: bias falls as O(1/k) for Lipschitz targets while variance grows as O(k/n). Splines (piecewise polynomials with continuity constraints) dominate plain binning when smoothness is plausible; binning survives because it is easy to explain and robust to any monotone data corruption within a bin.

Supervised discretisation

Unsupervised edges (width, quantile, k-means on the single feature — the three KBinsDiscretizer strategies) ignore the target. Supervised methods place edges to maximise target purity:

  • MDLP (Fayyad and Irani, 1993, Multi-interval discretization of continuous-valued attributes): recursive entropy-minimising splits with an MDL stopping criterion.
  • ChiMerge (Kerber, 1992): bottom-up merging of adjacent intervals whose class distributions are statistically indistinguishable by chi-squared test.
  • Decision trees are implicit supervised discretisers: every split threshold is a learned bin edge. Fitting a shallow tree on one feature and using its thresholds as edges is a practical, leakage-prone-if-misused trick — do it inside cross-validation.

Binning inside modern gradient boosting

Histogram-based gradient boosting (LightGBM; sklearn's HistGradientBoostingClassifier) quantile-bins every feature to at most 255 levels before training. Split search then costs O(bins) instead of O(n log n) per feature. Accuracy is usually unchanged — empirical support for the claim that fine-grained precision in tabular features rarely carries signal. Ke et al. (2017), LightGBM: A highly efficient gradient boosting decision tree, NeurIPS.

Information-theoretic view

Discretisation is lossy compression of a feature: it bounds the mutual information I(binned X; Y) <= I(X; Y). The design question is how much information about the target survives k bins. Equal-frequency binning maximises the entropy of the binned feature itself, which is a sensible default when the target is unknown at design time but optimal for no particular target.

What to learn next