Preprocessing and Feature Selection

Filter feature selection

Filter methods score each column on its own — with statistics like the F-test or mutual information — and drop the weak ones before any model is trained.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Filter feature selection scores every column on its own, cheaply, and throws away the ones that carry no information about the answer.

Think of hiring for one job from five hundred CVs. You will not interview five hundred people — interviews are expensive. First you shortlist: does the CV mention the right skills, the right experience? Each CV is judged alone, by quick rules, without ever meeting the person.

Filters treat columns the same way. Each column gets a quick score for "does this, by itself, relate to the thing we predict?". Low scorers are dropped before any model — the expensive interview — is trained.

Why it exists

More columns is not more knowledge. Useless columns cost three ways:

  1. Noise gets mistaken for signal. Give a model 200 random columns and a few will correlate with the answer by pure luck. The model learns the luck, then fails on new data — the story of overfitting.
  2. Everything gets slower — training, prediction, storage, debugging.
  3. Nobody can explain the model. Forty columns can be discussed in a meeting. Four thousand cannot.

Filters are the cheapest defence: one quick statistical look per column, no model training involved. That cheapness is also their limit — a filter looks at each column alone, so it cannot see teamwork between columns.

How it works

column      quick score against the target       verdict
--------    -------------------------------      -------
age            strong link                        keep
income         weak link                          keep
noise          no link                            drop
lucky_no       no link                            drop

Score, rank, keep the top few. That is the whole method.

A real example you have seen

A doctor ordering tests works as a filter. Before any complex diagnosis, cheap single measurements — temperature, blood pressure, one blood test — rule out most possibilities. Expensive whole-body scans are reserved for what survives the screening.

Remember this

  • Filters score each column independently against the target — fast and model-free.
  • They catch outright junk but miss columns that only work in teams.
  • Selection must happen on training data only, or the test score becomes a lie.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Outputs verified with scikit-learn 1.7.2.

Two real signals, two impostors

filter_fs.py
import numpy as np
from sklearn.feature_selection import SelectKBest, f_classif, mutual_info_classif

rng = np.random.default_rng(0)
n = 200
age = rng.uniform(20, 60, n)
income = rng.uniform(20000, 90000, n)
noise = rng.normal(size=n)                       # pure junk column
lucky = rng.integers(0, 10, n).astype(float)     # also junk

# buying insurance depends on age, weakly on income, never on the junk
y = (age + income / 4000 + rng.normal(0, 6, n) > 55).astype(int)

X = np.column_stack([age, income, noise, lucky])
names = ["age", "income", "noise", "lucky"]

scores = f_classif(X, y)[0]
mi = mutual_info_classif(X, y, random_state=0)
for name, f, m in zip(names, scores, mi):
    print(f"{name:7} F = {f:7.1f}   MI = {m:.3f}")

keep = SelectKBest(f_classif, k=2).fit(X, y)
print("kept:", [n_ for n_, k_ in zip(names, keep.get_support()) if k_])
Output
age     F =   281.9   MI = 0.445
income  F =     7.9   MI = 0.008
noise   F =     0.0   MI = 0.000
lucky   F =     1.4   MI = 0.000
kept: ['age', 'income']

The walkthrough

Two scorers, two personalities. The F-test (f_classif) asks: do the class averages of this column differ more than chance would allow? It is fast and sharp for straight-line relationships. Mutual information (mi) asks a broader question: does knowing this column reduce uncertainty about the target in any way, curved or straight? It is slower and needs more data, but catches patterns the F-test is blind to.

Both agree here. Age dominates, income is weak but real, the junk columns score at the floor. SelectKBest(k=2) keeps the top two and gives you a transformer that drops the rest everywhere downstream.

The F-test's blind spot is worth seeing. A U-shaped relationship — target high at both extremes of a feature — has equal class means, so F scores near zero while MI still fires. If you suspect curved effects, run both and investigate disagreements; they are the interesting columns.

Choosing k. There is no formula. Treat k as a tuning knob inside cross-validation, or use SelectPercentile, or plot sorted scores and cut at the visible cliff.

Common mistakes

Selecting features using all the data, then splitting. The selector saw the test rows when ranking columns, so the test set is contaminated and scores read high. This is the classic silent leakage bug in published work. Fit the selector on training data only — inside the cross-validation loop, ideally in a Pipeline.

Keeping both of two near-duplicate columns. Filters score columns one at a time, so height-in-cm and height-in-inches both score high and both get kept. Filters deduplicate nothing. Check pairwise correlations separately.

Dropping a column that only works in a team. Two columns can be useless alone and decisive together — each alone scores zero. Filters will drop both. When teamwork matters, you need wrapper methods like recursive feature elimination.

Using chi2 on negative values. The chi-squared scorer assumes counts or frequencies and raises an error on negatives. It is for non-negative data like word counts, not standardised features.

Try it yourself

Add a fifth column xor = ((age > 40) != (income > 55000)).astype(float) and make the target depend on it. Score everything again and watch the F-test shrug at a column that genuinely drives the answer.

What to learn next

Researcher — Mathematics and papers.

The two scorers, precisely

ANOVA F-statistic. For feature x and K classes, F = (between-class variance) / (within-class variance):

F = [ sum_k n_k (xbar_k - xbar)^2 / (K - 1) ] / [ sum_k sum_{i in k} (x_i - xbar_k)^2 / (n - K) ]

Where n_k is the size of class k, xbar_k its mean of x, xbar the grand mean, n the total count. Under Gaussian within-class x and equal variances, F follows an F(K-1, n-K) distribution, giving p-values. It detects only first-moment (mean-shift) dependence.

Mutual information. I(X; Y) = sum over x, y of p(x, y) log [ p(x, y) / (p(x) p(y)) ] — zero if and only if X and Y are independent, capturing arbitrary dependence. For continuous features sklearn estimates it with the Kraskov-Stogbauer-Grassberger k-nearest-neighbour estimator (Kraskov et al., 2004), inheriting its variance at small n and its n_neighbors sensitivity.

Redundancy-aware filters

Univariate scoring ignores inter-feature structure. mRMR (minimum-redundancy maximum-relevance; Peng et al., 2005) greedily selects features maximising I(x_j; Y) minus the average I(x_j; x_selected) — relevance minus redundancy. Correlation-based feature selection (Hall, 1999) pursues the same trade with correlation ratios. These sit between pure filters and wrappers in cost.

Selection bias, the recurring scandal

Ambroise and McLachlan (2002, PNAS) re-analysed gene-expression classifiers that reported near-zero error: selecting genes on the full dataset before cross-validation produced wildly optimistic estimates — bias large enough to make random data look predictive, since with p much greater than n, some of thousands of features always correlate with any labels by chance. The corrected protocol — selection refit inside every CV fold — remains the single most important practical rule in this topic. Guyon and Elisseeff (2003), An introduction to variable and feature selection, JMLR, is the standard survey and includes the constructed examples (XOR, useless-alone pairs) behind the teamwork caveat.

Where filters sit

Filters are O(p) scoring passes, embarrassingly parallel, model-agnostic. Wrappers (RFE, forward selection) query a model repeatedly. Embedded methods make selection part of training — L1 regularisation zeroing coefficients, tree ensembles ranking split gains. On modern tabular pipelines a common sequence is: variance/near-duplicate pruning, a filter pass to cut the field, then an embedded or wrapper method for the final squad.

What to learn next