Preprocessing and Feature Selection
Recursive feature elimination
RFE trains a model, drops the weakest column, retrains, and repeats — an elimination tournament that judges every column by how the model actually uses it.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Recursive feature elimination is a knockout tournament for columns: train the model, drop the weakest column, retrain, repeat.
Think of selecting a cricket team from thirty players. You do not judge each player from paper statistics alone. You play practice matches, watch who contributes least in the actual game, drop that one player, and play again. Each new match re-tests everyone in the changed team.
That re-testing is the whole point. A player might have looked replaceable while a star occupied their role — with the star dropped, they shine. Columns behave the same way inside a model.
Why it exists
Filter methods judge each column alone, without a model. Cheap, but blind to two things:
- Teamwork. Two columns can be individually weak but decisive together. Filters drop both.
- The model's own opinion. A column useful to one model can be useless to another. Filters cannot know, because they never train one.
RFE fixes both by asking the model itself, repeatedly. The price is honest: training the model many times over. RFE is a wrapper method — a selection method wrapped around real model training.
How it works
round 1: train on 10 columns -> weakest column out (9 left)
round 2: train on 9 columns -> weakest out (8 left)
round 3: train on 8 columns -> weakest out (7 left)
...continue until the wanted number remain...Why not drop the seven weakest at once? Because eliminations change the game. A column can inherit importance when its overlapping partner leaves. One elimination per round lets the ranking update after every change — exactly like the practice matches.
A real example you have seen
Reality-show eliminations work this way, and for the same reason. One contestant leaves per week, and the dynamics re-form around the survivors. A singer who was overshadowed suddenly stands out in week five. Eliminating half the cast in week one, by first impressions, would have missed them.
Remember this
- RFE ranks columns by how the trained model uses them, not by solo statistics.
- Dropping one at a time lets remaining columns reveal hidden value.
- The cost is many model trainings — the price of asking the model's opinion.
What to learn next
- Boruta and shadow features — selection with a built-in control experiment.
- Filter feature selection — the cheap first pass that keeps RFE affordable.
- Random forest — the estimator most often wrapped when patterns are non-linear.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2.
Ten columns, three real, and a referee
RFECV runs the tournament and additionally uses cross-validation to decide how many columns to keep — the number you would otherwise have to guess.
from sklearn.datasets import make_classification
from sklearn.feature_selection import RFECV
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
# 10 columns: the first 3 carry signal, the other 7 are pure noise
X, y = make_classification(n_samples=500, n_features=10, n_informative=3,
n_redundant=0, n_repeated=0, class_sep=2.0,
random_state=0, shuffle=False)
X = StandardScaler().fit_transform(X)
model = LogisticRegression(max_iter=1000)
rfe = RFECV(model, step=1, cv=5).fit(X, y) # cross-validation picks how many
print("columns kept:", rfe.support_.astype(int))
print("ranking: ", rfe.ranking_)
print("best number of features:", rfe.n_features_)columns kept: [1 1 1 0 0 0 0 0 0 0] ranking: [1 1 1 8 7 2 3 6 4 5] best number of features: 3
The walkthrough
It found exactly the right three. The dataset was built with signal in the first three columns (shuffle=False keeps them first). support_ marks precisely those, and ranking_ gives every survivor rank 1 while the noise columns get their elimination order — rank 8 fell first.
How "weakest" is measured. After each training, RFE reads the model's importance signal: coefficient sizes (coef_) for linear models, feature_importances_ for trees. The column with the smallest value is out. This is why the estimator you wrap matters — you are inheriting its notion of importance, including its blind spots.
Scaling is not optional here. With unscaled features, a logistic regression coefficient's size reflects the feature's units, not its importance. Income-in-rupees would get a tiny coefficient and be eliminated first, regardless of merit. Standardise before wrapping a linear model.
step controls how many columns fall per round. step=1 is the careful tournament; step=5 trades care for speed on wide data. RFECV then evaluates each squad size with 5-fold cross-validation and keeps the size with the best average score.
Common mistakes
Running RFE on all rows, then cross-validating the survivor set. The selection saw every row, so the later scores are contaminated — the same leakage trap as with filters, and it inflates results dramatically on noisy data. Put the whole RFECV inside a Pipeline, or select on training folds only.
Wrapping a model with no importance signal. KNeighborsClassifier has neither coef_ nor feature_importances_; RFE raises an error. Either wrap a model that exposes importances or pass importance_getter with a custom function.
Ignoring the cost. RFE with step=1 on p columns trains the model roughly p times, and RFECV multiplies that by the folds. On 5,000 columns with a slow model, that is a lunch break, and possibly a weekend. Cut the field first with a cheap filter, then run RFE on the shortlist.
Treating the kept set as the truth. Rerun with a different seed or fold split and the boundary picks can change, especially among correlated columns. The stable core matters; the edge cases deserve scepticism.
Try it yourself
Change to n_redundant=2 so two columns are copies of signal columns in disguise. Rerun and study ranking_: watch how RFE handles the duplicates, and check whether the chosen n_features_ grows.
What to learn next
- Boruta and shadow features — selection with a built-in control experiment.
- Filter feature selection — the cheap first pass that keeps RFE affordable.
- Random forest — the estimator most often wrapped when patterns are non-linear.
Researcher — Mathematics and papers.
Origin: SVM-RFE
Guyon, Weston, Barnhill and Vapnik (2002), Gene selection for cancer classification using support vector machines, Machine Learning, introduced RFE for linear SVMs on microarray data with p around 7,000 genes and n around 60 samples. The ranking criterion is w_j^2 — the squared weight of feature j in the max-margin hyperplane — motivated as a first-order estimate of the increase in the objective when feature j is removed (an OBD-style sensitivity argument, after LeCun's Optimal Brain Damage). Eliminating one feature per iteration and re-fitting captures conditional relevance: the criterion for j is evaluated in the context of the current surviving set.
Cost and structure
With p features and elimination step s, RFE performs ceil(p/s) model fits; RFECV multiplies by k folds, then one final fit. For an O(C(n, p)) trainer the total is sum over rounds of C(n, p_t) — for linear models effectively O(k * n * p^2 / s). The procedure is greedy backward selection: it never revisits an eliminated feature, so it inherits the usual greedy failure mode — a feature eliminated early cannot return, even if it becomes conditionally valuable later. Floating variants (SFFS/SFBS; Pudil et al., 1994) reintroduce features to patch exactly this.
Relation to embedded selection
L1-regularised models perform selection during a single fit by driving coefficients to zero along the regularisation path — the lasso (Tibshirani, 1996). The elastic-net compromise handles correlated groups better, whereas both lasso and RFE-with-linear-models split importance across correlated features and can arbitrarily keep one of a pair. Stability selection (Meinshausen and Buhlmann, 2010) — running selection over subsamples and keeping features chosen often — converts either method's brittleness into a controllable false-discovery guarantee, and is the standard remedy when the identity of selected features matters scientifically.
Choosing the criterion honestly
Coefficient magnitude and impurity-based feature_importances_ are both biased criteria: the former by scale and collinearity, the latter toward high-cardinality features. Permutation importance (Breiman, 2001) measured on held-out folds is slower but model-faithful, and can be plugged into RFE via importance_getter. If the estimator's importances are untrustworthy, RFE ranks noise carefully — the tournament is only as good as its referee.
What to learn next
- Boruta and shadow features — selection with a built-in control experiment.
- Filter feature selection — the cheap first pass that keeps RFE affordable.
- Random forest — the estimator most often wrapped when patterns are non-linear.