Ensembles and Gradient Boosting

Stacking

Stacking trains a second-stage model whose only inputs are the predictions of your first-stage models, learning whose opinion to trust and when.

On this page 5
  1. Why this had to be invented
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Stacking trains an extra model whose inputs are the predictions of your other models, so the combining itself is learned instead of fixed.

Think of a hospital case conference. A cardiologist, a radiologist and a surgeon each give an opinion on a difficult patient. Chairing the meeting is a senior physician who has worked with all three for years. She knows the radiologist is superb on scans but overcautious with older patients, and she weighs each voice accordingly.

A voting ensemble counts the three opinions equally. Stacking hires the chairperson. The extra model is called the meta-learner — a model that learns about other models' outputs — and the models under it are base models.

Why this had to be invented

Equal votes waste information. Maybe your forest is excellent except on small-value transactions, where the linear model quietly does better. A fixed vote cannot express "trust A, except when B and C agree against it". A learned combiner can — that rule is a pattern in the base models' outputs, and finding patterns is what models do.

There is one trap, and it defines the method. The chairperson must judge the specialists by their performance on patients they had not already studied. If she rates the radiologist on cases he memorised, the most overconfident memoriser wins her trust. In model terms, the meta-learner must train on predictions the base models made for rows they never saw. These are out-of-fold predictions. They come from the rotation trick used in k-fold cross-validation.

How it works

                training data
              /       |       \
        forest      knn       svm          stage 1: each base model
              \       |       /                     predicts (on rows it
            [0.91,  0.70,  0.85]                    did not train on)
                      |
             logistic regression           stage 2: the meta-learner
                      |                             learns to combine
                final answer

At prediction time the flow is the same without the rotation: each base model predicts, the meta-learner combines.

A real example you have seen

The Netflix Prize (2006–2009), the contest that popularised industrial-scale model combining. The winning entries blended hundreds of models with learned combiners. That taught the field both lessons at once. Stacking wins competitions, and a hundred-model stack can be far too unwieldy to ship. Kaggle leaderboards have run on stacked ensembles ever since.

Remember this

  • Stacking = base models + a meta-learner trained on their predictions.
  • The meta-learner must see only out-of-fold predictions, or it rewards memorisers.
  • Gains are real but usually small — the price is complexity at prediction time.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU. Runs in about half a minute.

Three specialists and a chairperson

stacking_demo.py
from sklearn.datasets import make_classification
from sklearn.ensemble import RandomForestClassifier, StackingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.svm import SVC

X, y = make_classification(n_samples=1200, n_features=16, n_informative=6,
                           flip_y=0.05, random_state=13)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.4, random_state=13)

base = [("forest", RandomForestClassifier(n_estimators=200, random_state=13)),
        ("knn", KNeighborsClassifier(n_neighbors=15)),
        ("svm", SVC(random_state=13))]

for name, model in base:
    print(f"{name:7s} {model.fit(Xtr, ytr).score(Xte, yte):.3f}")

stack = StackingClassifier(base, final_estimator=LogisticRegression(), cv=5)
print(f"stack   {stack.fit(Xtr, ytr).score(Xte, yte):.3f}")
Output
forest  0.848
knn     0.850
svm     0.865
stack   0.875

One point over the best specialist. That is a typical stacking margin: real, worth having in a competition, and worth questioning in production — more on that below.

The walkthrough

cv=5 is the anti-cheating machinery. During fit, scikit-learn splits the training data five ways, trains each base model on four parts, and collects its predictions on the fifth — rotating until every row has a prediction from a model that never saw it. Those out-of-fold predictions train the logistic regression. Then the base models are refitted on all the training data for use at prediction time. You get the entire discipline by passing one integer.

The meta-learner should be humble. Its input is three numbers per row. LogisticRegression — effectively learned weights per specialist — is the standard choice. A deep meta-forest on three columns invites the stack to overfit the folds themselves.

stack_method defaults to probabilities. Where a base model offers predict_proba or decision_function, the meta-learner sees those graded scores rather than hard yes/no votes — the chairperson hears how confident each specialist is. That nuance is a main source of stacking's edge over hard voting.

Diversity still rules. The base models here fail differently by construction — boxes, neighbourhoods, margins. Stacking three tuned forests yields three near-identical columns, and the meta-learner has nothing to work with. The first lesson's law applies to inputs of the meta-learner too.

Common mistakes

Hand-rolling the fold logic wrongly. The classic bug: fit base models on all training data, feed their training predictions to the meta-learner. The most overfit base model looks perfect, the meta-learner trusts it completely, and the stack underperforms on real data. If you build stacking manually, out-of-fold discipline is the entire difficulty; StackingClassifier exists because so many hand-rolled versions got it wrong.

Stacking clones. Base models that agree everywhere give the meta-learner constant columns. Check pairwise disagreement of base predictions before bothering with a stack.

A clever meta-learner. Gradient boosting as the final estimator over three columns will happily memorise fold noise. Linear or logistic stays the default until proven insufficient.

Ignoring the serving bill. Every prediction now runs a forest, a kNN search, an SVM, then a combiner — for one extra point of accuracy. In production serving terms, that is three models' latency and maintenance for a margin your product may not feel. Competitions and production have different budgets; decide which one you are in.

Try it yourself

Add passthrough=True to the StackingClassifier — the meta-learner then sees the original 16 features alongside the three predictions. Does it help here? Then replace all three base models with RandomForestClassifier variants differing only in random_state, and watch the stack's edge evaporate.

What to learn next

Researcher — Mathematics and papers.

Formulation

Level-0 models ${g_m}_{m=1}^{M}$ and a level-1 combiner $h$. With $k$-fold partition ${F_1, \dots, F_k}$ of the training set, the out-of-fold prediction matrix $Z \in \mathbb{R}^{n \times M}$ has entries

$$ Z_{im} = g_m^{(-\kappa(i))}(x_i) $$

Where $\kappa(i)$ is the fold containing row $i$, and $g_m^{(-\kappa(i))}$ is model $m$ trained on all folds except $\kappa(i)$. The combiner solves $\hat{h} = \arg\min_h \sum_i L\big(y_i, h(Z_{i\cdot})\big)$, and the deployed predictor is $h(g_1(x), \dots, g_M(x))$ with each $g_m$ refitted on the full data.

Origin: Wolpert (1992), Stacked generalization, Neural Networks — framed as reducing the generalisation bias of any single hypothesis. Breiman (1996), Stacked regressions, made it practical by restricting $h$ to non-negative linear combinations, $h(z) = \sum_m \beta_m z_m$, $\beta_m \geq 0$: the constraint is a regulariser that prevents the combiner exploiting cancelling overfits, and it frequently outperforms unconstrained least squares.

The super learner result

Van der Laan, Polley and Hubbard (2007), Super learner (Statistical Applications in Genetics and Molecular Biology), prove an oracle inequality: the cross-validation-selected convex combination of base learners performs asymptotically as well as the best possible combination in the library, up to a term of order $\sqrt{\log M / n}$. Practical reading: adding candidate models to the library is nearly free asymptotically, so include cheap diverse baselines. The result assumes the out-of-fold discipline exactly as formulated — the leakage-free $Z$ is the theorem's object.

Blending versus stacking

Blending replaces the $k$-fold rotation with a single held-out split: base models train on part A, the combiner trains on their predictions over part B. Cheaper, no fold bookkeeping, but the combiner sees fewer rows and the split adds variance. Stacking with $k$-fold uses all data for both stages and is the default; blending survives where retraining base models $k$ times is unaffordable.

Restacking / multi-level stacks feed level-1 outputs into a level-2 combiner. Netflix-Prize-era systems went three levels deep (Töscher, Jahrer and Bell, 2009, The BigChaos solution to the Netflix Grand Prize). Returns diminish sharply after level 1 on most problems, while leakage risk and serving cost compound.

When stacking pays

Empirically, stacking's edge over the best single model correlates with the spread and decorrelation of base-model errors — the ambiguity decomposition again, applied to the meta-features. On tabular problems where one gradient-boosted model dominates, stacks of {GBM, linear, kNN, NN} typically add a fraction of a percent; modern AutoML systems (AutoGluon: Erickson et al., 2020) nonetheless stack by default, because at their scale the fraction is systematic and the fold plumbing is amortised. The counterweight is operational: latency multiplies with $M$, and failure analysis must now trace through two stages — one reason distillation of a stack into a single model is a common deployment pattern.

What to learn next

What to learn next

These follow on from what you just read.

  • Preprocessing and Feature Selection

    Feature scaling

    Feature scaling squeezes every column onto a similar range so that no column drowns out the others by accident of its units.

  • Preprocessing and Feature Selection

    Power transforms

    Power transforms like log and Yeo-Johnson reshape lopsided columns into balanced ones, so a few giant values stop dominating everything the model learns.

  • Preprocessing and Feature Selection

    Binning and discretisation

    Binning chops a continuous number like age or income into a handful of buckets, trading fine detail for robustness and rules a human can read.