Ensembles and Gradient Boosting

Why ensembles work

Combining many imperfect models cancels their private mistakes, which is why a committee of average models beats one careful model so often.

On this page 5
  1. Why this had to be invented
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An ensemble is a team of imperfect models whose combined answer is more reliable than any single member's answer.

Think about a serious medical decision. You would not act on one doctor's opinion alone. You would visit two or three, and go with what most of them agree on. Each doctor makes occasional mistakes — but they make different mistakes, so the agreement between them is safer than any one of them.

That is the entire idea. In machine learning, an ensemble means several trained models whose answers are averaged or put to a vote.

Why this had to be invented

A single model has a personality. A decision tree carves the data into sharp boxes. A linear model draws one straight boundary. Each personality suits some parts of a problem and fails on others.

Worse, a flexible model is twitchy. Train it on a slightly different sample of the same data and it gives noticeably different answers. That twitchiness is called variance — the model's tendency to change its mind when the data changes a little.

For years the response was to hunt for one better model. The ensemble insight goes the other way. Keep the imperfect models. Make sure they are imperfect in different ways. Then let them vote, and watch their private errors cancel.

How it works

 question:  "is this transaction fraud?"

 model A  →  yes      (wrong on some cases others get right)
 model B  →  yes      (wrong on different cases)
 model C  →  no       (wrong on different cases again)
              |
        majority vote
              |
            "yes"     ← more often correct than A, B or C alone

Two conditions must hold, and both matter.

First, each member must be better than random guessing. A committee of coin-flippers agrees on nothing useful.

Second, the members must disagree sometimes. Eleven copies of the same model produce the same answer eleven times. The vote adds nothing. Diversity is the fuel — different algorithms, different samples of the data, or different features.

A real example you have seen

On the game show Kaun Banega Crorepati, a stuck contestant can ask the studio audience. Hundreds of ordinary people press a button, and the most popular option is right far more often than any single audience member would be. No one in the crowd is an expert. The crowd is.

Weather forecasts work the same way. The "70% chance of rain" you see comes from running many slightly different simulations and counting how many of them predict rain.

Remember this

  • An ensemble combines several models by voting or averaging.
  • It works because members make different mistakes, and different mistakes cancel.
  • Members must each beat random guessing, and they must sometimes disagree.

What to learn next

  • Bagging — the standard recipe for manufacturing diverse models from one dataset.
  • Random forest — bagging plus feature randomness, the workhorse ensemble.
  • XGBoost — the other family: models built in sequence, each fixing the last.

Developer — Code and libraries.

Before touching any real algorithm, it is worth seeing the effect with nothing but random numbers. The whole phenomenon fits in ten lines.

Setup

bash
pip install numpy scikit-learn

Outputs verified with numpy 1.26.4 and scikit-learn 1.7.2 on CPU. Everything here runs in seconds.

The committee effect, from scratch

Eleven voters. Each answers 10,000 questions, and each is right 65% of the time, independently of the others.

committee.py
import numpy as np

rng = np.random.default_rng(42)
n_voters, n_questions = 11, 10_000

# each voter answers each question correctly with probability 0.65, independently
correct = rng.random((n_voters, n_questions)) < 0.65
votes_right = correct.sum(axis=0)              # how many of the 11 got each question right
print("one voter alone:      ", correct[0].mean())
print("majority of 11 voters:", (votes_right >= 6).mean())

# now remove the independence: all 11 votes are copies of voter 0
copies = np.repeat(correct[:1], n_voters, axis=0)
print("11 copies of one voter:", (copies.sum(axis=0) >= 6).mean())
Output
one voter alone:       0.6543
majority of 11 voters: 0.8493
11 copies of one voter: 0.6543

Eleven mediocre voters, each right about 65% of the time, reach 85% together. The same eleven votes, when they are copies of one voter, gain exactly nothing. Independence did all the work.

The same effect with real models

Three genuinely different algorithms on the same noisy two-moons data.

vote.py
from sklearn.datasets import make_moons
from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier

X, y = make_moons(n_samples=400, noise=0.35, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)

models = [("tree", DecisionTreeClassifier(random_state=0)),
          ("knn", KNeighborsClassifier()),
          ("logreg", LogisticRegression())]

for name, model in models:
    print(f"{name:7s} {model.fit(Xtr, ytr).score(Xte, yte):.3f}")

vote = VotingClassifier(models).fit(Xtr, ytr)
print(f"vote    {vote.score(Xte, yte):.3f}")
Output
tree    0.840
knn     0.870
logreg  0.860
vote    0.880

The vote beats every member. An honest note: with other random seeds the vote sometimes ties the best member instead of beating it. The gain depends on how differently the members fail, and on this small dataset the margin is a percent or two. On harder problems the margin grows.

The walkthrough

correct.sum(axis=0) counts, for every question, how many of the 11 voters got it right. Comparing against 6 asks whether a majority did. If axis reads strangely, the dim rule is the same idea in PyTorch clothing.

VotingClassifier(models) trains each named model and, at prediction time, takes the majority answer. Passing voting="soft" averages predicted probabilities instead, which usually works a little better when members produce calibrated probabilities.

Why these three models? A tree carves boxes, k-nearest-neighbours reads local neighbourhoods, and logistic regression draws a straight line. Three different personalities, three different failure patterns. That disagreement is the raw material.

Common mistakes

Combining copies. Three runs of the same algorithm with the same data and settings vote identically. You saw the number: no gain at all. Diversity has to come from somewhere — different algorithms, different data samples, or different features.

Adding a much weaker member and expecting the vote to absorb it. A member far below the others drags the committee down, because its errors are too frequent to be outvoted reliably. Members should be at least roughly comparable.

Expecting an ensemble to fix a shared blind spot. If every model trains on the same misleading feature or the same leaky data, they are all wrong in the same direction, and the vote confidently repeats the shared mistake. Ensembles cancel independent errors only.

Reaching for an ensemble before checking one model honestly. Measure a single model with proper evaluation first. An ensemble of unmeasured models is a slower way to be unsure.

Try it yourself

In committee.py, drop the accuracy from 0.65 to 0.55, then to 0.51. Predict what happens to the majority before running. Then raise n_voters to 101 at accuracy 0.55 and watch the committee recover. The lesson: many barely-better-than-chance voters still add up, as long as they stay independent.

What to learn next

  • Bagging — the standard recipe for manufacturing diverse models from one dataset.
  • Random forest — bagging plus feature randomness, the workhorse ensemble.
  • XGBoost — the other family: models built in sequence, each fixing the last.

Researcher — Mathematics and papers.

Condorcet's jury theorem

The oldest result here predates machine learning by two centuries. For $M$ independent voters, each correct with probability $p$, the majority is correct with probability

$$ P_{maj} = \sum_{k=\lceil M/2 \rceil}^{M} \binom{M}{k} p^k (1-p)^{M-k} $$

Where:

  • $M$ — the number of voters (odd, to avoid ties).
  • $p$ — each voter's probability of being correct, assumed identical.
  • $k$ — the number of voters who happen to be correct.

If $p > 0.5$, then $P_{maj} \to 1$ as $M \to \infty$; if $p < 0.5$, it tends to 0 (Condorcet, 1785, Essai sur l'application de l'analyse). With $p = 0.65$, $M = 11$: $P_{maj} \approx 0.847$ — matching the simulation above to within sampling noise.

The theorem's independence assumption is exactly what real ensembles violate, which motivates the next decomposition.

Variance of a correlated average

For $M$ regression estimators, each with variance $\sigma^2$ and average pairwise correlation $\rho$, the variance of their mean is

$$ \operatorname{Var}\left(\bar{f}\right) = \rho \sigma^2 + \frac{1-\rho}{M} \sigma^2 $$

Where:

  • $\sigma^2$ — the variance of a single member's prediction across training samples.
  • $\rho$ — the average correlation between two members' errors.
  • $M$ — the ensemble size.

The second term dies as $M$ grows. The first term does not. Ensemble size attacks only the uncorrelated part; correlation sets the floor. Every practical ensemble method — bagging, random feature selection, different algorithms — is a device for pushing $\rho$ down.

The ambiguity decomposition

Krogh and Vedelsby (1995), Neural network ensembles, cross validation, and active learning, show for a weighted ensemble average under squared error:

$$ E = \bar{E} - \bar{A} $$

Where:

  • $E$ — the ensemble's squared error at a point.
  • $\bar{E}$ — the weighted average of the members' individual squared errors.
  • $\bar{A}$ — the ambiguity: the weighted average squared deviation of each member from the ensemble output.

Since $\bar{A} \geq 0$ always, the ensemble is never worse than the average member, and the gain equals the disagreement. This is an identity, not a bound — disagreement is not a proxy for the benefit, it is the benefit. The generalisation to bias-variance-covariance is Ueda and Nakano (1996).

Why single models fail: Dietterich's three reasons

Dietterich (2000), Ensemble methods in machine learning, gives three orthogonal reasons ensembles help:

  1. Statistical — with limited data, many hypotheses fit equally well; averaging them hedges the risk of picking a wrong one.
  2. Computational — training is a local search that gets stuck; different starting points reach different local optima, and averaging approximates the search done right.
  3. Representational — the true function may sit outside any single member's hypothesis class, but inside the span of several.

Current state

On tabular data, ensembles of trees remain the strongest general-purpose models — Grinsztajn, Oyallon and Varoquaux (2022), Why do tree-based models still outperform deep learning on typical tabular data? (NeurIPS). In deep learning, deep ensembles — several networks trained from different initialisations — are a strong baseline for uncertainty estimation (Lakshminarayanan et al., 2017, Simple and scalable predictive uncertainty estimation using deep ensembles).

The two dominant construction strategies get their own lessons: independent members averaged (bagging, Breiman 1996) and sequential members added (boosting, Freund and Schapire 1997).

What to learn next

  • Bagging — the standard recipe for manufacturing diverse models from one dataset.
  • Random forest — bagging plus feature randomness, the workhorse ensemble.
  • XGBoost — the other family: models built in sequence, each fixing the last.

What to learn next

These follow on from what you just read.

  • Ensembles and Gradient Boosting

    Bagging

    Bagging trains many copies of one model on random re-samples of the same data and averages them, trading a twitchy model for a steady one.

  • Ensembles and Gradient Boosting

    Out-of-bag evaluation

    Every bootstrap sample leaves about a third of the rows unseen, and scoring each row only with the trees that never saw it gives you validation for free.

  • Ensembles and Gradient Boosting

    AdaBoost

    AdaBoost trains weak models one at a time and makes every example it got wrong count for more in the next round, which was the first proof that weak learners add up to a strong one.