Ensembles and Gradient Boosting
Bagging
Bagging trains many copies of one model on random re-samples of the same data and averages them, trading a twitchy model for a steady one.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Bagging trains many copies of the same model, each on a different random re-sample of the data, then averages their answers.
Picture a cloth bag holding ten numbered chits. You draw a chit, note the number, and put it back. You draw again, note it, put it back — ten draws in total. Your list of ten numbers will repeat some chits and miss others entirely.
That list is called a bootstrap sample: a re-sample of the data, the same size as the original, drawn with replacement. Bagging is short for bootstrap aggregating — make many bootstrap samples, train one model on each, and combine them.
Why this had to be invented
The previous lesson said ensembles need members that disagree. But you have one dataset and one favourite algorithm. Where does the disagreement come from?
Bagging's answer: shake the data. A deep decision tree is twitchy — retrain it on a slightly different sample and it reorganises itself. Bagging turns that twitchiness from a weakness into a supply of diverse committee members.
Each tree sees a different re-sample, so each tree carves the space a little differently. Patterns that are truly in the data show up in most of the trees. Patterns that were noise in one sample fail to repeat, and the averaging washes them away.
How it works
the training data
/ | \
draw with draw with draw with ← each draw repeats some
replacement replacement replacement rows, misses others
| | |
tree 1 tree 2 tree 3 ... tree 200
\ | /
majority vote / average
|
one steadier answerEvery model trains on the same kind of data but not the same data. What survives the averaging is what all the samples agree on.
A real example you have seen
Exit polls on election night. No polling agency asks every voter. Each agency samples different people, and each individual poll swings a few percent. News channels show the poll of polls — an average across agencies — because the average is steadier than any single poll.
And if this recipe sounds familiar: a random forest is bagging applied to trees, plus one extra trick with random features.
Remember this
- A bootstrap sample is drawn from the data with replacement — repeats allowed, gaps expected.
- Bagging = train one model per bootstrap sample, then vote or average.
- It calms twitchy, flexible models. It does nothing for steady ones, or for models already too simple.
What to learn next
- Out-of-bag evaluation — the free validation set hiding inside every bootstrap.
- Random forest — bagging plus feature randomness.
- AdaBoost — the sequential family, where members stop being independent.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnOutputs verified with numpy 1.26.4 and scikit-learn 1.7.2 on CPU.
What a bootstrap sample looks like
import numpy as np
rng = np.random.default_rng(0)
rows = np.arange(10)
sample = rng.choice(rows, size=10, replace=True)
print("bootstrap sample:", sorted(sample.tolist()))
print("distinct rows in it:", len(set(sample.tolist())))bootstrap sample: [0, 0, 0, 1, 2, 3, 5, 6, 8, 8] distinct rows in it: 7
Row 0 appears three times. Rows 4, 7 and 9 never appear. On average a bootstrap sample contains about 63% of the distinct rows — the missing third becomes very useful in the next lesson.
One twitchy tree versus 200 bagged trees
from sklearn.datasets import make_moons
from sklearn.ensemble import BaggingClassifier
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
X, y = make_moons(n_samples=500, noise=0.3, random_state=2)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.4, random_state=2)
tree = DecisionTreeClassifier(random_state=0).fit(Xtr, ytr)
print(f"one deep tree: {tree.score(Xte, yte):.3f}")
bag = BaggingClassifier(DecisionTreeClassifier(random_state=0),
n_estimators=200, random_state=0).fit(Xtr, ytr)
print(f"200 bagged trees: {bag.score(Xte, yte):.3f}")one deep tree: 0.840 200 bagged trees: 0.900
Six percentage points, from the same algorithm on the same data. Nothing about any individual tree improved — each bagged tree is as twitchy as the lone one. The averaging is doing all the work.
The walkthrough
replace=True is the entire bootstrap. Sampling without replacement and using all rows would hand every tree identical data — identical trees, useless vote.
BaggingClassifier(estimator, n_estimators=200) clones the estimator 200 times, draws 200 bootstrap samples, and fits one clone per sample. Prediction is a vote across clones. Add n_jobs=-1 to fit clones on all CPU cores at once — they are independent, so bagging parallelises for free.
Why the deep tree? An unpruned tree has low bias and high variance — it can fit anything, and it changes with every data shake. That is the ideal bagging ingredient. Bagging reduces variance and leaves bias almost alone.
More trees never overfit. Going from 200 to 2,000 trees costs time but does not hurt accuracy — the average only gets steadier. This is unlike boosting, where more rounds can overfit.
Common mistakes
Bagging a stable model. Wrap LogisticRegression in the bagger above and the gain vanishes. A linear model barely changes between bootstrap samples, so its clones agree, and a committee that agrees adds nothing. Bagging pays only for high-variance models.
Using bagging to fix underfitting. If the single model scores badly on training data, it is biased, not twitchy. Averaging many underfit models produces one confident underfit answer. Diagnose with overfitting and underfitting first.
Confusing bagging with boosting. Bagging trains members independently, in parallel, and averages equals. Boosting trains members one after another, each correcting the last. Similar names, opposite mechanics — the XGBoost lesson covers the other family.
Leaving the base tree shallow. max_depth=3 stumps have too little variance for averaging to remove. For bagging, grow trees deep and let the ensemble do the calming.
Try it yourself
Swap DecisionTreeClassifier(random_state=0) for LogisticRegression() in bagging_demo.py and rerun. Predict both numbers first. Then put the tree back and try n_estimators of 5, 20 and 500 — watch where the improvement saturates.
What to learn next
- Out-of-bag evaluation — the free validation set hiding inside every bootstrap.
- Random forest — bagging plus feature randomness.
- AdaBoost — the sequential family, where members stop being independent.
Researcher — Mathematics and papers.
Formal definition
Given training set $D = {(x_i, y_i)}_{i=1}^{n}$, draw $B$ bootstrap samples $D^{(b)}$, each of size $n$ sampled uniformly with replacement from $D$. Train $\hat{f}^{(b)}$ on $D^{(b)}$ and aggregate:
$$ \hat{f}{bag}(x) = \frac{1}{B} \sum{b=1}^{B} \hat{f}^{(b)}(x) $$
Where:
- $D^{(b)}$ — the $b$-th bootstrap sample.
- $B$ — the number of ensemble members.
- $\hat{f}^{(b)}$ — the model fitted to sample $b$; classification uses a plurality vote instead of the mean.
Bagging is Breiman (1996), Bagging predictors, Machine Learning 24. The bootstrap itself is Efron (1979), Bootstrap methods: another look at the jackknife.
The 63.2% property
The probability that a given row is absent from one bootstrap sample of size $n$ is
$$ \left(1 - \frac{1}{n}\right)^{n} \xrightarrow{n \to \infty} e^{-1} \approx 0.368 $$
So each member trains on roughly $63.2\%$ of the distinct rows. The excluded $36.8\%$ are that member's out-of-bag rows, the basis of out-of-bag evaluation.
Why it reduces variance
From the ensemble variance identity, an average of $B$ predictors with individual variance $\sigma^2$ and average correlation $\rho$ has variance $\rho\sigma^2 + \frac{1-\rho}{B}\sigma^2$. Bootstrap resampling lowers $\rho$ below 1 without raising $\sigma^2$ much; increasing $B$ then removes the $(1-\rho)/B$ term entirely. Bias is essentially unchanged, since each member sees a sample from (almost) the training distribution.
The deeper analysis is Bühlmann and Yu (2002), Analyzing bagging, Annals of Statistics: bagging's benefit concentrates on hard-threshold, unstable procedures — exactly decision-tree split points — where it acts as a smoothing operator on the decision function. For linear, stable estimators the smoothing has nothing to smooth, which formalises the "do not bag logistic regression" advice.
Variants
- Pasting — sampling without replacement, smaller-than-$n$ samples (Breiman, 1999). Useful when data is huge and each member should be cheap.
- Subagging — subsample aggregating at rate $n/2$ without replacement; asymptotically similar to bagging at lower cost (Bühlmann and Yu, 2002).
- Random subspaces — resample features instead of rows (Ho, 1998, The random subspace method).
- Random forest — bootstrap rows and restrict each split to a random feature subset, pushing $\rho$ lower still (Breiman, 2001). This is the variant that won: see random forest.
Cost
Training is $B$ times one model, embarrassingly parallel. Prediction is $B$ forward passes, which is the real production cost; distillation into a single model (Bucilă et al., 2006) recovers speed when latency matters — the idea behind knowledge distillation.
What to learn next
- Out-of-bag evaluation — the free validation set hiding inside every bootstrap.
- Random forest — bagging plus feature randomness.
- AdaBoost — the sequential family, where members stop being independent.