Ensembles and Gradient Boosting

Out-of-bag evaluation

Every bootstrap sample leaves about a third of the rows unseen, and scoring each row only with the trees that never saw it gives you validation for free.

Read these first

On this page 5
  1. Why this had to be invented
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Out-of-bag evaluation scores a bagged model by letting each row be judged only by the trees that never trained on it.

Think of a school competition where several teachers act as judges. There is one firm rule: a teacher never marks their own student. For each contestant, only the unconnected judges score, so every mark is a fair mark.

Bagging creates this situation automatically. Each tree trains on a bootstrap sample, and every bootstrap sample misses about a third of the rows. For any row, the trees that missed it are its unconnected judges.

Those missed rows are called out-of-bag rows — "not in that tree's bag". The score built from them is the OOB score.

Why this had to be invented

Honest evaluation normally costs data. You hold out a test set, and those rows never help the model learn. With small datasets that hurts — every row you reserve for judging is a row not teaching.

Out-of-bag evaluation dodges the cost. The random gaps in each tree's training sample already exist. Reuse them, and you get a fair score while training on all of your data. Nothing is held back, yet no row is ever judged by a tree that saw it.

How it works

              tree 1   tree 2   tree 3   tree 4   tree 5
row 17 seen?    yes      no       yes      no       yes

 → only tree 2 and tree 4 may vote on row 17
 → compare their combined vote with row 17's true answer

repeat for every row  →  OOB accuracy

Each row gets judged by roughly a third of the trees. Across hundreds of trees and all the rows, that is plenty of votes for a stable score.

A real example you have seen

Any situation where the connected judge steps aside. A professor does not review their own student's paper. A referee with a cousin on the team hands the whistle to someone else. The principle is identical: the evaluator must not have helped produce the thing being evaluated.

Remember this

  • Each bagged tree misses about a third of the rows — its out-of-bag rows.
  • Each row is scored only by trees that never saw it, so the score is honest.
  • You train on all the data and still get validation — no held-out set spent.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU. Runs in a few seconds.

OOB score versus cross-validation

The claim to verify: the free OOB score lands close to a 5-fold cross-validation score that costs five extra model fits.

oob_demo.py
from sklearn.datasets import make_classification
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score

X, y = make_classification(n_samples=600, n_features=12, n_informative=5,
                           flip_y=0.05, random_state=3)

forest = RandomForestClassifier(n_estimators=300, oob_score=True, random_state=3)
forest.fit(X, y)
print("OOB accuracy:", round(forest.oob_score_, 3))

fresh = RandomForestClassifier(n_estimators=300, random_state=3)
cv = cross_val_score(fresh, X, y, cv=5)
print("5-fold CV accuracy:", round(cv.mean(), 3))
Output
OOB accuracy: 0.908
5-fold CV accuracy: 0.897

One percentage point apart. The OOB number came from a single training run; the cross-validation number needed five more.

How big is "out of bag", exactly?

oob_fraction.py
n = 600
print("chance a row misses one bootstrap sample:", round((1 - 1/n)**n, 3))
Output
chance a row misses one bootstrap sample: 0.368

Each row is out-of-bag for about 36.8% of the trees. With 300 trees, each row is judged by roughly 110 of them.

The walkthrough

oob_score=True tells the forest to track, for every row, the votes of the trees whose bootstrap sample missed that row, and to report the resulting accuracy as forest.oob_score_. The cost is bookkeeping, not extra training.

Why the two numbers differ slightly. Each row's OOB vote uses about 110 of the 300 trees, not all 300 — a smaller committee, so a slightly noisier and usually slightly pessimistic verdict. The CV estimate instead trains on 80% of the rows per fold. Both are honest; neither is exact.

oob_score_ needs enough trees. With 20 trees, each row is judged by about 7 — a tiny, noisy jury. OOB scores stabilise as the forest grows. If your OOB score jumps around between reruns, add trees before concluding anything.

Common mistakes

Tuning on the OOB score, then reporting it. The moment you pick settings because they maximised OOB accuracy, that number becomes an optimistic advert for the winner. Tune on OOB freely — but report final performance on data no decision ever touched. The same trap exists for cross-validation, and the fix is the same.

Too few trees for a stable score. n_estimators=50 leaves each row with a jury of ~18 trees. Comparing two models on scores that noisy is reading tea leaves. Use a few hundred trees before trusting small differences.

Expecting OOB with bootstrap=False. Turn off bootstrap sampling and there are no out-of-bag rows — every tree saw everything. scikit-learn raises an error rather than reporting a meaningless score.

Using OOB where rows are not independent. Duplicate or near-duplicate rows (the same customer twice, overlapping time windows) leak between a tree's bag and its out-of-bag set, and OOB turns optimistic. The judge went to school with the contestant. Grouped or time-aware splits are the fix — see forecast evaluation for the time-series version.

Try it yourself

Run oob_demo.py with n_estimators set to 20, then 100, then 1000, each with three different random_state values. Watch the spread of oob_score_ shrink as trees are added. Find the point where the spread is smaller than 0.01 — that is roughly where the OOB score becomes worth quoting.

What to learn next

Researcher — Mathematics and papers.

Definition

For a bagged ensemble ${\hat{f}^{(b)}}_{b=1}^{B}$ with bootstrap samples ${D^{(b)}}$, the OOB prediction for row $i$ aggregates only the members that excluded it:

$$ \hat{f}_{oob}(x_i) = \operatorname{vote}\left{ \hat{f}^{(b)}(x_i) : (x_i, y_i) \notin D^{(b)} \right} $$

Where:

  • $B$ — ensemble size.
  • $D^{(b)}$ — the $b$-th bootstrap sample.
  • $\operatorname{vote}$ — plurality for classification, mean for regression.

The OOB error is the empirical error of $\hat{f}_{oob}$ over all $n$ rows. The construction is Breiman (1996), Out-of-bag estimation (technical report, Berkeley), building on Efron and Tibshirani's leave-one-out bootstrap.

Expected jury size

Row $i$ is out of bag for member $b$ with probability $(1 - 1/n)^n \to e^{-1}$. So the expected number of judging members is $B e^{-1} \approx 0.368 B$, and each OOB prediction behaves like a bagged ensemble of size $0.368B$ rather than $B$.

Bias properties

Two opposing biases, both usually small:

  1. Pessimistic — each OOB vote uses a $\sim 0.37 B$ sub-ensemble, weaker than the full ensemble it is meant to assess. The gap vanishes as $B$ grows, since bagged performance saturates in $B$.
  2. Optimistic under dependence — duplicated or clustered rows place near-copies of a test row inside the judging trees' training bags, inflating the score. This is a data-dependence failure, not a flaw in the estimator.

Empirical comparisons: OOB error tracks cross-validation closely for forests of adequate size (Breiman 2001; Janitza and Hornung 2018, On the overestimation of random forest's out-of-bag error, who document the classification-specific pathologies — notably with strongly unbalanced classes and small $n$).

What OOB rows are also used for

  • Permutation importance — permute feature $j$ within each tree's OOB rows and measure the vote degradation (Breiman 2001). This is the honest importance covered in feature importance done right.
  • Calibration-style diagnostics — OOB class-probability estimates give near-unbiased inputs for reliability curves without a held-out set.
  • Generalised random forests — OOB ("honest") predictions are central to the inference machinery of Wager and Athey (2018), Estimation and inference of heterogeneous treatment effects using random forests (JASA), where trees must not be scored or used on rows that influenced their splits.

Cost comparison

EstimatorExtra model fitsData spent
Held-out test set020–30% of rows
$k$-fold CV$k$none, but $k\times$ training time
OOB0none

OOB's price is that it exists only for bootstrap ensembles, and its jury-size pessimism at small $B$.

What to learn next

What to learn next

These follow on from what you just read.

  • Ensembles and Gradient Boosting

    AdaBoost

    AdaBoost trains weak models one at a time and makes every example it got wrong count for more in the next round, which was the first proof that weak learners add up to a strong one.

  • Ensembles and Gradient Boosting

    LightGBM

    LightGBM makes gradient boosting fast by sorting feature values into coarse bins and growing trees leaf by leaf, which is why it dominates on large tables.

  • Ensembles and Gradient Boosting

    CatBoost

    CatBoost feeds text-like category columns straight into gradient boosting, and its ordered trick stops a model from secretly grading its own answers.