Machine Learning

Random forest

A random forest grows hundreds of deliberately different decision trees and lets them vote, turning the twitchiness of a single tree into a steadier answer than any one tree can give.

On this page 8
  1. Why it exists
  2. How the trees are forced to differ
  3. How it works
  4. The free test you get for nothing
  5. Where you have already seen it
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A random forest grows many different decision trees and lets them vote on the answer.

Think of the jar of laddoos at a school fair. You pay ten rupees to guess how many are inside. Almost nobody guesses well. One person says 60, another says 200, and both are badly wrong.

Now add up every guess in the queue and take the average. That average lands remarkably close to the true count, closer than nearly every individual person managed. You have watched this happen. The errors point in different directions, so they cancel each other out.

A random forest is that queue, made of decision trees.

Why it exists

The decision trees lesson ended with a complaint. One tree is twitchy. Change a few rows and it rearranges itself. Let it grow freely and it memorises your data.

For a long time that looked like a defect to be fixed. The insight behind random forests is to stop fixing it and start using it.

If every tree makes different mistakes, then averaging their answers cancels the mistakes out. The part that all the trees agree on is the real pattern. The part they disagree on was noise, and it averages away.

That only works if the trees genuinely differ. This turns out to be the whole design problem.

How the trees are forced to differ

Trained on the same data in the same way, every tree would come out identical. Averaging a hundred copies of one opinion gives you that one opinion back.

So a random forest deliberately handicaps each tree, in two separate ways.

Each tree sees a different sample of the rows. Rows are drawn at random, with repeats allowed, until there are as many as the original. Some rows appear twice, and roughly a third are left out entirely. Every tree therefore grows up on a slightly different version of the data.

Each tree sees only some of the columns at each question. When a tree is choosing its next question, it is shown a random handful of the available clues and must pick from those. It cannot always reach for the strongest one.

That second rule feels wrong the first time you meet it. Why hide the best clue from the tree?

Because otherwise every tree opens with the same question, and from there they largely repeat each other. Hiding the strongest clue some of the time forces different trees to find different routes to the answer. Their disagreement is manufactured on purpose.

How it works

                     your data
                         |
     +-------------+-----+-----+-------------+
     |             |           |             |
  sample 1     sample 2    sample 3   ...  sample 300
  (some rows   (a different (another                 )
   twice, some  mix)         mix)
   missing)
     |             |           |             |
   tree 1       tree 2      tree 3        tree 300
     |             |           |             |
   "scam"       "safe"      "scam"        "scam"
     |             |           |             |
     +-------------+-----+-----+-------------+
                         |
                    count the votes
                         |
              "scam"  (271 of 300 trees agree)

Predicting a number instead of a group changes only the last step. Take the average of the trees rather than the majority.

That vote count carries real information. All three hundred trees agreeing is a different situation from a hundred and sixty of them agreeing. Both still come back as "scam".

The free test you get for nothing

Remember that about a third of the rows are left out of each tree's sample.

For any single row, roughly a third of the trees never saw it during training. Those trees can be asked about that row as though it were fresh data.

Doing this for every row gives you a score on data the model never trained on, without setting anything aside. It is called the out-of-bag score, and it comes free with the fitting. Few methods offer anything like it.

Where you have already seen it

  • Bank fraud scoring, where forests remain a workhorse for tabular data.
  • Loan and insurance risk models, often alongside a simpler model kept for explaining decisions.
  • Kinect body tracking, which used a forest to label body parts in depth images in real time.
  • Medical risk tools built from patient records with many columns.

The honest part

You lose the thing that made trees special. One tree could be printed and read. Three hundred trees cannot. You gain accuracy and you spend readability, and that is a genuine trade, not a free lunch.

It still cannot see past its data. Averaging trees that all flatline outside the training range gives you a flatline outside the training range. Forests inherit this weakness completely.

The built-in importance scores mislead. Random forests report which columns mattered, and that report is biased in a specific way you will measure in the Developer section. People quote those numbers in meetings constantly, and they are frequently wrong.

More trees never hurts accuracy, but it is not free. Doubling the trees doubles the memory and the prediction time. The gains flatten out quickly.

Remember this

  • A forest is many trees, each grown on a different random sample of rows and columns, voting together.
  • The disagreement between trees is the point, not a side effect — identical trees would gain you nothing.
  • You trade the readability of one tree for accuracy and steadiness.

What to learn next

  • XGBoost — building trees in sequence to fix each other's errors, instead of in parallel.
  • Model evaluation — reading the vote counts a forest gives you.
  • Feature engineering — the work a forest cannot do for you.

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

All five files below run in a few seconds on a laptop. Nothing is downloaded.

One tree against three hundred

The data: 400 patients, eight measurements each. Only the first three measurements mean anything — the other five are pure noise, which is exactly the situation real tabular data puts you in.

one_vs_many.py
import numpy as np
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(1)
X = rng.uniform(0, 10, size=(400, 8))
# Only the first three columns mean anything. The other five are noise.
signal = 0.6 * X[:, 0] + 0.9 * X[:, 1] - 0.7 * X[:, 2]
truth = (signal > np.median(signal)).astype(int)
# Eight percent of the labels were recorded wrongly
y = np.where(rng.random(400) < 0.08, 1 - truth, truth)

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.4, random_state=0, stratify=y)

one = DecisionTreeClassifier(random_state=0).fit(X_train, y_train)
forest = RandomForestClassifier(n_estimators=300, random_state=0).fit(X_train, y_train)
print(f"one tree   train {one.score(X_train, y_train):.3f}   test {one.score(X_test, y_test):.3f}")
print(f"300 trees  train {forest.score(X_train, y_train):.3f}   test {forest.score(X_test, y_test):.3f}")
print()
for n in (1, 3, 10, 30, 100, 300):
    f = RandomForestClassifier(n_estimators=n, random_state=0).fit(X_train, y_train)
    print(f"{n:4d} trees  test {f.score(X_test, y_test):.3f}")
Output
one tree   train 1.000   test 0.800
300 trees  train 1.000   test 0.831

   1 trees  test 0.650
   3 trees  test 0.738
  10 trees  test 0.750
  30 trees  test 0.794
 100 trees  test 0.825
 300 trees  test 0.831

Read the sweep from the top, because it contains a surprise.

A forest of one tree scores 0.650 — much worse than the single tree's 0.800. That is not an error. A forest's tree is handicapped: it sees a bootstrap sample instead of all the rows, and only a subset of columns at each split. Individually, it is a worse tree.

Yet 300 of these worse trees score 0.831, beating the one good tree.

That is the entire method in two lines of output. The forest does not build better trees. It builds worse trees that disagree, and harvests the disagreement.

Notice too that the gains flatten: 100 to 300 trees buys 0.006. More trees never hurt accuracy, but past a point you are paying memory and latency for nothing.

Proving the disagreement is what does the work

If the theory above is right, then switching the randomness off should destroy the benefit. bootstrap=False gives every tree all the rows. max_features=None gives every tree every column at every split.

why_it_works.py
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(1)
X = rng.uniform(0, 10, size=(400, 8))
signal = 0.6 * X[:, 0] + 0.9 * X[:, 1] - 0.7 * X[:, 2]
truth = (signal > np.median(signal)).astype(int)
y = np.where(rng.random(400) < 0.08, 1 - truth, truth)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.4, random_state=0, stratify=y)

for tag, kw in [("randomness ON ", dict()),
                ("randomness OFF", dict(bootstrap=False, max_features=None))]:
    f = RandomForestClassifier(n_estimators=50, random_state=0, **kw).fit(X_train, y_train)
    # ask every tree in the forest separately
    votes = np.array([t.predict(X_test) for t in f.estimators_])
    # how often do two trees, picked at random, disagree about the same patient?
    disagree = (votes[:, None, :] != votes[None, :, :]).mean()
    member = np.mean([(v == y_test).mean() for v in votes])
    print(f"{tag}  trees disagree {disagree:.3f}   "
          f"average single tree {member:.3f}   whole forest {f.score(X_test, y_test):.3f}")
Output
randomness ON   trees disagree 0.385   average single tree 0.660   whole forest 0.819
randomness OFF  trees disagree 0.016   average single tree 0.780   whole forest 0.775

This is the most important output in the lesson. Read the two rows against each other.

With randomness on: the trees disagree on 38.5 percent of patients. Each tree on its own is poor, averaging 0.660. The forest reaches 0.819 — about sixteen points better than its own average member.

With randomness off: the trees agree on 98.4 percent of patients. Each tree is individually better, averaging 0.780. The forest reaches 0.775 — no better than a member, and slightly worse.

Fifty strong trees that agree are worth less than fifty weak trees that argue. Averaging can only remove errors that differ between the things being averaged.

So the randomisation is not overhead to be switched off when you want more accuracy. It is the mechanism that produces the accuracy.

How much randomisation, though, is a real decision. Weakening the trees buys disagreement and costs individual skill, and the best setting is a balance between the two. The "Try it yourself" section at the end of this block puts a number on that trade-off, and the answer may not be the one you expect.

The free validation score

oob.py
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(1)
X = rng.uniform(0, 10, size=(400, 8))
signal = 0.6 * X[:, 0] + 0.9 * X[:, 1] - 0.7 * X[:, 2]
truth = (signal > np.median(signal)).astype(int)
y = np.where(rng.random(400) < 0.08, 1 - truth, truth)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.4, random_state=0, stratify=y)

# oob_score=True scores each row using only the trees that never saw it
f = RandomForestClassifier(n_estimators=300, oob_score=True, random_state=0).fit(X_train, y_train)
print("out-of-bag score:", round(f.oob_score_, 3))
print("held-out test   :", round(f.score(X_test, y_test), 3))
Output
out-of-bag score: 0.825
held-out test   : 0.831

0.825 against 0.831. The out-of-bag estimate came from the training call itself, at no extra cost, and it landed within a point of the real held-out score.

Two cautions. It needs enough trees to be stable — with 10 trees, some rows are out-of-bag for barely any of them. And it does not replace a genuine test set when you are also tuning hyperparameters, for the reasons in train, test and validation splits.

The importance scores that fool people

A forest will happily tell you which columns mattered. We know the ground truth here: columns 0, 1 and 2 are real, and columns 3 to 7 are noise we generated ourselves.

importance.py
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.inspection import permutation_importance

rng = np.random.default_rng(1)
X = rng.uniform(0, 10, size=(400, 8))
signal = 0.6 * X[:, 0] + 0.9 * X[:, 1] - 0.7 * X[:, 2]
truth = (signal > np.median(signal)).astype(int)
y = np.where(rng.random(400) < 0.08, 1 - truth, truth)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.4, random_state=0, stratify=y)

forest = RandomForestClassifier(n_estimators=300, random_state=0).fit(X_train, y_train)
names = [f"col{i}" for i in range(8)]

print("useful columns are col0, col1, col2 - the rest are pure noise")
print()
print("MDI (the built-in one):")
for n, v in zip(names, forest.feature_importances_):
    print(f"   {n}  {v:.3f}")

# permutation importance shuffles one column and measures the damage, on held-out data
perm = permutation_importance(forest, X_test, y_test, n_repeats=20, random_state=0)
print("permutation (measured on held-out data):")
for n, v in zip(names, perm.importances_mean):
    print(f"   {n}  {v:+.3f}")

print()
print("noise columns get", round(forest.feature_importances_[3:].sum(), 3),
      "of the MDI total, but", round(perm.importances_mean[3:].sum(), 3), "of the permutation total")
Output
useful columns are col0, col1, col2 - the rest are pure noise

MDI (the built-in one):
   col0  0.150
   col1  0.269
   col2  0.184
   col3  0.085
   col4  0.087
   col5  0.069
   col6  0.076
   col7  0.080
permutation (measured on held-out data):
   col0  +0.111
   col1  +0.222
   col2  +0.107
   col3  +0.005
   col4  +0.003
   col5  +0.002
   col6  +0.002
   col7  -0.006

noise columns get 0.397 of the MDI total, but 0.005 of the permutation total

Look at what feature_importances_ did. It handed 40 percent of the total importance to five columns of pure random noise. col4 scored 0.087, which is more than half of what genuinely-useful col0 received.

If you had presented that table to a doctor, you would have told them five meaningless measurements are worth collecting.

Permutation importance gets it right. It shuffles one column on held-out data and measures how much the score falls. The five noise columns collectively cost 0.005 — indistinguishable from zero, with one even slightly negative, which is what noise looks like.

The cause is structural. feature_importances_ counts how much each column reduced impurity on the training data, and a noise column with many distinct values always offers plenty of split points that happen to help on the rows in front of it. It is measuring opportunity, not usefulness.

Use permutation_importance on held-out data. This single habit will save you from more wrong conclusions than any other change in this section.

Forests still cannot extrapolate

no_extrapolation.py
import numpy as np
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import RandomForestRegressor

area = np.array([[6.0], [7.0], [8.0], [9.0], [10.0], [11.0], [12.0], [13.0]])
price = np.array([40.0, 46.0, 52.0, 58.0, 63.0, 70.0, 75.0, 82.0])

dt = DecisionTreeRegressor(max_depth=2).fit(area, price)
rf = RandomForestRegressor(n_estimators=200, random_state=0).fit(area, price)

outside = np.array([[16.0], [20.0], [30.0]])
print("area          :", outside.ravel())
print("one tree says :", dt.predict(outside).round(1))
print("forest says   :", rf.predict(outside).round(1))
Output
area          : [16. 20. 30.]
one tree says : [78.5 78.5 78.5]
forest says   : [79.3 79.3 79.3]

Two hundred trees, and a 3000 square foot flat is still priced identically to a 1600 square foot one.

Averaging things that all flatline gives you a flatline. If your target has a trend that continues past the range you observed, no tree-based method will follow it. Use a linear model, or model the trend and fit the forest to what is left over.

Common mistakes

Quoting feature_importances_. Demonstrated above with 40 percent of importance on noise. Use permutation_importance on held-out data.

Tuning n_estimators as if it were a capacity knob. More trees do not overfit a forest; the curve flattens and stays flat. Set it as high as your latency and memory allow, then tune max_features, min_samples_leaf and max_depth, which do control capacity.

Leaving class_weight alone on imbalanced data. Each tree votes by majority within its leaves, so a rare class can be outvoted everywhere. Pass class_weight="balanced_subsample".

Forgetting n_jobs=-1. The trees are independent, so training parallelises almost perfectly. The default of n_jobs=None uses one core. This is free speed you are leaving on the table.

Expecting a forest to rescue bad features. A forest reduces variance. It does not invent information. If none of your columns carry signal, 500 trees agree confidently on nothing.

Assuming a forest cannot overfit. It resists overfitting far better than one tree, and it is not immune. With very noisy labels and deep unconstrained trees, the training score of 1.000 above is a reminder that every tree still memorised its own sample.

Reaching for a forest first on tabular data. Gradient boosting usually scores higher. A forest is easier to tune and harder to get badly wrong, which makes it an excellent baseline — see XGBoost for the alternative.

Try it yourself

In why_it_works.py, add a third row to the loop that keeps the bootstrap but hands every tree all eight columns:

why_it_works.py (edit)
for tag, kw in [("randomness ON ", dict()),
                ("rows only     ", dict(max_features=None)),
                ("randomness OFF", dict(bootstrap=False, max_features=None))]:

Before running it, predict where the middle row will land. It keeps the row sampling but drops the column sampling, so it has one source of disagreement instead of two.

Output
randomness ON   trees disagree 0.385   average single tree 0.660   whole forest 0.819
rows only       trees disagree 0.285   average single tree 0.723   whole forest 0.844
randomness OFF  trees disagree 0.016   average single tree 0.780   whole forest 0.775

The disagreement landed in the middle, as expected. The forest score did not — the middle row is the best of the three, at 0.844.

Sit with that, because it corrects an over-simple reading of this lesson. More disagreement is not automatically better. Going from the middle row to the top row bought 0.100 more disagreement and cost 0.063 of individual tree skill, and that trade came out negative.

With only eight columns, of which three matter, the default of roughly two columns per split starves the trees. Many splits get offered nothing but noise columns.

The lesson to carry away is the shape of the trade, not a preferred setting. max_features is the knob that moves it, and it is worth tuning on any dataset with few columns. On data with hundreds of columns, the default is usually close to right.

What to learn next

  • XGBoost — building trees in sequence to fix each other's errors, instead of in parallel.
  • Model evaluation — reading the vote counts a forest gives you.
  • Feature engineering — the work a forest cannot do for you.

Researcher — Mathematics and papers.

Definition

A random forest (Breiman, 2001) is bagging with randomised split selection. Given training set D of size n:

for b = 1..B:
    D_b  <- bootstrap sample of size n drawn from D with replacement
    T_b  <- grow a CART tree on D_b, where at each node a random subset
            of m <= d features is considered for the split
    (trees are grown deep; typically unpruned)

Classification:  f(x) = majority vote over { T_b(x) }
Regression:      f(x) = (1/B) * SUM_b T_b(x)
  • B — number of trees, n — training set size, d — number of features
  • m — features sampled per split; mtry in the R literature. Defaults: sqrt(d) for classification, d/3 (R) or d (scikit-learn RandomForestRegressor) for regression

The two randomisation sources are separable, and this matters conceptually. Bootstrap alone gives Breiman's earlier bagging (1996). Feature subsampling alone gives Ho's random subspace method (1998). The combination is the random forest.

Why averaging works

For regression with squared loss, consider B predictors each with variance sigma^2 and pairwise correlation rho. The variance of their average is:

Var( (1/B) SUM_b T_b )  =  rho * sigma^2  +  ( (1 - rho) / B ) * sigma^2
  • rho — average pairwise correlation between tree predictions
  • sigma^2 — variance of an individual tree

This single expression explains the whole design.

The second term vanishes as B -> infinity. The first does not. The irreducible floor is rho * sigma^2, so the only way to improve a large forest is to reduce the correlation between its trees.

That is exactly what feature subsampling does. Lowering m reduces rho while raising sigma^2, since each tree is individually weaker. m trades these against each other, and the optimum is typically well below d.

The empirical numbers in the Developer section are this formula made visible: disagreement 0.385 with member accuracy 0.660 beats disagreement 0.016 with member accuracy 0.780.

Bagging alone reduces variance but leaves bias essentially unchanged, which is why the base learners are grown deep. Deep trees are low-bias and high-variance — precisely the failure mode averaging repairs.

Generalisation bound

Breiman's bound is stated in terms of margin:

PE*  <=  rho_bar * (1 - s^2) / s^2
  • PE* — generalisation error of the forest
  • s — the strength, the expected margin of individual trees
  • rho_bar — mean correlation between the margin functions of tree pairs

The bound is loose and rarely predictive of actual error. Its value is qualitative and it is the correct intuition: error falls with strength and rises with correlation.

The bound also implies forests do not overfit as B grows. As B -> infinity, PE converges almost surely to a limit — a consequence of the strong law of large numbers, since the trees are i.i.d. given the training set. This is a genuine and unusual guarantee: B is a computational budget, not a capacity parameter. Contrast boosting, where more rounds do eventually overfit.

Out-of-bag estimation

Under bootstrap sampling, the probability a given observation is omitted from a given tree is:

(1 - 1/n)^n  ->  1/e  ~=  0.368   as n -> infinity

So each tree omits roughly 36.8 percent of the data, and each observation is out-of-bag for roughly 36.8 percent of the trees. Aggregating the predictions of only those trees gives a nearly unbiased estimate of generalisation error.

Bylander (2002) and others note that OOB error is mildly pessimistic, since each observation is scored by an ensemble of about 0.368 * B trees rather than the full B. The bias shrinks as B grows.

OOB is not a substitute for a held-out set when hyperparameters are selected using it — the selection bias of Cawley & Talbot (2010) applies unchanged.

Variable importance, and why the default is wrong

Mean decrease in impurity (MDI). Sum the weighted impurity decrease over all nodes splitting on a feature, averaged over trees. Cheap, computed during fitting, and biased.

Strobl et al. (2007) identify two mechanisms. Features with more distinct values offer more candidate split points and win ties more often, so continuous and high-cardinality categorical features are inflated even under the null. And because impurity is measured on in-bag data, the bias is an in-sample overfitting artefact.

The Developer section quantifies this on a dataset with known ground truth: five pure-noise columns absorbed 39.7 percent of the MDI total.

Permutation importance (MDA). Shuffle feature j and measure the increase in error, ideally on held-out data:

Imp(j) = Err( X with column j permuted ) - Err( X )

Unbiased with respect to cardinality, but it has its own failure mode: under correlated features it evaluates the model on inputs off the data manifold, and importance is split arbitrarily among a correlated group. Conditional permutation importance (Strobl et al., 2008) permutes within strata of correlated features and addresses this.

SHAP. TreeSHAP (Lundberg et al., 2020) computes exact Shapley values for tree ensembles in O(TLD^2), giving local per-prediction attributions with additivity guarantees. It is the current default for serious attribution work on trees.

Consistency theory

Breiman's original algorithm resisted analysis for a decade, because the splits depend on the labels in a complicated way.

Biau (2012) proved consistency for a simplified variant. Scornet, Biau & Vert (2015) established the first consistency result for the original algorithm under an additive regression model. Wager & Athey (2018) proved asymptotic normality for honest forests — where the sample used to choose splits is disjoint from the sample used to fit leaf values — enabling valid confidence intervals and forming the basis of causal forests.

Honesty is not a technicality. Without it, the same data determines both the partition and the estimates within it, and the resulting intervals are invalid.

Cost

Training     O( B * m * n log^2 n )     embarrassingly parallel over B
Prediction   O( B * depth )             typically O(B log n)
Memory       O( B * number of nodes )

Training parallelises with near-linear speedup, since trees are independent. Prediction is B times a single tree's cost, which is the real deployment constraint: a 500-tree forest is 500 tree traversals per prediction.

Variants

Extremely randomised trees (Geurts et al., 2006) push randomisation further: split thresholds are drawn at random rather than optimised, and the full sample is used by default rather than a bootstrap. This lowers rho and variance further at some cost in bias, and it is substantially faster since no threshold search is performed. sklearn.ensemble.ExtraTreesClassifier. Frequently competitive with, occasionally better than, a random forest — and consistently under-tried.

Quantile regression forests (Meinshausen, 2006) keep the full empirical distribution of training values in each leaf rather than only the mean, yielding prediction intervals rather than point estimates.

Causal forests (Wager & Athey, 2018; Athey et al., 2019) adapt the splitting criterion to maximise heterogeneity in treatment effect rather than in outcome, estimating conditional average treatment effects with valid inference.

Isolation forests (Liu et al., 2008) invert the idea for anomaly detection: anomalies are isolated by fewer random splits, so average path length becomes an outlier score.

Key references

  • Ho, T. K. (1998). The Random Subspace Method for Constructing Decision Forests. IEEE TPAMI 20(8).
  • Breiman, L. (1996). Bagging Predictors. Machine Learning 24(2).
  • Breiman, L. (2001). Random Forests. Machine Learning 45(1). The founding paper.
  • Geurts, P., Ernst, D. & Wehenkel, L. (2006). Extremely Randomized Trees. Machine Learning 63(1).
  • Meinshausen, N. (2006). Quantile Regression Forests. JMLR 7.
  • Strobl, C. et al. (2007). Bias in Random Forest Variable Importance Measures. BMC Bioinformatics 8(25).
  • Biau, G. (2012). Analysis of a Random Forests Model. JMLR 13.
  • Scornet, E., Biau, G. & Vert, J.-P. (2015). Consistency of Random Forests. Annals of Statistics 43(4).
  • Wager, S. & Athey, S. (2018). Estimation and Inference of Heterogeneous Treatment Effects using Random Forests. JASA 113(523).
  • Fernández-Delgado, M. et al. (2014). Do We Need Hundreds of Classifiers to Solve Real World Classification Problems? JMLR 15.

Current state

Fernández-Delgado et al. (2014) compared 179 classifiers across 121 datasets and found random forests the strongest family overall. That result predates the maturity of modern gradient boosting; today tuned GBMs generally edge ahead on tabular benchmarks.

Random forests retain three practical advantages that keep them in production. They are close to hyperparameter-free — defaults are usually within a point of tuned performance, which is not remotely true of boosting. They parallelise with near-linear speedup. And OOB error removes the need for a separate validation split during development.

Grinsztajn et al. (2022) confirm that tree ensembles as a family still outperform deep learning on medium-sized tabular data, attributing this to robustness against uninformative features and to an inductive bias suited to irregular, non-smooth target functions. The eight-column experiment in the Developer section is a small instance of exactly that.

The honest summary: use a random forest as the baseline that is difficult to get wrong, then check whether gradient boosting earns the extra tuning effort.

What to learn next

  • XGBoost — building trees in sequence to fix each other's errors, instead of in parallel.
  • Model evaluation — reading the vote counts a forest gives you.
  • Feature engineering — the work a forest cannot do for you.