AI glossary

Random forest

In one sentence A random forest trains many decision trees on different random slices of the data and lets them vote, trading explainability for reliability.

By Updated

A random forest is a crowd of decision trees, each trained on a random slice of the data, whose votes are combined into one prediction.

Ask one person to guess the weight of an ox and you get one biased guess. Ask three hundred and average them, and the crowd famously lands near the truth — individual errors point in different directions and cancel. A random forest manufactures such a crowd out of decision-trees.

The crucial word is random, applied twice to force the trees to disagree. Each tree trains on a random resample of the rows (bootstrapping — sampling with replacement). And at every split, each tree may only consider a random subset of the features. Without this, every tree would latch onto the same strongest signal and the crowd would be three hundred copies of one opinion.

one deep tree   :  memorises its data — high variance
300 varied trees:  errors cancel in the vote — variance shrinks

The result is one of the most dependable algorithms in practice: strong on tabular data with minimal tuning, resistant to overfitting, no feature-scaling needed, and it ranks feature importance for free. The costs: you lose the single tree's print-out explainability, and prediction means running every tree. On tabular benchmarks its main rival is gradient-boosting, which builds its trees sequentially instead of independently and often edges ahead with tuning.

Where to go next