AI glossary

Gradient boosting

In one sentence Gradient boosting builds trees one after another, each trained to correct the combined errors of all the trees before it.

By Updated

Gradient boosting builds an ensemble sequentially: each new tree is trained on the mistakes the current ensemble is still making.

Picture a proofreading chain for a manuscript. The first reader fixes what she can. The second reads her corrected copy and fixes what she missed. The third reads theirs. No single reader is brilliant, but each one works only on the remaining errors, so the chain converges on a clean text. Gradient boosting is that chain built from small decision-trees.

Concretely: fit a weak tree, compute the errors (residuals), fit the next tree to those errors, add its output scaled by a learning-rate, and repeat hundreds of times. The "gradient" in the name is real — each round is a gradient-descent step, taken in the space of functions, on your chosen loss-function.

The contrast with random-forest is the heart of it: forests build independent trees in parallel and average away variance; boosting builds dependent trees in sequence and eats away bias. Boosting usually wins on accuracy but is easier to overfit and has more knobs — tree depth, number of rounds, learning rate.

The famous implementations are XGBoost, LightGBM and CatBoost. On medium-sized tabular data they remain the strongest practical choice, still routinely beating neural networks in competitions and production.

Where to go next