Ensembles and Gradient Boosting

Tuning gradient-boosted trees

Almost all gradient boosting tuning reduces to one trade — a smaller learning rate with more trees, cut off by early stopping — plus a short list of knobs in priority order.

On this page 5
  1. Why this had to be invented
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Tuning a gradient-boosted model is mostly one decision: cook on a lower flame for longer, and let a timer tell you when to stop.

Anyone who has made dal knows the trade. A high flame finishes fast and scorches easily. A low flame takes longer and forgives mistakes. The dish is better slow — up to a point, after which more simmering changes nothing.

In gradient boosting, the flame is the learning rate: how big a correction each new tree applies. The cooking time is the number of trees. And the timer is early stopping: watch a side dish of held-out data, and stop when it stops improving.

Why this had to be invented

Boosted trees have an intimidating settings page — twenty-plus knobs in every library. Beginners either freeze, or worse, search all of them at once and burn a week of compute learning nothing.

The escape is knowing that the knobs are not equals. One trade dominates. A short list matters after it. The rest is decoration on most problems.

How it works

Each tree adds a correction to the running answer. The learning rate scales that correction down.

learning rate 0.3   →  ██            14 corrections, each large
learning rate 0.1   →  ██████        59 corrections, medium
learning rate 0.03  →  ████████████  245 corrections, each tiny  ← best score

Big corrections overshoot — like turning a steering wheel in jerks. Small corrections follow the road closely but need more of them. Early stopping removes the guesswork about "how many": ask for far too many trees, and stop when the held-out score flattens.

After that trade, the knobs in order of importance:

1. flame + timer     learning rate, with early stopping
2. tree size         how much each tree can memorise
3. minimum leaf size how few examples may end up in one decision
4. data sampling     show each tree a random portion of rows/columns

A real example you have seen

Every leaderboard-topping tabular model — fraud detection, credit scoring, competition winners — went through exactly this loop. Teams that win are rarely using secret algorithms. They are using this checklist, patiently, with honest validation.

Remember this

  • Lower learning rate + more trees + early stopping is the core recipe.
  • Tree size (leaves or depth) is the second knob: it controls memorisation per tree.
  • Never hand-pick the number of trees — let early stopping cut it.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install lightgbm

Outputs verified with lightgbm 4.7.0 and scikit-learn 1.7.2 on CPU; the script runs in a few seconds. The same logic applies verbatim to XGBoost and CatBoost — only parameter names change.

The flame experiment

Same data, same model, three learning rates. Early stopping picks each one's tree count.

flame.py
import lightgbm as lgb
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=6000, n_features=20, n_informative=10,
                           flip_y=0.1, random_state=11)
Xtr, Xval, ytr, yval = train_test_split(X, y, test_size=0.3, random_state=11)

for lr in (0.3, 0.1, 0.03):
    model = lgb.LGBMClassifier(n_estimators=5000, learning_rate=lr,
                               num_leaves=31, random_state=11, verbose=-1)
    model.fit(Xtr, ytr, eval_X=Xval, eval_y=yval,
              callbacks=[lgb.early_stopping(100, verbose=False)])
    print(f"lr={lr:<5} kept {model.best_iteration_:4d} trees, "
          f"val accuracy {model.score(Xval, yval):.3f}")
Output
lr=0.3   kept   14 trees, val accuracy 0.873
lr=0.1   kept   59 trees, val accuracy 0.881
lr=0.03  kept  245 trees, val accuracy 0.887

Each halving-ish of the learning rate roughly quadrupled the trees kept and nudged accuracy up. That pattern — slower is slightly better, at proportionally higher cost — is the norm, not a quirk of this dataset.

The walkthrough

n_estimators=5000 is not a real request. It is a ceiling that early stopping will never hit. Tuning n_estimators by grid search while the learning rate sits fixed is redundant work — the callback finds the right count for every learning rate automatically.

Why not learning rate 0.001? Follow the trend and the returns shrink while cost explodes: ten times the trees for a fraction of a point. Common practice: explore with 0.1, do final training at 0.03–0.05, and go lower only when a competition decimal pays for the compute.

The knobs after the flame, with names across libraries:

knobLightGBMXGBoosttypical range
tree sizenum_leavesmax_depth16–256 leaves / depth 3–8
min rows per leafmin_child_samplesmin_child_weight5–100
row samplingsubsample + subsample_freqsubsample0.6–1.0
column samplingcolsample_bytreecolsample_bytree0.6–1.0
L2 strengthreg_lambdalambda0–10

Search those with random search or Optuna rather than by hand — while the learning rate stays fixed at your exploration value and early stopping stays on.

Common mistakes

Tuning against the test set. Every knob you adjust because a score went up spends that score's honesty. Keep three pools: training data, a validation set for early stopping and tuning, and a final test set touched once.

Early stopping on training loss. Training loss falls forever; the callback would never fire usefully. The eval data must be data the trees never trained on.

Turning every knob at once. Change five settings, watch the score move, learn nothing about why. Fix the flame, tune tree size; fix that, tune sampling. One variable at a time is slow but compounds; a proper search strategy automates it honestly.

Trusting one validation split for small data. With a few thousand rows, a lucky split lies. Use k-fold cross-validation for the final comparison of candidate settings.

Try it yourself

Add num_leaves values 8 and 128 as an inner loop to flame.py at lr=0.1. Before running, predict which combination overfits (hint: watch the gap between training and validation accuracy — add a model.score(Xtr, ytr) print). Then check how many trees early stopping keeps for big trees versus small ones.

What to learn next

Researcher — Mathematics and papers.

Shrinkage as regularisation

The boosted model after $M$ rounds is $F_M(x) = \sum_{m=1}^{M} \nu \, f_m(x)$ where $\nu \in (0, 1]$ is the learning rate and $f_m$ the $m$-th tree fitted to the current negative gradient. Friedman (2001), Greedy function approximation: a gradient boosting machine, introduced $\nu$ ("shrinkage") and reported the persistent empirical law: smaller $\nu$ with correspondingly larger $M$ never hurts test error, at compute cost $M \propto 1/\nu$.

The theoretical picture: boosting is coordinate-descent-like optimisation in function space (Mason et al., 1999, Boosting algorithms as gradient descent), and shrinkage lengthens the regularisation path traversed slowly — early stopping then selects a point on that path. For linear weak learners, boosting with shrinkage tracks the $\ell_1$ regularisation path (Rosset, Zhu and Hastie, 2004; Efron et al., 2004 relate it to LARS). Early stopping itself is a well-characterised regulariser: for gradient-type methods the number of iterations plays the role of an inverse penalty strength (Yao, Rosasco and Caponnetto, 2007; Zhang and Yu, 2005 for boosting specifically).

Stochastic gradient boosting

Row subsampling per tree is Friedman (2002), Stochastic gradient boosting: fit each $f_m$ on a random fraction $\eta$ of rows. Benefits are variance reduction through decorrelation (the bagging mechanism imported into boosting) plus constant-factor speedups. Column subsampling per tree/node came from random forests via XGBoost practice. Typical productive ranges ($\eta \in [0.5, 1]$) interact weakly with $\nu$, which is why they can be tuned after it.

Which knobs actually matter

Tunability studies quantify the intuition. Probst, Boulesteix and Bischl (2019), Tunability: importance of hyperparameters of machine learning algorithms (JMLR), measure per-hyperparameter gain over defaults across many datasets: for gradient boosting the learning-rate/iterations pair and tree complexity dominate; sampling fractions and L2 terms contribute smaller, dataset-dependent gains. This ordering justifies the staged search: a full joint grid mostly re-measures the flat directions, the failure mode analysed in random search.

Recommended search distributions (log-uniform where noted):

  • $\nu$: log-uniform over $[10^{-2}, 0.3]$ — see Bayesian optimisation for why log scale.
  • leaves: log-uniform integers $[8, 256]$; or depth uniform $[3, 8]$.
  • min_child_samples: log-uniform integers $[5, 100]$.
  • subsample, colsample: uniform $[0.5, 1.0]$.
  • $\lambda$: log-uniform $[10^{-3}, 10]$ (with a point mass at 0).

Interactions worth knowing

  • $\nu$ × $M$: near-perfect trade along $\nu M \approx$ const until the small-$\nu$ plateau.
  • Tree size × min_child_samples: both cap leaf granularity; tune the second only after the first.
  • Subsampling × $\nu$: heavy subsampling adds gradient noise, which small $\nu$ smooths — aggressive values of both is a known-good corner (it is the default philosophy of LightGBM's GOSS).

Bentéjac et al. (2021) provide cross-library tuning benchmarks; the ranking of knob importance replicates across XGBoost, LightGBM and CatBoost.

What to learn next