Ensembles and Gradient Boosting
Tuning gradient-boosted trees
Almost all gradient boosting tuning reduces to one trade — a smaller learning rate with more trees, cut off by early stopping — plus a short list of knobs in priority order.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Tuning a gradient-boosted model is mostly one decision: cook on a lower flame for longer, and let a timer tell you when to stop.
Anyone who has made dal knows the trade. A high flame finishes fast and scorches easily. A low flame takes longer and forgives mistakes. The dish is better slow — up to a point, after which more simmering changes nothing.
In gradient boosting, the flame is the learning rate: how big a correction each new tree applies. The cooking time is the number of trees. And the timer is early stopping: watch a side dish of held-out data, and stop when it stops improving.
Why this had to be invented
Boosted trees have an intimidating settings page — twenty-plus knobs in every library. Beginners either freeze, or worse, search all of them at once and burn a week of compute learning nothing.
The escape is knowing that the knobs are not equals. One trade dominates. A short list matters after it. The rest is decoration on most problems.
How it works
Each tree adds a correction to the running answer. The learning rate scales that correction down.
learning rate 0.3 → ██ 14 corrections, each large
learning rate 0.1 → ██████ 59 corrections, medium
learning rate 0.03 → ████████████ 245 corrections, each tiny ← best scoreBig corrections overshoot — like turning a steering wheel in jerks. Small corrections follow the road closely but need more of them. Early stopping removes the guesswork about "how many": ask for far too many trees, and stop when the held-out score flattens.
After that trade, the knobs in order of importance:
1. flame + timer learning rate, with early stopping
2. tree size how much each tree can memorise
3. minimum leaf size how few examples may end up in one decision
4. data sampling show each tree a random portion of rows/columnsA real example you have seen
Every leaderboard-topping tabular model — fraud detection, credit scoring, competition winners — went through exactly this loop. Teams that win are rarely using secret algorithms. They are using this checklist, patiently, with honest validation.
Remember this
- Lower learning rate + more trees + early stopping is the core recipe.
- Tree size (leaves or depth) is the second knob: it controls memorisation per tree.
- Never hand-pick the number of trees — let early stopping cut it.
What to learn next
- Random search — the honest default for searching the second-tier knobs.
- Optuna — automating the whole loop, pruning bad trials early.
- K-fold cross-validation — the measurement layer every search stands on.
Developer — Code and libraries.
Setup
pip install lightgbmOutputs verified with lightgbm 4.7.0 and scikit-learn 1.7.2 on CPU; the script runs in a few seconds. The same logic applies verbatim to XGBoost and CatBoost — only parameter names change.
The flame experiment
Same data, same model, three learning rates. Early stopping picks each one's tree count.
import lightgbm as lgb
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=6000, n_features=20, n_informative=10,
flip_y=0.1, random_state=11)
Xtr, Xval, ytr, yval = train_test_split(X, y, test_size=0.3, random_state=11)
for lr in (0.3, 0.1, 0.03):
model = lgb.LGBMClassifier(n_estimators=5000, learning_rate=lr,
num_leaves=31, random_state=11, verbose=-1)
model.fit(Xtr, ytr, eval_X=Xval, eval_y=yval,
callbacks=[lgb.early_stopping(100, verbose=False)])
print(f"lr={lr:<5} kept {model.best_iteration_:4d} trees, "
f"val accuracy {model.score(Xval, yval):.3f}")lr=0.3 kept 14 trees, val accuracy 0.873 lr=0.1 kept 59 trees, val accuracy 0.881 lr=0.03 kept 245 trees, val accuracy 0.887
Each halving-ish of the learning rate roughly quadrupled the trees kept and nudged accuracy up. That pattern — slower is slightly better, at proportionally higher cost — is the norm, not a quirk of this dataset.
The walkthrough
n_estimators=5000 is not a real request. It is a ceiling that early stopping will never hit. Tuning n_estimators by grid search while the learning rate sits fixed is redundant work — the callback finds the right count for every learning rate automatically.
Why not learning rate 0.001? Follow the trend and the returns shrink while cost explodes: ten times the trees for a fraction of a point. Common practice: explore with 0.1, do final training at 0.03–0.05, and go lower only when a competition decimal pays for the compute.
The knobs after the flame, with names across libraries:
| knob | LightGBM | XGBoost | typical range |
|---|---|---|---|
| tree size | num_leaves | max_depth | 16–256 leaves / depth 3–8 |
| min rows per leaf | min_child_samples | min_child_weight | 5–100 |
| row sampling | subsample + subsample_freq | subsample | 0.6–1.0 |
| column sampling | colsample_bytree | colsample_bytree | 0.6–1.0 |
| L2 strength | reg_lambda | lambda | 0–10 |
Search those with random search or Optuna rather than by hand — while the learning rate stays fixed at your exploration value and early stopping stays on.
Common mistakes
Tuning against the test set. Every knob you adjust because a score went up spends that score's honesty. Keep three pools: training data, a validation set for early stopping and tuning, and a final test set touched once.
Early stopping on training loss. Training loss falls forever; the callback would never fire usefully. The eval data must be data the trees never trained on.
Turning every knob at once. Change five settings, watch the score move, learn nothing about why. Fix the flame, tune tree size; fix that, tune sampling. One variable at a time is slow but compounds; a proper search strategy automates it honestly.
Trusting one validation split for small data. With a few thousand rows, a lucky split lies. Use k-fold cross-validation for the final comparison of candidate settings.
Try it yourself
Add num_leaves values 8 and 128 as an inner loop to flame.py at lr=0.1. Before running, predict which combination overfits (hint: watch the gap between training and validation accuracy — add a model.score(Xtr, ytr) print). Then check how many trees early stopping keeps for big trees versus small ones.
What to learn next
- Random search — the honest default for searching the second-tier knobs.
- Optuna — automating the whole loop, pruning bad trials early.
- K-fold cross-validation — the measurement layer every search stands on.
Researcher — Mathematics and papers.
Shrinkage as regularisation
The boosted model after $M$ rounds is $F_M(x) = \sum_{m=1}^{M} \nu \, f_m(x)$ where $\nu \in (0, 1]$ is the learning rate and $f_m$ the $m$-th tree fitted to the current negative gradient. Friedman (2001), Greedy function approximation: a gradient boosting machine, introduced $\nu$ ("shrinkage") and reported the persistent empirical law: smaller $\nu$ with correspondingly larger $M$ never hurts test error, at compute cost $M \propto 1/\nu$.
The theoretical picture: boosting is coordinate-descent-like optimisation in function space (Mason et al., 1999, Boosting algorithms as gradient descent), and shrinkage lengthens the regularisation path traversed slowly — early stopping then selects a point on that path. For linear weak learners, boosting with shrinkage tracks the $\ell_1$ regularisation path (Rosset, Zhu and Hastie, 2004; Efron et al., 2004 relate it to LARS). Early stopping itself is a well-characterised regulariser: for gradient-type methods the number of iterations plays the role of an inverse penalty strength (Yao, Rosasco and Caponnetto, 2007; Zhang and Yu, 2005 for boosting specifically).
Stochastic gradient boosting
Row subsampling per tree is Friedman (2002), Stochastic gradient boosting: fit each $f_m$ on a random fraction $\eta$ of rows. Benefits are variance reduction through decorrelation (the bagging mechanism imported into boosting) plus constant-factor speedups. Column subsampling per tree/node came from random forests via XGBoost practice. Typical productive ranges ($\eta \in [0.5, 1]$) interact weakly with $\nu$, which is why they can be tuned after it.
Which knobs actually matter
Tunability studies quantify the intuition. Probst, Boulesteix and Bischl (2019), Tunability: importance of hyperparameters of machine learning algorithms (JMLR), measure per-hyperparameter gain over defaults across many datasets: for gradient boosting the learning-rate/iterations pair and tree complexity dominate; sampling fractions and L2 terms contribute smaller, dataset-dependent gains. This ordering justifies the staged search: a full joint grid mostly re-measures the flat directions, the failure mode analysed in random search.
Recommended search distributions (log-uniform where noted):
- $\nu$: log-uniform over $[10^{-2}, 0.3]$ — see Bayesian optimisation for why log scale.
- leaves: log-uniform integers $[8, 256]$; or depth uniform $[3, 8]$.
min_child_samples: log-uniform integers $[5, 100]$.- subsample, colsample: uniform $[0.5, 1.0]$.
- $\lambda$: log-uniform $[10^{-3}, 10]$ (with a point mass at 0).
Interactions worth knowing
- $\nu$ × $M$: near-perfect trade along $\nu M \approx$ const until the small-$\nu$ plateau.
- Tree size ×
min_child_samples: both cap leaf granularity; tune the second only after the first. - Subsampling × $\nu$: heavy subsampling adds gradient noise, which small $\nu$ smooths — aggressive values of both is a known-good corner (it is the default philosophy of LightGBM's GOSS).
Bentéjac et al. (2021) provide cross-library tuning benchmarks; the ranking of knob importance replicates across XGBoost, LightGBM and CatBoost.
What to learn next
- Random search — the honest default for searching the second-tier knobs.
- Optuna — automating the whole loop, pruning bad trials early.
- K-fold cross-validation — the measurement layer every search stands on.