Machine Learning

XGBoost

XGBoost grows trees one after another, each one trained on the mistakes the earlier trees left behind, which is why it usually wins on tabular data and why it will overfit if you let it.

Read these first

On this page 8
  1. The short answer
  2. The tailor's chalk
  3. Why it exists
  4. How it works
  5. Where you have already seen it
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

XGBoost builds decision trees one after another, and every new tree is trained to fix what the earlier trees got wrong.

The tailor's chalk

Think about getting a kurta stitched. The tailor takes your measurements and stitches a first version. You try it on. It pulls across the shoulder and hangs loose at the waist.

He does not throw it away and start again. He marks those two places with chalk and takes them in. You try it again, and there is one small thing left. He marks that.

Each fitting corrects what is left over from the last one. After three or four rounds the shirt fits.

That is boosting. The chalk marks have a name too. They are the residuals, meaning the part of the answer the model has not got right yet.

Why it exists

A random forest grows hundreds of trees at the same time, on different random samples, and averages them. The trees never see each other's work. Averaging cancels their random errors, which is a real and useful thing to do.

But averaging can only remove errors that differ between trees. If every tree is wrong about the same thing, the average is wrong about it too. A forest reduces the twitchiness. It does not chase down a mistake all the trees share.

Boosting attacks that directly. Tree two is not shown the original problem. It is shown what is still wrong after tree one. Tree three is shown what is still wrong after both.

Nothing is being averaged. The trees are being added up.

How it works

   the data
       |
   guess the average for everybody          error: large
       |
   tree 1 learns "where am I still wrong?"
       |  add a fraction of its answer
       v
   guess is better now                      error: smaller
       |
   tree 2 learns "where am I STILL wrong?"
       |  add a fraction of its answer
       v
   guess is better again                    error: smaller still
       |
      ...  a few hundred times ...
       v
   the final answer is the average, plus a bit of tree 1,
   plus a bit of tree 2, plus a bit of tree 3, and so on

Two words in that picture carry most of the weight.

"a fraction" — you never accept a whole correction. You take a slice of it, often five or ten percent. This is the learning rate, meaning how much of each correction you keep.

The tailor does the same thing. He takes a seam in a little at a time. He does not guess the final width in one cut. Small corrections, repeated, land far closer than one big one.

"still wrong" — each tree works on the leftovers, not on the original target. This is what makes the trees a team instead of a crowd.

Where you have already seen it

  • Credit scoring, where the inputs are a table of income, age and repayment history.
  • Fraud checks on a card swipe, decided in a few milliseconds.
  • Delivery time estimates in food and cab apps.
  • Search and advertisement ranking, deciding what to put at the top.
  • Machine learning competitions. For a decade, tabular contests were won by gradient boosting far more often than by anything else.

If your data lives in a spreadsheet with named columns, this family of models is the thing to try first. Neural networks are not the default there, and pretending otherwise wastes people's time.

What is honestly hard here

More rounds will eventually make it worse. This is the biggest single difference from a random forest, and it catches everyone.

In a forest, more trees never hurt. You add trees until you run out of memory or patience. In boosting, each new tree is trained on the leftovers. Once the real pattern is used up, the only leftovers are noise. The model starts learning the noise. That is overfitting, and you will watch it happen in the Developer section.

It needs tuning, and a forest mostly does not. A random forest with default settings is usually close to its best. XGBoost with default settings is usually not. That extra work is the price of the extra accuracy.

It still cannot see past its data. Boosting is made of trees, so it inherits the tree's flat horizon. Ask about a flat twice the size of anything you trained on. You get back the same answer as for the largest one you saw.

It is not a neural network and it does not want to be. For pictures, sound and language, use a neural network. For tables, use this.

Remember this

  • Boosting adds trees in sequence, each one trained on what is still wrong.
  • The learning rate decides how much of each correction you accept, and small is usually better.
  • More trees help a forest forever, and hurt a boosted model eventually. Watch a held-out score and stop when it stops improving.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install xgboost scikit-learn numpy

Everything here is generated in memory and runs in a few seconds on a CPU. XGBoost has no GPU requirement.

Boosting with your own hands, in twelve lines

Before touching the library, build the loop. Eight flats, one clue, and a stack of depth-1 trees. A depth-1 tree is called a stump: it asks one question and gives two answers.

by_hand.py
import numpy as np
from sklearn.tree import DecisionTreeRegressor

# Eight flats. Size in hundreds of square feet -> rent in thousands.
size = np.array([[4.], [5.], [6.], [7.], [8.], [9.], [10.], [11.]])
rent = np.array([9., 11., 14., 16., 20., 23., 27., 30.])

pred = np.full_like(rent, rent.mean())      # round 0: everybody gets the average
lr = 0.5                                    # accept half of each correction
print(f"round 0  guess = the average {rent.mean():.2f}   error {np.abs(rent - pred).mean():.3f}")

for r in range(1, 6):
    residual = rent - pred                  # the chalk marks: what is still wrong
    stump = DecisionTreeRegressor(max_depth=1).fit(size, residual)
    pred = pred + lr * stump.predict(size)  # take the seam in, part of the way
    print(f"round {r}  error {np.abs(rent - pred).mean():.3f}   predictions {np.round(pred, 1)}")
Output
round 0  guess = the average 18.75   error 6.250
round 1  error 3.688   predictions [15.6 15.6 15.6 15.6 21.9 21.9 21.9 21.9]
round 2  error 2.583   predictions [14.5 14.5 14.5 14.5 20.8 20.8 25.2 25.2]
round 3  error 1.830   predictions [12.3 12.3 15.3 15.3 21.5 21.5 25.9 25.9]
round 4  error 1.180   predictions [11.6 11.6 14.6 14.6 20.9 22.6 27.  27. ]
round 5  error 0.959   predictions [11.4 11.4 14.4 14.4 20.7 22.4 26.8 28.5]

Watch the predictions column build a staircase.

After round 1 there are two distinct values: one stump, one question, two answers. After round 2 there are three, after round 3 there are four, and after round 5 there are six. The error has fallen from 6.250 to 0.959.

Each stump on its own is close to useless. A stump cannot describe rising rent — it can only say "small" or "large". Stacked with corrections, the same useless stumps trace the trend.

That is the whole algorithm. XGBoost is this loop, written carefully, with regularisation and a great deal of engineering.

Boosting against a forest, on the same data

Now the real thing. Two thousand rows, ten columns, of which four carry signal — including one effect that only appears when two columns are combined.

head_to_head.py
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split
from xgboost import XGBRegressor

rng = np.random.default_rng(7)
N = 2000
X = rng.uniform(-3, 3, size=(N, 10))
y = (np.sin(X[:, 0] * 1.5) * 3 + 0.6 * X[:, 1] ** 2 + 1.2 * X[:, 2] * X[:, 3]
     + rng.normal(0, 0.8, N))

# Three piles, not two: boosting needs a validation set to decide when to stop.
X_train, X_rest, y_train, y_rest = train_test_split(X, y, test_size=0.4, random_state=0)
X_val, X_test, y_val, y_test = train_test_split(X_rest, y_rest, test_size=0.5, random_state=0)

rf = RandomForestRegressor(n_estimators=400, random_state=0, n_jobs=-1).fit(X_train, y_train)
print(f"random forest, 400 trees   test error {mean_absolute_error(y_test, rf.predict(X_test)):.3f}")

gb = XGBRegressor(n_estimators=3000, learning_rate=0.05, max_depth=4,
                  subsample=0.8, colsample_bytree=0.8,
                  early_stopping_rounds=50, random_state=0, verbosity=0)
gb.fit(X_train, y_train, eval_set=[(X_val, y_val)], verbose=False)
print(f"xgboost, early stopping    test error {mean_absolute_error(y_test, gb.predict(X_test)):.3f}")
print(f"   rounds actually used: {gb.best_iteration + 1} out of 3000")
Output
random forest, 400 trees   test error 1.976
xgboost, early stopping    test error 1.155
   rounds actually used: 1460 out of 3000

Your last two numbers will differ a little. XGBoost sums gradient histograms across however many CPU threads your machine has, and floating-point addition is not associative, so the stopping round lands somewhere around 1450 to 1550 and the error moves in the third decimal. The forest line is deterministic and will match exactly.

That is roughly a 40 percent reduction in error, on identical data, in about two seconds.

early_stopping_rounds=50 is doing the important work. It scores the model on X_val after every round and stops when 50 rounds pass with no improvement. You ask for 3000 trees and the model stops at the few hundred to a couple of thousand it actually needs.

Use a separate validation set for this, not the test set. The stopping point is a decision made from data, and any data used to make a decision stops being an honest measure. This is the whole argument of train, test and validation splits.

The thing a forest does not do

Run this next, and read it against the forest column.

more_rounds_hurt.py
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split
from xgboost import XGBRegressor

rng = np.random.default_rng(0)
N = 250                                              # a small, noisy dataset
X = rng.uniform(-3, 3, size=(N, 6))                  # only 2 of 6 columns matter
y = np.sin(X[:, 0] * 1.5) * 3 + 0.6 * X[:, 1] ** 2 + rng.normal(0, 1.5, N)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.4, random_state=0)

print("rounds   forest test   boosting train   boosting test")
for n in (5, 20, 60, 150, 400, 1000):
    rf = RandomForestRegressor(n_estimators=n, random_state=0).fit(X_train, y_train)
    gb = XGBRegressor(n_estimators=n, learning_rate=0.05, max_depth=4,
                      random_state=0, verbosity=0).fit(X_train, y_train)
    print(f"{n:6d}   {mean_absolute_error(y_test, rf.predict(X_test)):11.3f}   "
          f"{mean_absolute_error(y_train, gb.predict(X_train)):14.3f}   "
          f"{mean_absolute_error(y_test, gb.predict(X_test)):13.3f}")
Output
rounds   forest test   boosting train   boosting test
     5         1.752            1.966           2.179
    20         1.636            1.288           1.820
    60         1.636            0.659           1.674
   150         1.594            0.335           1.670
   400         1.610            0.088           1.730
  1000         1.604            0.003           1.743

This is the most important output on the page. Three columns, three different behaviours.

The forest column barely moves. From 5 trees to 1000 trees it goes 1.752 down to about 1.60 and then stays there. More trees are a computation budget, not a capacity setting.

The boosting train column falls to nothing. 1.966 down to 0.003. By round 1000 the model reproduces its training rows almost exactly.

The boosting test column is a U. It improves to 1.670 at 150 rounds, then gets worse: 1.730, then 1.743. The best model on this data is not the biggest one.

Rounds 150 to 1000 were spent memorising 250 noisy rows. The training error kept falling the whole time, which is why the training score can never tell you when to stop.

n_estimators in XGBoost is a capacity knob. In a random forest it is not. Confusing the two is the most common mistake people bring from forests to boosting.

Gaps in your data, handled properly

Real tables have blanks. XGBoost accepts NaN directly and learns, at each split, which side a missing value should go.

Here is a loan dataset where the missing values are not random — people who defaulted are far more likely to have left the income field blank.

gaps.py
import numpy as np
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from xgboost import XGBClassifier

rng = np.random.default_rng(3)
N = 1200
income = rng.uniform(10, 100, N)
years = rng.uniform(0, 20, N)
default = (rng.random(N) < 1 / (1 + np.exp((income - 45) / 12))).astype(int)

X = np.column_stack([income, years])
blank = rng.random(N) < np.where(default == 1, 0.45, 0.05)   # blanks are not random
X[blank, 0] = np.nan
print(f"{blank.mean():.0%} of income values are missing, and defaulters "
      f"leave the field blank far more often")
print()

X_train, X_test, y_train, y_test = train_test_split(X, default, test_size=0.3, random_state=0)


def score(name, xtr, xte, ytr=y_train):
    model = XGBClassifier(n_estimators=200, max_depth=3, learning_rate=0.1,
                          random_state=0, verbosity=0).fit(xtr, ytr)
    print(f"{name:38s} accuracy {accuracy_score(y_test, model.predict(xte)):.3f}")


score("gaps left as NaN, XGBoost decides", X_train, X_test)

fill = np.nanmean(X_train[:, 0])
a, b = X_train.copy(), X_test.copy()
a[np.isnan(a[:, 0]), 0] = fill
b[np.isnan(b[:, 0]), 0] = fill
score("filled with the mean", a, b)

observed = X_train[~np.isnan(X_train[:, 0]), 0]
r = np.random.default_rng(0)
c, d = X_train.copy(), X_test.copy()
c[np.isnan(c[:, 0]), 0] = r.choice(observed, np.isnan(c[:, 0]).sum())
d[np.isnan(d[:, 0]), 0] = r.choice(observed, np.isnan(d[:, 0]).sum())
score("filled with a plausible random income", c, d)

keep = ~np.isnan(X_train[:, 0])
score("rows with gaps deleted", X_train[keep], b, y_train[keep])
Output
20% of income values are missing, and defaulters leave the field blank far more often

gaps left as NaN, XGBoost decides      accuracy 0.839
filled with the mean                   accuracy 0.839
filled with a plausible random income  accuracy 0.761
rows with gaps deleted                 accuracy 0.694

Four ways to handle the same gaps, and a fourteen-point spread between the best and the worst.

Deleting the rows was worst, at 0.694. Those rows were the most informative in the dataset, because the blank itself predicts default. Dropping them throws away a signal and biases what is left.

Filling with a plausible random income scored 0.761. This looks like the careful, statistically respectable choice. It is the one that erases the evidence. After the fill, a blank row is indistinguishable from a genuine mid-income row.

Mean filling tied with NaN at 0.839. This deserves an honest explanation rather than a victory lap. Every blank got the same value, so "income equals exactly 54.9" became a sentinel the tree could split on. It worked by accident. Fill with a per-row value instead, or with a number that collides with real incomes, and it stops working.

Leaving the gap as NaN is the option that does not depend on luck. XGBoost learns a default direction per split from the data itself.

If you must fill, add a was_missing column alongside. That keeps the evidence explicitly rather than by accident.

Common mistakes

Treating n_estimators like a forest's. Demonstrated above. Set it high and let early_stopping_rounds choose, or tune it against a validation set.

Early stopping against the test set. It makes the test score optimistic and you will not notice. Keep three piles.

A large learning rate with few rounds. learning_rate=0.3 with 50 rounds trains fast and generalises worse than 0.05 with 500. Lower the rate, raise the rounds, use early stopping. Time spent here beats time spent on anything else.

Passing a pandas DataFrame with string columns and no dtype. XGBoost 2.0 and later handle categoricals natively, but only if the columns are category dtype and you pass enable_categorical=True. Otherwise you get ValueError: DataFrame.dtypes for data must be int, float, bool or category.

One-hot encoding a column with thousands of levels. It produces thousands of sparse columns, each with too few positives to split on well. Use native categorical support, or target encoding done inside the folds — see feature engineering for how that leaks if you do it carelessly.

Reading feature_importances_ as truth. The same in-sample bias described in random forest applies here. Use permutation importance on held-out data, or shap.TreeExplainer if you need per-prediction attribution.

Reaching for XGBoost on 200 rows. With very little data, a regularised linear model or a forest is often better and always easier to defend. Boosting earns its keep from a few thousand rows upwards.

Try it yourself

In more_rounds_hurt.py, change learning_rate from 0.05 to 0.3 and rerun.

Predict what happens first. A larger learning rate means bigger corrections, so the model reaches the bottom of the U earlier and passes it sooner. Find the number of rounds where the test error is lowest at each learning rate, and note how the two settings trade against each other.

Then set learning_rate=0.01 and go up to 5000 rounds. Slower is usually better, and the exercise shows you what "usually" is worth in real numbers.

What to learn next

Researcher — Mathematics and papers.

The regularised objective

Chen & Guestrin (2016) frame boosting as additive training on a regularised objective. The model after K rounds is a sum of trees:

ŷ_i = SUM_{k=1..K} f_k(x_i) ,       f_k in F, the space of regression trees

L = SUM_i l(y_i, ŷ_i)  +  SUM_k Ω(f_k)

Ω(f) = γ T  +  (1/2) λ ||w||^2
  • l — any twice-differentiable loss (squared error, logistic, ranking objectives)
  • T — the number of leaves in tree f
  • w ∈ R^T — the leaf output values, called leaf weights
  • γ — the cost charged per additional leaf; gamma in the API
  • λ — L2 penalty on leaf weights; reg_lambda in the API

The explicit Ω term is the substantive departure from Friedman (2001). Classical gradient boosting controls complexity through tree depth and shrinkage; XGBoost puts it in the objective, where it participates in the split decision itself.

Second-order expansion

At round t, holding the previous t−1 trees fixed, take a second-order Taylor expansion of the loss around ŷ^(t−1):

L^(t)  ≈  SUM_i [ g_i f_t(x_i)  +  (1/2) h_i f_t(x_i)^2 ]  +  Ω(f_t)  +  const

g_i = ∂ l(y_i, ŷ) / ∂ŷ        evaluated at ŷ = ŷ_i^(t−1)      (gradient)
h_i = ∂² l(y_i, ŷ) / ∂ŷ²      evaluated at ŷ = ŷ_i^(t−1)      (hessian)

For squared error, g_i = 2(ŷ_i − y_i) and h_i = 2, which recovers "fit the next tree to the residuals" — the loop in the Developer section. For logistic loss, g_i = p_i − y_i and h_i = p_i(1 − p_i), so the hessian downweights confidently-classified points automatically.

Using h_i is what "second order" buys. Friedman's GBM uses only g_i and applies a separate line search for leaf values; XGBoost gets the step size analytically per leaf.

Optimal leaf weights and the split gain

Fix the tree structure and let I_j be the set of training rows landing in leaf j. Write G_j = SUM_{i ∈ I_j} g_i and H_j = SUM_{i ∈ I_j} h_i. The objective becomes a sum of independent quadratics in w_j, so the minimiser is closed form:

w_j*  =  − G_j / ( H_j + λ )

L*    =  − (1/2) SUM_{j=1..T}  G_j^2 / ( H_j + λ )   +  γ T

L* scores a tree structure. The gain from splitting a leaf into left and right children follows directly:

Gain = (1/2) [  G_L^2/(H_L + λ)  +  G_R^2/(H_R + λ)  −  (G_L+G_R)^2/(H_L+H_R+λ)  ]  −  γ

Three consequences are worth naming.

  • γ is a genuine threshold. A split whose gain does not exceed γ is not taken. This is pre-pruning built into the criterion, unlike CART's impurity decrease, which is always non-negative.
  • λ shrinks leaf weights towards zero and, because it appears in the denominator, penalises leaves with small H_j more. Leaves supported by few or low-confidence rows get pulled harder.
  • The gain can be negative. XGBoost grows to max_depth and then prunes bottom-up, removing splits with negative gain. A depth-first greedy stop would miss a strong split hiding behind a weak one.

Shrinkage and subsampling

ŷ^(t) = ŷ^(t−1)  +  η · f_t(x)

η is the learning rate (eta, learning_rate). Friedman (2001) established that small η with many rounds beats large η with few, and the standard operating range is 0.01 to 0.1. The mechanism is regularisation by slow fitting: each tree contributes a small fraction, so the ensemble approaches the fit along a smoother path.

Friedman (2002) added row subsampling (subsample), sampling without replacement rather than the bootstrap used by bagging. XGBoost adds column subsampling per tree, per level and per split (colsample_bytree, colsample_bylevel, colsample_bynode), imported from random forests. Column subsampling is frequently the more effective of the two on wide tabular data.

Sparsity-aware split finding

The Developer section shows NaN handled without imputation. The mechanism: at each split, XGBoost enumerates candidate thresholds twice — once sending all missing values left, once sending them right — and stores the better direction as that node's default direction.

Missing rows are excluded from the threshold enumeration entirely, so the cost is linear in the number of non-missing entries. On genuinely sparse data this yields the large speedups reported in the paper. The learned default direction also means "missingness" carries information through the model rather than being erased by an imputation step.

Approximate split finding

Exact greedy split finding enumerates every distinct feature value, costing O(n d) per level after sorting. XGBoost's approximate algorithm proposes candidate split points from quantiles of the feature distribution, then aggregates gradient statistics into those buckets.

Because the second-order objective weights each row by h_i, the quantiles must be weighted by h_i too. The paper's weighted quantile sketch supports merge and prune operations with a provable error bound, which the standard unweighted sketches do not. Setting sketch_eps = ε yields roughly 1/ε candidate splits.

Modern practice defaults to tree_method="hist": bin each feature into a fixed number of bins (max_bin, default 256) once, then accumulate gradient histograms per node. Cost per level drops to O(n_bins · d) and the histogram of a child is obtained by subtracting a sibling's from the parent's.

exact greedy    O(n d)         per level, after a global pre-sort
histogram       O(n d)         to bin, then O(n_bins · d) per node
memory          O(n d)         compressed column blocks, cache-aware prefetch

The three implementations

XGBoostLightGBMCatBoost
Growthdepth-wise, then pruneleaf-wise, best-firstoblivious (symmetric) trees
Split findinghistogram or exacthistogram + GOSShistogram
Categoricalsnative since 2.0nativeordered target statistics
Notable extrasparsity-aware default directionEFB feature bundlingordered boosting

LightGBM (Ke et al., 2017) grows leaf-wise: it expands whichever leaf offers the largest gain, anywhere in the tree. That gives lower loss per leaf and a stronger tendency to overfit small datasets, so num_leaves must be controlled. GOSS keeps all large-gradient rows and subsamples small-gradient ones with a compensating weight. EFB bundles mutually exclusive sparse features into single columns.

CatBoost (Prokhorenkova et al., 2018) targets target leakage in categorical encoding, the exact failure demonstrated in feature engineering. Its ordered target statistics compute a category's encoding using only rows preceding it in a random permutation. Ordered boosting extends the same permutation trick to the gradient estimates themselves, removing the prediction shift that arises from computing gradients on the same rows used to fit the tree. Its oblivious trees — the same split condition at every node of a level — act as strong regularisation and make inference exceptionally fast.

Boosting theory

Boosting predates gradient boosting. Schapire (1990) answered Kearns and Valiant's question of whether weak learnability implies strong learnability, in the affirmative. Freund & Schapire (1997) gave AdaBoost, with a training error bound decreasing exponentially in the number of rounds.

AdaBoost's resistance to overfitting was explained by Schapire et al. (1998) through the margin distribution: rounds after training error reaches zero continue to increase margins, and the generalisation bound depends on the margin distribution rather than on the number of rounds.

Friedman, Hastie & Tibshirani (2000) reinterpreted AdaBoost as forward stagewise additive modelling under exponential loss, which opened the door to arbitrary losses. Friedman (2001) then generalised it to gradient descent in function space — the framework everything above sits inside.

The empirical picture on tabular data remains stable. Grinsztajn, Oyallon & Varoquaux (2022) find tree ensembles still outperform deep learning on medium-sized tabular datasets, attributing it to rotation-invariance being the wrong inductive bias for tabular features, and to robustness against uninformative columns. Shwartz-Ziv & Armon (2022) reach the same conclusion and add that proposed deep tabular architectures often fail to reproduce outside their own papers' datasets.

Papers

  • Schapire, R. (1990). The Strength of Weak Learnability. Machine Learning 5(2).
  • Freund, Y. & Schapire, R. (1997). A Decision-Theoretic Generalization of On-Line Learning. JCSS 55(1).
  • Friedman, J., Hastie, T. & Tibshirani, R. (2000). Additive Logistic Regression: A Statistical View of Boosting. Annals of Statistics 28(2).
  • Friedman, J. (2001). Greedy Function Approximation: A Gradient Boosting Machine. Annals of Statistics 29(5).
  • Friedman, J. (2002). Stochastic Gradient Boosting. Computational Statistics and Data Analysis 38(4).
  • Chen, T. & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. KDD — arxiv.org/abs/1603.02754
  • Ke, G. et al. (2017). LightGBM: A Highly Efficient Gradient Boosting Decision Tree. NeurIPS.
  • Prokhorenkova, L. et al. (2018). CatBoost: Unbiased Boosting with Categorical Features. NeurIPS — arxiv.org/abs/1706.09516
  • Lundberg, S. et al. (2020). From Local Explanations to Global Understanding with Explainable AI for Trees. Nature Machine Intelligence 2.
  • Grinsztajn, L., Oyallon, E. & Varoquaux, G. (2022). Why Do Tree-Based Models Still Outperform Deep Learning on Tabular Data? NeurIPS — arxiv.org/abs/2207.08815

What to learn next