Ensembles and Gradient Boosting

CatBoost

CatBoost feeds text-like category columns straight into gradient boosting, and its ordered trick stops a model from secretly grading its own answers.

Read these first

On this page 5
  1. Why this had to be invented
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

CatBoost is a gradient boosting library that accepts category columns — city names, product types — directly, and encodes them without letting answers leak into features.

Picture a queue of people rating a new dish, one by one. Each person may hear the verdicts of people ahead of them in the queue, never behind. Nobody's own verdict can influence the information they received before tasting.

Hold that picture. It is the exact trick — the "cat" is for categorical — that makes CatBoost's handling of category columns safe where the naive approach quietly cheats.

Why this had to be invented

Models eat numbers. A column like city = mumbai / delhi / pune has to become numeric somehow. Numbering the cities 1, 2, 3 invents a fake ordering. One-hot columns explode when a column has thousands of values, as covered in feature engineering.

The tempting third option: replace each city with the average answer for that city. Replace mumbai with the fraction of Mumbai customers who cancelled. Informative, compact — and subtly poisoned. Each row's own answer is inside its own average. The feature contains a trace of the very thing the model must predict. Training scores look great; real-world scores do not. This trap is called target leakage.

CatBoost's fix is the queue. Shuffle the rows, then encode each row's city using only the answers of rows earlier in the queue. A row's own answer never touches its own feature.

How it works

shuffled rows:      r1    r2    r3    r4    r5
encoding r4's city  ↑     ↑     ↑
"mumbai" uses the answers of r1..r3 only
r4's own answer:  never used for r4's feature

(and the whole thing is repeated with several shuffles)

The rest is gradient boosting much like XGBoost: trees built one after another, each correcting the last. CatBoost also applies the same queue discipline inside boosting itself, which reduces another, subtler self-grading effect.

A real example you have seen

Category-heavy tables are everywhere money moves: which city, which phone model, which app version, which cuisine. Ride apps predicting cancellations and food apps predicting delivery delays run on exactly such columns. CatBoost was built at Yandex — which runs search, taxi and delivery businesses — precisely because their tables were mostly categories.

Remember this

  • CatBoost eats category columns directly — no manual encoding step.
  • Averaging answers per category leaks unless done in queue order; CatBoost does it in queue order.
  • Everything else is familiar gradient boosting, tuned to be strong out of the box.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install catboost

Honest warning: CatBoost is a heavy install. The wheel is roughly 100 MB, and it unpacks to about 345 MB on disk (measured for catboost 1.2.10, Windows, Python 3.10 — your platform will differ). It bundles compiled training code and a visualisation UI. Outputs verified with catboost 1.2.10 and pandas 2.2.3 on CPU.

Categories in, predictions out — no encoding step

A tiny made-up churn table. Note that two columns are raw strings.

catboost_demo.py
import pandas as pd
from catboost import CatBoostClassifier

df = pd.DataFrame({
    "city":   ["mumbai", "delhi", "mumbai", "pune", "delhi", "pune",
               "mumbai", "delhi", "pune", "mumbai", "delhi", "pune"],
    "plan":   ["basic", "pro", "pro", "basic", "basic", "pro",
               "basic", "pro", "basic", "pro", "basic", "pro"],
    "months": [3, 24, 18, 2, 5, 30, 1, 12, 4, 20, 2, 26],
    "left":   [1, 0, 0, 1, 1, 0, 1, 0, 1, 0, 1, 0],
})
X, y = df[["city", "plan", "months"]], df["left"]

model = CatBoostClassifier(iterations=200, verbose=0, random_seed=7)
model.fit(X, y, cat_features=["city", "plan"])

new_customer = pd.DataFrame({"city": ["pune"], "plan": ["pro"], "months": [22]})
print("training accuracy:", model.score(X, y))
print("will this customer leave?", model.predict(new_customer)[0])
print("probability of leaving:", round(model.predict_proba(new_customer)[0, 1], 3))
Output
training accuracy: 1.0
will this customer leave? 0
probability of leaving: 0.109

No LabelEncoder, no one-hot columns, no preprocessing pipeline for the categories. The strings went straight in.

One honest reading of that output: training accuracy 1.0 on 12 rows means the model memorised this toy table, as any boosted model would. The demonstration here is the API, not the score. Real evaluation needs held-out data and realistic row counts.

The walkthrough

cat_features=["city", "plan"] is the entire integration. Named columns are treated as categories and get the ordered encoding described above, plus automatic category combinations (city × plan) that CatBoost tries as candidate features.

verbose=0 matters: the default prints a line per iteration.

Strong defaults are a design goal. CatBoost aims to perform well untuned — symmetric trees, ordered boosting, and a tuned learning-rate schedule are all on by default. It is a reasonable first model on a category-heavy table before any tuning.

Prediction speed. CatBoost's trees are symmetric — every node at the same depth asks the same question — which turns prediction into a handful of table lookups. Among the big three libraries it typically serves fastest, a point that matters more in production than leaderboard decimals.

Common mistakes

Encoding categories yourself first. Running LabelEncoder or one-hot before CatBoost throws away its main advantage and, with label encoding, invents a fake numeric order. Hand it strings and declare cat_features.

Forgetting cat_features entirely. String columns will error, but integer-coded categories (like pincode) will be silently treated as numbers — the model learns that pincode 400001 is "less than" 500001. Declare integer categories explicitly.

Rolling your own target encoding with pandas. df.groupby("city")["left"].mean() computed on the full training set is the classic self-grading leak. If you must hand-encode for another library, compute encodings inside cross-validation folds — or let CatBoost do it properly.

Judging models on 12 rows. The toy table above shows API shape only. CatBoost's ordered machinery earns its keep from thousands of rows upward; on tiny data its extra machinery can even cost accuracy against a plain model.

Try it yourself

Add a fourth column phone = ["android", "iphone", ...] with values of your choosing, declare it in cat_features, and check the prediction changes. Then deliberately mis-declare months as categorical and observe that nothing errors — but consider what the model can no longer learn about month ordering once you did.

What to learn next

Researcher — Mathematics and papers.

Target statistics and prediction shift

The naive greedy target statistic for category value $k$ on row $i$ is

$$ \hat{x}i = \frac{\sum{j \neq i} \mathbb{1}[x_j = k] \, y_j + a p}{\sum_{j \neq i} \mathbb{1}[x_j = k] + a} $$

Where:

  • $y_j$ — the label of row $j$.
  • $a$ — a smoothing weight pulling rare categories toward the prior.
  • $p$ — the prior (the global mean label).

Even with the row excluded ($j \neq i$), computing this on the full training set makes the distribution of $\hat{x}_i$ differ between training and inference — a conditional shift that Prokhorenkova et al. (2018), CatBoost: unbiased boosting with categorical features (NeurIPS), name target leakage in the feature construction, and quantify with explicit toy-model bias terms.

Ordered target statistics fix this by drawing a random permutation $\sigma$ and restricting each row's sum to its predecessors, ${j : \sigma(j) < \sigma(i)}$ — the "queue" of the beginner section. Early rows get noisy encodings, so CatBoost averages over several permutations across iterations.

Ordered boosting and prediction shift in the ensemble

The same paper identifies a second, boosting-level bias: standard gradient boosting computes residuals for row $i$ using models that were themselves fitted on row $i$. The residuals are therefore systematically underestimated — prediction shift. Ordered boosting maintains, conceptually, a sequence of models where the residual of row $i$ comes from a model trained only on rows before $i$ in the permutation. The exact scheme is memory-hungry; the implementation approximates it with $O(\log n)$ model states per permutation. On small-to-medium datasets the correction shows measurable gains; on very large data the bias it removes shrinks naturally.

Symmetric (oblivious) trees

CatBoost's base learner is the oblivious decision tree: one splitting feature-threshold pair per depth level, shared by all nodes at that level. A depth-$D$ tree is a table of $2^D$ leaves indexed by a $D$-bit key. Consequences:

  • Inference is $D$ comparisons and a lookup — SIMD-friendly, and the reason for CatBoost's serving speed.
  • The constraint acts as regularisation; individual trees are weaker, which suits boosting.

Oblivious trees trace to Kohavi (1994) and were used in Yandex's earlier MatrixNet ranker.

Empirical position

Bentéjac et al. (2021) and the original paper's benchmarks place CatBoost at or near parity with LightGBM and XGBoost on numeric tables, and typically ahead when high-cardinality categoricals dominate — the regime its machinery targets. Training is generally slower than LightGBM; inference is generally faster. GPU training is first-class.

All three libraries have since converged on features pioneered across the trio: histogram splits everywhere, native categorical support in LightGBM, and categorical handling arriving in XGBoost. The durable CatBoost lessons are the two ordered constructions — statistically clean encodings and residuals.

What to learn next