Imbalanced, Multi-class and Multi-label
Lift and gain charts
Sort customers by model score, take the top slice, and ask — how many times better than random guessing did we do? That number is lift.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A lift chart answers one business question: if I can only act on my model's top picks, how much better is that than picking at random?
Think of a mango seller with one crate to fill and a thousand mangoes to choose from. An experienced picker fills the crate with mostly ripe fruit. A blindfolded picker gets the ordinary mix. The experienced picker's advantage — ripe fruit in the crate versus ripe fruit in a random crate — is the lift.
Why it exists
The usual scores grade all predictions at once. But many businesses cannot act on all predictions. A telecom can call 10% of customers about leaving. A charity can mail 20,000 donors. Capacity is fixed; the model's job is choosing who.
So the question changes. Not "how accurate is the model?" but "if I take the model's top 10%, how many of the people I care about are in it?"
Gain answers with a share. The top 10% contains, say, a third of all the responders. Lift answers with a multiplier: that is three-and-a-half times what random selection gets you.
How it works
Score everyone with the model. Sort from most to least confident. Cut the sorted list into ten equal slices, called deciles. Then count what you caught in each slice.
sorted customers → [ top 10% ][ next 10% ] ... [ bottom 10% ]
responders caught: 36% 21% ... 1%
lift: 3.6x 2.2x ... 0.1xA good model piles the responders into the early slices. A useless model spreads them evenly, and every slice has a lift near one.
A real example you have seen
Those "special recharge offer" SMS messages do not go to every subscriber. A churn model ranks everyone by risk of leaving, and the retention budget covers the top decile or two. The chart that convinced a manager to fund that campaign was almost certainly a gain chart — it converts model quality into "calls saved per rupee".
Remember this
- Gain: what share of all targets the top slices captured.
- Lift: how many times better that is than random picking.
- Use them when you can only act on the top of the ranking.
What to learn next
- One-vs-rest and one-vs-one — from rare-class problems to many-class problems.
- Model evaluation — precision, recall and AUC, the metrics under this chart.
- Recommender ranking metrics — the same "top of the list" thinking, applied to recommendations.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 and NumPy 1.26, CPU only.
Deciles by hand — no plotting library needed
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=5000, n_features=8, weights=[0.9, 0.1],
class_sep=0.6, random_state=42)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=42)
proba = LogisticRegression(max_iter=1000).fit(X_tr, y_tr).predict_proba(X_te)[:, 1]
order = np.argsort(-proba) # best scores first
deciles = np.array_split(y_te[order], 10) # ten equal slices of customers
base_rate = y_te.mean()
print("decile responders gain% lift")
found = 0
for i, d in enumerate(deciles, 1):
found += d.sum()
gain = 100 * found / y_te.sum()
lift = d.mean() / base_rate
print(f" {i:2d} {d.sum():4d} {gain:5.1f} {lift:4.1f}")decile responders gain% lift 1 46 35.7 3.6 2 28 57.4 2.2 3 12 66.7 0.9 4 14 77.5 1.1 5 12 86.8 0.9 6 3 89.1 0.2 7 8 95.3 0.6 8 3 97.7 0.2 9 1 98.4 0.1 10 2 100.0 0.2
The walkthrough
Read the first row as a sentence. "Call the top 10% of customers and you reach 35.7% of all likely responders — 3.6 times what random calling achieves." That sentence wins budget meetings; an AUC of 0.85 does not.
Gain is cumulative, lift here is per-decile. Gain climbs to 100% by construction. Per-decile lift shows where the model stops helping: after decile 2 it hovers near or below 1.0, meaning those slices are no better than random. Cumulative lift (gain% ÷ percent contacted) is the other common convention — decile 2's would be 57.4 ÷ 20 ≈ 2.9.
Only the ranking matters. predict_proba values could be badly calibrated and this chart would not change, because sorting ignores the absolute numbers. That makes lift robust for models trained on resampled data, whose probabilities are skewed — as noted in balanced ensembles.
np.array_split, not np.split. The test set has 1250 rows, which does not divide by 10 evenly. array_split tolerates unequal slices; split raises an error.
Common mistakes
Building the chart on training data. The ranking must come from data the model never saw, or every decile flatters. Same discipline as resampling inside cross-validation.
Ignoring ties. A model that outputs few distinct scores (a small tree, for instance) puts thousands of customers on the same score. How ties are ordered then decides decile membership; shuffle within ties or use more granular scores.
Comparing lifts across datasets. Lift is relative to the base rate. A lift of 3 on a 10% base rate and a lift of 3 on a 0.1% base rate are wildly different absolute results. Report the base rate next to the chart, always.
Stopping at decile 1. The full curve shows where returns flatten. Here, contacting past decile 2 buys little — that is the business recommendation.
Try it yourself
Compute cumulative lift and find the contact depth where it drops below 2.0. Then retrain with class_weight="balanced" and check whether the deciles change at all — reason about why from the "only the ranking matters" note above.
What to learn next
- One-vs-rest and one-vs-one — from rare-class problems to many-class problems.
- Model evaluation — precision, recall and AUC, the metrics under this chart.
- Recommender ranking metrics — the same "top of the list" thinking, applied to recommendations.
Researcher — Mathematics and papers.
Definitions
Let scores s(x) rank samples, π be the positive base rate, and t ∈ (0, 1] a contact depth (fraction of the population, taken from the top of the ranking). With TPR(t) denoting the fraction of all positives captured within depth t:
- Cumulative gain: G(t) = TPR(t)
- Cumulative lift: L(t) = G(t) / t
- Per-bucket lift: precision within the bucket ÷ π
The gain curve is the CAP (cumulative accuracy profile). Relations to standard curves: the ROC plots TPR against FPR, while the gain curve plots TPR against t = π·TPR + (1−π)·FPR — an axis reparameterisation. Consequently the accuracy ratio of the CAP equals 2·AUC − 1, i.e. the Gini coefficient. A model's gain curve is bounded above by the perfect model's curve (slope 1/π until saturation) and below by the diagonal G(t) = t.
Connection to precision–recall
L(t) = precision(t) / π at every depth, so a lift chart is a precision–recall curve wearing business clothes: recall on one axis, precision expressed as a multiple of the base rate. For rare classes, PR-space arguments (Davis and Goadrich, 2006, The relationship between precision-recall and ROC curves, ICML) carry over: lift comparisons between models can flip when the base rate changes, and interpolation between chart points is nonlinear.
Optimal targeting depth
With per-contact cost c and per-conversion value v, expected profit at depth t on population N is: Profit(t) = N·t·(π·L(t)·v − c). Maximising over t gives the economically optimal campaign size — the point where marginal precision π·L'(t)·... falls to c/v. This decision-theoretic reading (Elkan, 2001) is the rigorous version of "stop at decile 2".
Practical notes
Estimate uncertainty by bootstrap on the test ranking; decile gains on rare classes have wide intervals (the decile-1 count above is 46 — Poisson-level noise of ±7 changes lift by ±0.5). For model monitoring in production, per-decile lift over time is a sensitive drift detector: ranking degradation appears in the top deciles before global AUC moves. Weighted variants (expected value per contact rather than binary response) generalise the chart to uplift modelling — Radcliffe (2007), Using control groups to target on predicted lift, where the quantity ranked is the causal effect of contacting, not the response probability.
What to learn next
- One-vs-rest and one-vs-one — from rare-class problems to many-class problems.
- Model evaluation — precision, recall and AUC, the metrics under this chart.
- Recommender ranking metrics — the same "top of the list" thinking, applied to recommendations.