Data Engineering for AI

Imbalanced data

When one outcome is rare, accuracy stops meaning anything, and rebalancing only moves the threshold while a new column adds real information.

Read these first

On this page 8
  1. Why it exists as a problem
  2. The number that lies
  3. The two scores that tell the truth
  4. The fix people reach for, and what it really does
  5. Somewhere you have seen this
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Imbalanced data means one answer is far rarer than the other, which makes the usual accuracy score useless.

Think about a security guard at a mall gate. Almost nobody carries anything dangerous. A guard who waves every single person through will be right more than 999 times out of a thousand.

He has a wonderful score and does no work at all. He would miss the one person who mattered.

That is what happens to a model when one outcome is rare.

Why it exists as a problem

Most interesting predictions are about rare things. Fraud. A machine about to fail. A rare disease. A customer about to leave.

Rare is the whole point. And a model trained to maximise how often it is right learns the same trick as the lazy guard: always say no.

The number that lies

   10,000 payments
       120 are fraud          (about 1 in 83)
     9,880 are fine

   model that always says "not fraud"
       accuracy: 98.8%       <- looks excellent
       frauds caught: 0      <- catches nothing at all

Nobody would ship that model on purpose. Plenty of teams have shipped it by accident, because the dashboard showed 98.8% and everyone went home happy.

The two scores that tell the truth

Recall answers: of all the real frauds, how many did we catch? The lazy guard scores zero.

Precision answers: of everything we flagged, how much was really fraud? Flag every payment and precision collapses.

You cannot maximise both. Catching more fraud always means troubling more honest customers. That trade is a business decision, not a technical one.

The fix people reach for, and what it really does

The usual advice is to rebalance: duplicate the rare rows, or tell the model to count them more heavily.

That does help the model stop saying "no" to everything. But it does not teach the model anything new about fraud. It changes how cautious the model is, not how much it knows.

The thing that genuinely helps is a new column. One extra fact about each payment — was this device ever seen before — carries information no amount of rebalancing can invent.

Somewhere you have seen this

The one-time password your bank sends for an unusual payment. Behind it is a model that decided this looked different enough to check.

Notice it does not block the payment. It asks a cheap question instead. Good rare-event systems are designed around what to do when unsure, not around being certain.

What is honestly hard here

With 120 examples of fraud, you have 120 examples. That is a small number to learn a complicated pattern from, and no clever trick changes it.

You will also be told that some resampling method "fixes" imbalance. Be sceptical. It usually moves where the model draws its line, which you could have done directly by changing the threshold.

Remember this

  • With rare outcomes, accuracy is meaningless. Use precision and recall.
  • Rebalancing mostly moves the cut-off. It does not add knowledge.
  • A new, genuinely informative column is worth more than every resampling trick combined.

What to learn next

Developer — Code and libraries.

This experiment separates two things that get confused constantly: moving the threshold and adding information. One is free and changes nothing you did not already have. The other is the actual fix.

The data is card payments with a fraud rate near 1.2%. There is a new_device column being logged that nobody wired into the model.

Setup

bash
pip install numpy pandas scikit-learn

Rebalance, or log one more column

imbalance.py
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (accuracy_score, average_precision_score,
                             precision_score, recall_score, roc_auc_score)
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

rng = np.random.default_rng(31)
n = 20000
amount = rng.gamma(2.0, 900.0, n)
night = ((rng.integers(0, 24, n) < 5)).astype(int)
new_device = rng.binomial(1, 0.08, n)                    # logged, but nobody sent it to the model
risk = 0.0006 * amount + 1.4 * night + 3.0 * new_device - 7.2
fraud = rng.binomial(1, 1 / (1 + np.exp(-risk)))
df = pd.DataFrame({"amount": amount, "night": night, "new_device": new_device, "fraud": fraud})
train, test = train_test_split(df, test_size=0.3, random_state=0, stratify=df["fraud"])

print("rows:", len(df), " fraud rows:", int(df["fraud"].sum()),
      " fraud rate:", round(df["fraud"].mean(), 4))
print()


def report(name, pred, prob=None):
    line = (f"{name:34s}{accuracy_score(test['fraud'], pred):>9.4f}"
            f"{precision_score(test['fraud'], pred, zero_division=0):>10.3f}"
            f"{recall_score(test['fraud'], pred, zero_division=0):>8.3f}")
    if prob is not None:
        line += f"{roc_auc_score(test['fraud'], prob):>9.3f}{average_precision_score(test['fraud'], prob):>9.3f}"
    print(line)


def fit(cols, **kw):
    model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000, **kw))
    model.fit(train[cols], train["fraud"])
    return model.predict_proba(test[cols])[:, 1]


print(f"{'':34s}{'accuracy':>9}{'precision':>10}{'recall':>8}{'ROC-AUC':>9}{'PR-AUC':>9}")
report("always says 'not fraud'", np.zeros(len(test), dtype=int))
p_plain = fit(["amount", "night"])
report("2 columns, plain", (p_plain >= 0.5).astype(int), p_plain)
p_bal = fit(["amount", "night"], class_weight="balanced")
report("2 columns, class_weight balanced", (p_bal >= 0.5).astype(int), p_bal)
p_more = fit(["amount", "night", "new_device"])
report("3 columns (device flag added)", (p_more >= 0.5).astype(int), p_more)
p_more_bal = fit(["amount", "night", "new_device"], class_weight="balanced")
report("3 columns, class_weight balanced", (p_more_bal >= 0.5).astype(int), p_more_bal)

print()
print("the plain 2-column model, with only the threshold moved:")
print(f"  {'cut-off':>9}{'precision':>11}{'recall':>8}{'alerts raised':>15}")
for cut in [0.5, 0.2, 0.1, 0.05, 0.02]:
    pred = (p_plain >= cut).astype(int)
    print(f"  {cut:>9}{precision_score(test['fraud'], pred, zero_division=0):>11.3f}"
          f"{recall_score(test['fraud'], pred, zero_division=0):>8.3f}{int(pred.sum()):>15}")
Output
rows: 20000  fraud rows: 244  fraud rate: 0.0122

                                   accuracy precision  recall  ROC-AUC   PR-AUC
always says 'not fraud'              0.9878     0.000   0.000
2 columns, plain                     0.9880     1.000   0.014    0.798    0.120
2 columns, class_weight balanced     0.7140     0.031   0.753    0.799    0.119
3 columns (device flag added)        0.9878     0.500   0.068    0.909    0.285
3 columns, class_weight balanced     0.8432     0.060   0.808    0.908    0.277

the plain 2-column model, with only the threshold moved:
    cut-off  precision  recall  alerts raised
        0.5      1.000   0.014              1
        0.2      0.444   0.055              9
        0.1      0.196   0.123             46
       0.05      0.091   0.260            209
       0.02      0.050   0.616            902

Four things in that table

The trained model beat the lazy baseline by 0.0002 accuracy. 0.9880 against 0.9878. If accuracy is your dashboard metric, weeks of work are invisible on it.

Rebalancing changed the ranking by nothing. PR-AUC went 0.120 to 0.119; ROC-AUC 0.798 to 0.799. Recall rose from 0.014 to 0.753 and accuracy collapsed to 0.714. Read that carefully: class_weight="balanced" moved where the line sits, and left what the model knows untouched.

Look at the threshold table for proof. The plain model, unweighted, reaches recall 0.616 at a cut-off of 0.02. You do not need reweighting to trade precision for recall. You need one number.

Logging one more column more than doubled PR-AUC. 0.120 to 0.285, and ROC-AUC 0.798 to 0.909. That is new information entering the system, and it is the only change here that did.

Why PR-AUC, not ROC-AUC

PR-AUC (average precision) summarises the precision-recall curve. ROC-AUC summarises the true-positive against false-positive curve.

On rare-event problems ROC-AUC is over-optimistic. Its x-axis is the false positive rate, and with several thousand negatives in the test set, 200 false alarms barely move it. Precision divides by the number of alerts raised, so those same 200 false alarms are impossible to hide.

Report both. Trust PR-AUC for the decision. Davis and Goadrich (2006) showed a curve dominating in ROC space dominates in PR space too, but small ROC differences correspond to large, operationally decisive PR differences.

Reading the threshold table like an operator

That table is the actual product decision. At a cut-off of 0.05 you raise 209 alerts and catch 26% of fraud, with precision 0.091 — about ten alerts in eleven are false.

If a review costs 40 rupees and a missed fraud costs 8,000, you can compute the cut-off that minimises total cost directly. That calculation, not a model comparison, is what should set the threshold.

Note the top row: cut-off 0.5 gives precision 1.000 from a single alert. Perfect precision on one prediction is not a result.

Common mistakes

Resampling before splitting. Oversample first and copies of the same fraud row land in both train and test. Your recall will look extraordinary and be fictional. Split first, resample the training half only.

Using accuracy anywhere on a rare-event problem. It has one use: comparing against the always-negative baseline to see how little it says.

Believing SMOTE is free. SMOTE interpolates between a minority point and its neighbours. On high-dimensional or categorical data those interpolations land in regions where no real fraud lives, and they blur the boundary. Benchmark it against threshold tuning before adopting it; the threshold often wins.

Tuning the threshold on the test set. Pick it on a validation split, then measure once on test. See train, test and validation splits.

Optimising a single operating point. Report the whole curve. A model that is better at recall 0.3 and worse at recall 0.8 is not "better".

Try it yourself

Set new_device's coefficient in risk from 3.0 to 0.2, so the column carries little information. Re-run.

The three-column model should collapse back towards the two-column one. That is the control experiment: it proves the earlier gain came from the information in the column, not from having one more column.

Then compute the expected cost per cut-off with your own numbers for review cost and fraud loss, and find the minimum.

What to learn next

Researcher — Mathematics and papers.

Imbalance is a threshold problem before it is a learning problem

For a probabilistic classifier, the Bayes-optimal decision under asymmetric costs is a threshold on the posterior. With cost $c_{01}$ for a false negative and $c_{10}$ for a false positive, predict positive when:

$$ p(y=1 \mid x) \;>\; \tau^{*} \;=\; \frac{c_{10}}{c_{10} + c_{01}} $$

The class prior enters through $p(y=1 \mid x)$, not through $\tau^{*}$. Imbalance therefore does not require a change of algorithm. It requires that you stop using $\tau = 0.5$, which silently asserts equal costs.

Elkan (2001), The Foundations of Cost-Sensitive Learning, IJCAI, gives the correction for training on a resampled prior. If the base rate is changed from $\pi$ to $\pi'$, calibrated probabilities are recovered by:

$$ p(y=1 \mid x) \;=\; \frac{\pi \, p'(y=1\mid x)}{\pi\,p'(y=1\mid x) + (1-\pi)\big(1 - p'(y=1\mid x)\big)} \cdot \frac{1}{Z} $$

with the appropriate normalisation $Z$ from the sampling ratio. For logistic regression, undersampling the majority class by factor $\beta$ shifts only the intercept, by $\log \beta$ — every other coefficient is unchanged in expectation. This is the analytic reason class_weight="balanced" left ROC-AUC and PR-AUC untouched in the developer experiment: for a linear model it is close to an intercept shift.

Where imbalance genuinely becomes a learning problem

Three regimes where reweighting does more than move a threshold:

  1. Absolute scarcity. Not the ratio but the count. Fifty positives is a small-sample estimation problem, and no reweighting adds examples.
  2. Non-convex losses and deep networks. With finite optimisation budgets, gradients dominated by the majority class can stall minority-class learning before convergence. Cui et al. (2019), Class-Balanced Loss Based on Effective Number of Samples, CVPR, reweight by $(1-\beta)/(1-\beta^{n_c})$ rather than by $1/n_c$, arguing for diminishing returns from additional samples of the same class.
  3. Separable-but-imbalanced margins. Cao et al. (2019), Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss, NeurIPS, derive a class-dependent margin $\Delta_j \propto n_j^{-1/4}$ from a generalisation bound, and show gains that threshold tuning cannot reproduce.

Focal loss (Lin et al., 2017, Focal Loss for Dense Object Detection) down-weights easy examples by $(1-p_t)^{\gamma}$. It was designed for extreme foreground-background imbalance in dense detection, roughly 1000:1, where almost every anchor is an easy negative. Applying it to mildly imbalanced tabular data is cargo cult.

Rare-events bias in logistic regression

King and Zeng (2001), Logistic Regression in Rare Events Data, Political Analysis, show that maximum likelihood underestimates event probabilities when events are rare, with a bias of order $O(1/n)$ that is not negligible at small event counts. The bias-corrected estimator is:

$$ \tilde{\beta} = \hat{\beta} - \text{bias}(\hat{\beta}), \qquad \text{bias}(\hat{\beta}) = (X^\top W X)^{-1} X^\top W \xi $$

with $\xi_i = 0.5 \, Q_{ii}\big[(1+w_1)\hat{\pi}_i - w_1\big]$, $Q$ the hat matrix, $W$ the weight matrix and $w_1$ the sampling weight for events. This matters when the number of positives is in the tens or low hundreds, which is common in fraud and rare-disease work.

Resampling methods, assessed honestly

  • Random oversampling duplicates minority rows. Equivalent to instance reweighting for most losses. Increases overfitting risk on the duplicated points.
  • Random undersampling discards majority rows. Fast, and throws away real information; competitive when the majority is enormous and redundant.
  • SMOTE (Chawla et al., 2002, JAIR) interpolates linearly between a minority point and one of its $k$ minority neighbours. The synthetic points lie on segments inside the minority convex hull, so it cannot extrapolate to unseen regions. In high dimensions, distances concentrate and neighbour choice becomes near-arbitrary. For categorical or one-hot features, interpolation produces fractional values with no meaning.
  • ADASYN, Borderline-SMOTE focus synthesis near the boundary, which also concentrates synthesis near the noisiest points.

Elor and Averbuch-Elor (2022), To SMOTE, or not to SMOTE?, evaluate across many datasets and find that for strong classifiers with properly tuned thresholds, balancing gives no consistent gain in AUC or in cost-sensitive metrics. The gains reported in older papers largely disappear once the threshold is tuned rather than fixed at 0.5.

That is the finding the developer experiment reproduces in miniature.

Metrics

  • Average precision $= \sum_n (R_n - R_{n-1}) P_n$, the step-wise integral of the precision-recall curve. Its baseline is the positive class prevalence, so 0.285 against a prevalence of 0.0122 is roughly a 23-fold lift.
  • Precision@k when the review capacity is fixed at $k$ alerts per day. This is usually the honest operational metric and it is rarely reported.
  • Expected cost $= c_{10}\,\text{FP} + c_{01}\,\text{FN}$, evaluated over the full threshold sweep. If you can name the two costs, this dominates every other choice.
  • Calibration. Resampling and reweighting break calibration by construction. If a downstream system consumes the probability rather than the decision, recalibrate on the original prior with Platt scaling or isotonic regression.

Reading

  • Elkan, The Foundations of Cost-Sensitive Learning, IJCAI 2001.
  • King and Zeng, Logistic Regression in Rare Events Data, Political Analysis 2001.
  • Chawla et al., SMOTE: Synthetic Minority Over-sampling Technique, JAIR 2002 — arxiv.org/abs/1106.1813
  • Davis and Goadrich, The Relationship Between Precision-Recall and ROC Curves, ICML 2006.
  • Saito and Rehmsmeier, The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets, PLOS ONE 2015.
  • Cao et al., Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss, NeurIPS 2019 — arxiv.org/abs/1906.07413

What to learn next