Imbalanced, Multi-class and Multi-label

Undersampling strategies

Instead of inventing rare examples, undersampling throws away common ones — randomly, or with care about which ones carry information.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Undersampling balances a dataset by removing examples of the common class, instead of inventing examples of the rare one.

Think of preparing for an exam from a huge pile of past questions. Nine hundred questions cover one easy chapter, and forty cover the hard one. You do not need all nine hundred. You keep a sample of them, study all forty hard ones, and your practice time gets balanced.

Why it exists

SMOTE fixes imbalance by inventing rare examples. That worries people in serious settings: synthetic patients, synthetic frauds. Undersampling takes the opposite road. Every row the model sees is real — there are fewer of them.

It also solves a practical problem: size. Ten million normal transactions and ten thousand frauds make training slow. Keeping a fraction of the normal ones can shrink training time enormously while losing little.

The nagging worry is real too: thrown-away rows might have carried information. The clever strategies below choose which rows to discard, not how many.

How it works

before:   ████████████████████ common (many)
          ██ rare (few)

random undersampling — keep a random slice of the common class:
after:    ██ common
          ██ rare

Random undersampling keeps a random subset. Smarter strategies look at the border between classes:

  • Tomek links finds pairs of enemies that are each other's closest neighbours. Such pairs sit right on the boundary, confusing the model. It deletes the common-class member of each pair — a light cleanup.
  • Edited nearest neighbours removes common examples whose neighbours mostly disagree with them — likely noise.
  • NearMiss keeps only the common examples closest to the rare ones, where the decision actually happens.

A real example you have seen

Email spam filters train on billions of normal emails and far fewer spam campaigns of each new type. Nobody trains on every normal email ever sent. A sample of "normal" is more than enough, and the engineering effort goes into the rare, changing class.

Remember this

  • Undersampling removes common-class rows to balance the data. Nothing is invented.
  • Random removal is the baseline; Tomek links and friends choose which rows to remove.
  • The price is thrown-away information — safe when the common class is huge and repetitive.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn imbalanced-learn

Outputs verified with scikit-learn 1.7.2 and imbalanced-learn 0.14.2, CPU only.

Four strategies, one dataset

undersample.py
from collections import Counter

from imblearn.under_sampling import (EditedNearestNeighbours, NearMiss,
                                     RandomUnderSampler, TomekLinks)
from sklearn.datasets import make_classification

X, y = make_classification(n_samples=2000, n_features=6, weights=[0.95, 0.05],
                           class_sep=0.8, random_state=42)
print("original:", dict(Counter(y)))

samplers = [
    ("random", RandomUnderSampler(random_state=42)),
    ("tomek links", TomekLinks()),
    ("edited NN", EditedNearestNeighbours()),
    ("nearmiss-1", NearMiss(version=1)),
]
for name, s in samplers:
    _, y_res = s.fit_resample(X, y)
    print(f"{name:12s} -> {dict(Counter(y_res))}")
Output
original: {0: 1888, 1: 112}
random       -> {0: 112, 1: 112}
tomek links  -> {0: 1851, 1: 112}
edited NN    -> {0: 1704, 1: 112}
nearmiss-1   -> {0: 112, 1: 112}

The walkthrough

Two families are visible in the counts. Random and NearMiss force exact balance: 112 vs 112. Tomek links removed only 37 rows and edited nearest neighbours only 184 — they are cleaners, not balancers. They delete confusing boundary points and leave the ratio mostly alone.

Cleaners combine well with oversampling. The popular combination SMOTETomek runs SMOTE, then sweeps away the boundary noise SMOTE can create. SMOTEENN does the same with edited nearest neighbours.

RandomUnderSampler accepts sampling_strategy. Passing 0.5 keeps a 2-to-1 ratio instead of 1-to-1. Full balance is rarely the sweet spot; treat the ratio as a dial to tune.

Common mistakes

Undersampling a small dataset. With 2000 rows total, keeping 224 throws away 89% of your information, and scores get noisy. Undersampling shines when the majority class has hundreds of thousands of redundant rows.

Resampling before the train/test split. Same leak as with SMOTE — the test set must keep its natural imbalance, or your metrics describe a world that does not exist. See resampling inside cross-validation.

Expecting NearMiss to always help. NearMiss keeps majority points nearest the minority. If the minority class contains label noise, NearMiss faithfully keeps the majority points nearest the noise. Check its output with a quick scatter plot before trusting it.

Forgetting the probability shift. Train on a balanced diet and predicted probabilities no longer match real-world rates. If you need honest probabilities, recalibrate on untouched data.

Try it yourself

Train LogisticRegression on the output of each sampler and compare recall and precision on an untouched test split, reusing the pattern from the SMOTE lesson. Then try NearMiss(version=3) and read the imblearn docs to see how the three versions differ.

What to learn next

Researcher — Mathematics and papers.

Formal definitions

Tomek link (Tomek, 1976): a pair (x_i, x_j) with opposite labels such that no third sample x_k satisfies d(x_i, x_k) < d(x_i, x_j) or d(x_j, x_k) < d(x_i, x_j). Each is the other's nearest neighbour, so the pair straddles the boundary or one member is noise. Removing the majority member sharpens the margin.

Edited nearest neighbours (Wilson, 1972): delete every majority sample whose k nearest neighbours (k = 3 by default) majority-vote against its label. Repeated application gives RENN; AllKNN sweeps k = 1…k_max.

NearMiss (Mani and Zhang, 2003): version 1 keeps majority samples with the smallest mean distance to their m nearest minority samples; version 2 uses the m farthest minority samples; version 3 keeps a fixed number of majority guards around each minority point.

Condensed nearest neighbour (Hart, 1968): the earliest member — grow a subset that still classifies everything correctly under 1-NN, discarding the interior of homogeneous regions.

Cost

All neighbour-based strategies pay O(n² d) naive pairwise distance, or O(n log n · d) with ball trees, on n total samples and d features — the majority class dominates n, which is exactly the class you wanted to avoid iterating over. Random undersampling is O(n) and embarrassingly cheap. For datasets above a few million rows, random sampling per training batch (as in weighted samplers) replaces one-shot dataset surgery.

What the evidence says

Batista, Prati and Monard (2004), A study of the behavior of several methods for balancing machine learning training data, remains the standard comparison: combinations (SMOTE + Tomek, SMOTE + ENN) tended to beat single methods on small, hard datasets. Later work with strong boosted-tree learners finds the same caveat as for SMOTE — threshold tuning and class weighting close most of the gap. Undersampling's durable advantages are computational: it is the only balancing method that makes training cheaper, and it composes naturally into ensembles, one random subset per member — the subject of balanced bagging and EasyEnsemble.

Estimated probabilities shift under resampling: training on ratio r instead of true prior π biases the posterior. The correction p = p̂π(1−r) / (r(1−π) + p̂(π − r)) (Elkan, 2001, The foundations of cost-sensitive learning) restores calibration, or recalibrate empirically on an untouched validation set.

What to learn next