Imbalanced, Multi-class and Multi-label
SMOTE and its variants
SMOTE invents new examples of a rare class by blending pairs of real ones, so the model finally gets enough practice on the cases that matter.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
SMOTE invents extra examples of a rare class by blending pairs of real examples, the way you mix two paint shades to get one in between.
Picture mixing paint. You have two similar shades of blue, and you stir a little of each together. You get a new shade that sits between them — new, but believable. SMOTE does this with data points instead of paint.
Why it exists
Imagine teaching a bank's computer to spot fraud. You show it a million payments, and only two thousand are fraud. The computer learns a lazy trick: say "not fraud" every time. It is right 99.8% of the time and completely useless.
The class you care about — fraud, rare disease, machine failure — is called the minority class, the one with few examples. The problem of one class drowning out another is called class imbalance. We met this problem in imbalanced data. SMOTE is one of the standard fixes.
You could copy the rare examples again and again. But exact copies teach the model to memorise those specific cases. SMOTE — Synthetic Minority Over-sampling Technique — makes new examples instead, each one a blend of two real neighbours.
How it works
real fraud A ●───────────● real fraud B
▲
new point ◆
(somewhere on the line
between A and B)Take a rare example. Find one of its closest rare neighbours. Place a new point at a random spot between them. Repeat until the rare class is as big as you want.
The blend only ever happens between two rare examples. The common class is left alone.
A real example you have seen
Hospitals train models to flag rare diseases from scan results. A rare condition might appear in one scan out of a thousand. Teams routinely balance such training data with SMOTE-style methods before training. The same idea appears in fraud detection at banks and in predicting machine breakdowns in factories.
Remember this
- SMOTE invents new rare-class examples by blending pairs of real neighbours.
- It exists because models ignore a class they rarely see.
- Blending beats copying: copies invite memorising, blends add variety.
What to learn next
- Undersampling strategies — the opposite move: shrink the majority instead.
- Class weights — the zero-new-data fix to attempt first.
- Resampling inside cross-validation — the leak that makes SMOTE results look better than they are.
Developer — Code and libraries.
Setup
pip install scikit-learn imbalanced-learnOutputs verified with scikit-learn 1.7.2 and imbalanced-learn 0.14.2, CPU only. The exact scores below depend on those versions; small shifts after upgrades are normal.
Balance the training set, never the test set
from collections import Counter
from imblearn.over_sampling import SMOTE
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import precision_score, recall_score
from sklearn.model_selection import train_test_split
# 1000 transactions, 5% fraud — small enough to run instantly on CPU
X, y = make_classification(n_samples=1000, n_features=6, weights=[0.95, 0.05],
class_sep=0.8, random_state=42)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=42)
# SMOTE touches ONLY the training set — the test set stays real
X_res, y_res = SMOTE(random_state=42).fit_resample(X_tr, y_tr)
print("before:", dict(Counter(y_tr)))
print("after: ", dict(Counter(y_res)))
plain = LogisticRegression(max_iter=1000).fit(X_tr, y_tr)
smoted = LogisticRegression(max_iter=1000).fit(X_res, y_res)
for name, model in [("plain", plain), ("smote", smoted)]:
pred = model.predict(X_te)
print(f"{name}: recall={recall_score(y_te, pred):.2f}"
f" precision={precision_score(y_te, pred):.2f}")before: {0: 711, 1: 39}
after: {0: 711, 1: 711}
plain: recall=0.08 precision=1.00
smote: recall=0.85 precision=0.20The walkthrough
fit_resample returns a bigger dataset. 39 real fraud rows became 711 — the originals plus synthetic blends, matching the majority count.
Read the two result lines together. The plain model caught 8% of frauds. After SMOTE it caught 85% — but precision fell from 1.00 to 0.20. SMOTE did not make the model smarter. It moved the trade-off: many more frauds caught, many more false alarms raised. Whether that trade is good depends on what a missed fraud costs versus a false alarm. Model evaluation covers these two metrics in depth.
stratify=y in the split keeps the fraud percentage identical in train and test. Without it, a small test set can end up with almost no fraud cases at all.
Common mistakes
Resampling before splitting. Run SMOTE on the full dataset and synthetic points leak into the test set. Blends of training rows sit in your test data, and scores inflate. The fix, with proof, is in resampling inside cross-validation.
Using SMOTE on categorical columns. Blending "city = Delhi" with "city = Pune" produces a meaningless decimal. Use SMOTENC for mixed columns or SMOTEN for all-categorical data. Plain SMOTE assumes every feature is numeric and continuous.
Reporting accuracy afterwards. On imbalanced data, accuracy was already lying before SMOTE. Report recall, precision, or balanced accuracy instead.
Reaching for SMOTE first. Attempt class_weight="balanced" before any resampling — the class weights lesson shows it matching SMOTE here without inventing a single row.
Try it yourself
Change weights=[0.95, 0.05] to [0.99, 0.01] and rerun. Watch what happens to precision when the class gets rarer. Then pass sampling_strategy=0.5 to SMOTE, which stops at a 2-to-1 ratio instead of full balance, and compare.
What to learn next
- Undersampling strategies — the opposite move: shrink the majority instead.
- Class weights — the zero-new-data fix to attempt first.
- Resampling inside cross-validation — the leak that makes SMOTE results look better than they are.
Researcher — Mathematics and papers.
The interpolation rule
For a minority sample x_i, SMOTE picks one of its k nearest minority neighbours x_zi (default k = 5, Euclidean metric), draws λ ~ Uniform(0, 1), and emits:
x_new = x_i + λ · (x_zi − x_i)
Where x_i is the seed sample, x_zi the chosen neighbour, and λ the random blend fraction. Every synthetic point lies on a line segment between two real minority points, so SMOTE populates the convex hull of existing minority clusters. It cannot generate outside that hull, and it ignores where the majority class sits.
Chawla, Bowyer, Hall and Kegelmeyer (2002), SMOTE: Synthetic Minority Over-sampling Technique, JAIR 16.
Cost
Neighbour search dominates: O(n_min² · d) naive, or O(n_min log n_min · d) with tree indexes, for n_min minority samples and d features. Generation itself is O(G · d) for G synthetic points. Memory grows by the synthetic count. For high-dimensional data (d in the hundreds), distance concentration makes "nearest" neighbours nearly arbitrary — a known failure mode.
Variants worth knowing
- Borderline-SMOTE (Han et al., 2005) — only seeds from minority points near the class boundary, where the classifier is actually contested.
- ADASYN (He et al., 2008) — generates more synthetics for minority points surrounded by majority neighbours, shifting density toward hard regions.
- SVM-SMOTE (Nguyen et al., 2011) — uses support vectors to find the boundary before interpolating.
- SMOTENC / SMOTEN (Chawla et al., 2002, §6; imblearn implementations) — handle mixed and purely categorical features via mode-voting on categories.
- KMeans-SMOTE (Douzas et al., 2018) — clusters first, then oversamples within clusters, avoiding blends across disjoint minority islands.
What the evidence actually says
Large-scale comparisons temper the enthusiasm. Elor and Averbuch-Elor (2022), To SMOTE, or not to SMOTE?, find that with strong classifiers (gradient-boosted trees) and proper threshold tuning, SMOTE rarely beats class weighting or plain threshold moving on tabular benchmarks. SMOTE also distorts calibrated probabilities: the model sees a 50/50 world and its scores must be recalibrated (van den Goorbergh et al., 2022, The harm of class imbalance corrections for risk prediction models). Prefer it when the minority class is tiny in absolute terms (tens of rows) and the learner cannot weight classes; otherwise weighting is the better default.
What to learn next
- Undersampling strategies — the opposite move: shrink the majority instead.
- Class weights — the zero-new-data fix to attempt first.
- Resampling inside cross-validation — the leak that makes SMOTE results look better than they are.