Class imbalance
In one sentence Class imbalance is when one outcome vastly outnumbers the other in your data, letting lazy models score high while learning nothing.
Updated
Class imbalance is when the classes in your data are wildly unequal in count — 999 normal cases for every 1 interesting one.
Training a fraud model on real payments is like teaching someone to spot counterfeit notes by showing them ten thousand genuine notes and three fakes. The overwhelming lesson absorbed is "notes are genuine". The rare class — the entire reason the system exists — barely features in the training signal. Fraud, disease screening, defect detection and churn all live in this regime, because the events worth catching are rare by nature.
Two failures follow. The model under-learns the minority class. And your metrics lie: 99.7% accuracy is achievable by predicting "normal" always. The first fix is therefore measurement: judge by recall, precision, the confusion-matrix and precision-recall curves, never accuracy alone.
Then rebalance the learning pressure:
class weights penalise minority mistakes more heavily (one argument in sklearn)
oversampling repeat or synthesise minority examples (SMOTE)
undersampling drop majority examples
threshold tuning lower the decision cutoff for the rare classClass weights are the usual first move — cheap and effective. One warning: oversample inside the training folds only. Duplicating minority rows before the train-test-split puts copies of the same example on both sides — data-leakage in a new costume.
Where to go next
- Full lesson: Model evaluation
- Related terms: accuracy, recall, data-augmentation, confusion-matrix