AI glossary

Class imbalance

In one sentence Class imbalance is when one outcome vastly outnumbers the other in your data, letting lazy models score high while learning nothing.

By Updated

Class imbalance is when the classes in your data are wildly unequal in count — 999 normal cases for every 1 interesting one.

Training a fraud model on real payments is like teaching someone to spot counterfeit notes by showing them ten thousand genuine notes and three fakes. The overwhelming lesson absorbed is "notes are genuine". The rare class — the entire reason the system exists — barely features in the training signal. Fraud, disease screening, defect detection and churn all live in this regime, because the events worth catching are rare by nature.

Two failures follow. The model under-learns the minority class. And your metrics lie: 99.7% accuracy is achievable by predicting "normal" always. The first fix is therefore measurement: judge by recall, precision, the confusion-matrix and precision-recall curves, never accuracy alone.

Then rebalance the learning pressure:

class weights      penalise minority mistakes more heavily (one argument in sklearn)
oversampling       repeat or synthesise minority examples (SMOTE)
undersampling      drop majority examples
threshold tuning   lower the decision cutoff for the rare class

Class weights are the usual first move — cheap and effective. One warning: oversample inside the training folds only. Duplicating minority rows before the train-test-split puts copies of the same example on both sides — data-leakage in a new costume.

Where to go next