Imbalanced, Multi-class and Multi-label
Multi-label classification
When one example can carry several correct answers at once, the target becomes a row of yes/no switches — and the metrics change with it.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Multi-label classification lets one example have several correct answers at the same time, instead of exactly one.
Think of a restaurant menu card. Each dish carries little symbols: spicy, vegetarian, contains nuts. A dish can be all three at once, or none. Nobody asks "is this dish spicy OR vegetarian?" — the tags are independent switches, each on or off.
Why it exists
Ordinary classification is a single-choice exam: cat or dog or horse, pick one. That is called multi-class classification — many options, one answer.
But a news article can be about politics and business. A film can be crime and romance. A chest X-ray can show two conditions at once. Forcing one answer throws information away, and the model gets punished for being half right.
Multi-label classification changes the target. Instead of one answer, each example gets a row of yes/no switches — one per possible tag. The model's job is setting every switch correctly.
How it works
action crime romance
"spies chase a bomb" → 1 0 0
"a heist love story" → 0 1 1
"detective in the rain"→ 0 1 0The most common recipe is familiar from one-vs-rest: train one yes/no specialist per tag. The action specialist ignores the other columns, and so on. At prediction time, every specialist answers, and you collect the yeses.
Grading also changes. Getting two switches right out of three is worth something now. A prediction can be partly correct — and metrics must decide how much partial credit to give.
A real example you have seen
Gmail can mark one email as Promotions and Social. YouTube tags a video with several topics at once. Netflix files one film under thriller, crime and drama together. Every tagging system where labels can pile up on one item is doing multi-label classification.
Remember this
- Multi-class = pick exactly one answer. Multi-label = set many switches, each yes or no.
- The target becomes a 0/1 grid: one row per example, one column per tag.
- Partial credit exists now, and choosing how to grade it is part of the job.
What to learn next
- Classifier chains — letting each label peek at the previous labels' answers.
- Macro, micro and weighted averaging — the averaging decisions behind every multi-label score.
- Text classification — the single-label pipeline this lesson generalised.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2, CPU only.
Films with many genres
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.multiclass import OneVsRestClassifier
from sklearn.preprocessing import MultiLabelBinarizer
films = [
"spies chase a stolen bomb across rooftops",
"a couple falls in love during a heist gone wrong",
"detectives hunt a serial killer in the rain",
"two friends road trip and fall in love",
"an agent defuses a bomb and saves the city",
"a killer stalks a detective who loves her",
]
genres = [{"action"}, {"romance", "crime"}, {"crime"},
{"romance"}, {"action"}, {"crime", "romance"}]
mlb = MultiLabelBinarizer()
Y = mlb.fit_transform(genres) # one 0/1 column per genre
print("label columns:", list(mlb.classes_))
print(Y)
vec = TfidfVectorizer()
X = vec.fit_transform(films)
model = OneVsRestClassifier(LogisticRegression()).fit(X, Y)
new = ["a detective falls in love during a bomb scare"]
pred = model.predict(vec.transform(new))
print("prediction:", mlb.inverse_transform(pred))label columns: ['action', 'crime', 'romance']
[[1 0 0]
[0 1 1]
[0 1 0]
[0 0 1]
[1 0 0]
[0 1 1]]
prediction: [('crime', 'romance')]The walkthrough
MultiLabelBinarizer builds the switch grid. Sets of tag names in, 0/1 matrix out — and inverse_transform turns predictions back into readable tags. Column order is alphabetical, stored in mlb.classes_.
Y is a matrix, not a column. That is the entire structural difference from ordinary classification. OneVsRestClassifier sees three columns and trains three independent logistic regressions, one per genre.
The prediction picked two tags. "Detective" pulled crime; "falls in love" pulled romance. The bomb was not enough for action — six training films is a tiny world. Real systems need hundreds of examples per tag.
Grading needs new metrics. For a labelled test set, three standards matter:
from sklearn.metrics import hamming_loss, accuracy_score, f1_score
# three films, three genre switches: columns are action / crime / romance
Y_true = [[1, 0, 0], [0, 1, 1], [0, 1, 0]]
Y_pred = [[1, 0, 0], [0, 1, 0], [0, 1, 1]] # row 2 misses romance, row 3 invents it
print(f"hamming loss {hamming_loss(Y_true, Y_pred):.3f} fraction of switches set wrong")
print(f"subset accuracy {accuracy_score(Y_true, Y_pred):.3f} fraction of rows perfect")
print(f"micro F1 {f1_score(Y_true, Y_pred, average='micro'):.3f} pooled over all switches")hamming loss 0.222 fraction of switches set wrong subset accuracy 0.333 fraction of rows perfect micro F1 0.750 pooled over all switches
Two wrong switches out of nine, yet subset accuracy reports two rows out of three as failures. Same predictions, three very different verdicts.
accuracy_score on multi-label targets is subset accuracy — brutally strict, all-or-nothing per row. Hamming loss gives per-switch partial credit. Micro-F1 sits between, and macro vs micro applies with full force since tag frequencies are always skewed.
Common mistakes
Reaching for LabelEncoder or one-hot encoding. Those encode one answer per row. Multi-label needs MultiLabelBinarizer, which allows several 1s per row. Feeding label lists into ordinary encoders raises errors or silently makes each unique combination its own fake class.
Stratified splitting the usual way. train_test_split(stratify=...) expects single labels. With label combinations, naive stratification breaks; use ShuffleSplit or the iterative stratification from the skmultilearn package for rare-tag balance.
Reporting subset accuracy alone. With 20 tags, even a strong model rarely gets all 20 switches right per row. Subset accuracy near zero can hide genuinely good per-tag performance — pair it with micro- and macro-F1.
Ignoring the empty prediction. All specialists can say no, predicting zero tags. Decide what your product does then: force the top-scoring tag, or allow "untagged" honestly.
Try it yourself
Add a fourth genre, {"comedy"}, with two new film descriptions, and retrain. Then compute hamming_loss and subset accuracy on the training set itself and observe how strict subset accuracy is even in-sample.
What to learn next
- Classifier chains — letting each label peek at the previous labels' answers.
- Macro, micro and weighted averaging — the averaging decisions behind every multi-label score.
- Text classification — the single-label pipeline this lesson generalised.
Researcher — Mathematics and papers.
Problem formulation
Given X and label space {0,1}^L for L labels, learn h: X → {0,1}^L from pairs (x_i, y_i). The joint distribution P(Y|x) has 2^L outcomes; every practical method trades expressiveness of label dependence against tractability:
- Binary relevance (the OvR construction above): model each P(y_j|x) independently. Cost O(L) models; ignores label correlations entirely.
- Label powerset: one multiclass problem over observed label combinations. Captures dependence exactly but classes explode combinatorially and rare combinations starve.
- Classifier chains (Read et al., 2009): factorise P(Y|x) = Π_j P(y_j | x, y_1…y_{j−1}) — the subject of the next lesson.
- Structured/embedding methods: low-rank label embeddings (SLEEC — Bhatia et al., 2015), label trees (Parabel, Bonsai), and shared-encoder neural nets with L sigmoid outputs and binary cross-entropy — today's default at scale.
Loss functions and their optimal predictors
Dembczyński et al. (2012), On label dependence and loss minimization in multi-label classification, is the key theory paper. Its central result: the Bayes-optimal predictor for Hamming loss needs only the marginals P(y_j|x) — thresholded at 1/2 — while the optimal predictor for subset 0/1 loss is the joint mode, argmax_y P(y|x). The two coincide only under label independence or degenerate distributions. Practical translation: binary relevance is provably sufficient for Hamming loss, and chain/powerset methods pay off only for subset-style objectives. Choosing an evaluation metric therefore chooses your architecture.
Per-label F-measures add further structure: optimal decisions threshold each marginal at a label-specific value tied to that label's F1 (Koyejo et al., 2015, Consistent multilabel classification, NeurIPS) — motivating per-tag threshold tuning on validation data rather than a global 0.5.
Evaluation zoo
Example-based: Hamming loss, subset accuracy, example-F1 (average="samples"). Label-based: micro/macro precision-recall-F1. Ranking-based (score vectors): ranking loss, coverage, average precision; at extreme scale, precision@k and propensity-scored variants (Jain et al., 2016) dominate, since only the top of each label ranking is ever consumed. Tag distributions are long-tailed in every real corpus, so macro metrics and tail-propensity corrections decide whether rare labels count at all.
State of the art
Extreme multi-label classification (L in the millions — ad queries, product tagging, Wikipedia labels) is its own subfield: tree-based Parabel/Bonsai, and transformer encoders with hierarchical label attention (X-Transformer, LightXML) lead benchmarks like Amazon-670K. For ordinary L (tens to thousands), a pretrained text encoder with L sigmoid heads and tuned per-label thresholds remains the strong, boring baseline — see BERT fine-tuning with problem_type="multi_label_classification" in HuggingFace.
What to learn next
- Classifier chains — letting each label peek at the previous labels' answers.
- Macro, micro and weighted averaging — the averaging decisions behind every multi-label score.
- Text classification — the single-label pipeline this lesson generalised.