Imbalanced, Multi-class and Multi-label
Classifier chains
Labels often travel together — a chain lets each label's model see the previous labels' answers, recovering the correlations independent models throw away.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A classifier chain predicts labels one after another, letting each prediction see the answers that came before it.
Think of doctors filling in a patient form as a relay. The first doctor answers "diabetes: yes/no" and passes the form on. The second answers "blood pressure: yes/no" — while reading the first answer, because diabetes makes high blood pressure likelier. Each specialist inherits every answer so far.
Why it exists
The standard multi-label recipe trains one independent specialist per label. Independence is its weakness. Real labels travel in packs: news tagged "politics" is often tagged "government"; a film tagged "horror" is rarely tagged "family".
Independent specialists cannot use this. The horror specialist never learns that the family switch matters to its own answer, because it never sees that switch.
A classifier chain fixes the blindness with one structural change. Label predictions are made in a fixed order. Each model receives the original inputs plus all previous labels' answers as extra features.
How it works
input features ──→ [ model A ] ──→ action? yes
input features + "action=yes" ──→ [ model B ] ──→ crime? no
input features + "action=yes, crime=no" ──→ [ model C ] ──→ romance? noDuring training, each model down the chain trains with the true earlier labels attached. During prediction, true labels do not exist — each model receives the chain's own earlier guesses. A wrong guess early on can mislead everything after it, like one wrong entry on the relay form. That risk is real, and the fix is running several chains in different orders and averaging.
A real example you have seen
Music apps tag songs with mood labels: energetic, sad, acoustic, danceable. These are heavily linked — "sad" and "danceable" rarely co-occur, "energetic" and "danceable" often do. Systems that respect label relationships produce tag sets that feel coherent, and chained prediction is one of the standard ways to get that.
Remember this
- A chain feeds earlier labels' answers into later labels' models.
- It exists because labels are correlated, and independent models waste that signal.
- Early mistakes can cascade — averaging chains in several random orders is the standard defence.
What to learn next
- Ordinal targets — when the labels are ordered instead of independent.
- Multi-label classification — the setting this lesson upgraded.
- Sequence labelling — predictions that depend on previous predictions, over text.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2, CPU only.
Chain versus independent, on correlated labels
from sklearn.datasets import make_multilabel_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import f1_score
from sklearn.model_selection import train_test_split
from sklearn.multioutput import ClassifierChain, MultiOutputClassifier
X, Y = make_multilabel_classification(n_samples=600, n_features=10,
n_classes=4, n_labels=2,
random_state=42)
X_tr, X_te, Y_tr, Y_te = train_test_split(X, Y, random_state=42)
base = LogisticRegression(max_iter=1000)
indep = MultiOutputClassifier(base).fit(X_tr, Y_tr) # labels ignore each other
chain = ClassifierChain(base, order=[0, 1, 2, 3],
random_state=42).fit(X_tr, Y_tr) # each label sees the previous ones
for name, m in [("independent", indep), ("chain ", chain)]:
pred = m.predict(X_te)
exact = (pred == Y_te).all(axis=1).mean() # whole row must match
print(f"{name} micro-F1={f1_score(Y_te, pred, average='micro'):.2f}"
f" exact-match={exact:.2f}")independent micro-F1=0.79 exact-match=0.49 chain micro-F1=0.78 exact-match=0.55
The walkthrough
Look at which metric moved. Micro-F1 — a per-switch score — barely changed. Exact-match — the whole row must be right — climbed six points. This is the textbook signature of chains: they improve joint consistency of the label set, not each switch in isolation. Theory says exactly this should happen; see the researcher block.
ClassifierChain versus MultiOutputClassifier. Same constructor pattern, one difference inside: the chain appends columns of previous label outputs to X for each successive model. Model k trains on 10 + k features here.
order controls the relay sequence. [0, 1, 2, 3] is set explicitly so results are reproducible. Order matters: labels that are easy to predict and informative about others belong early. Unknown? Use random orders — see below.
Ensemble of chains. The standard production form averages 10 chains with shuffled orders:
import numpy as np
chains = [ClassifierChain(base, order="random", random_state=i).fit(X_tr, Y_tr)
for i in range(10)]
avg = np.mean([c.predict_proba(X_te) for c in chains], axis=0)
pred = (avg > 0.5).astype(int)
exact = (pred == Y_te).all(axis=1).mean()
print(f"chain ensemble micro-F1={f1_score(Y_te, pred, average='micro'):.2f}"
f" exact-match={exact:.2f}")chain ensemble micro-F1=0.80 exact-match=0.58
Averaging washes out any single unlucky ordering. The ensemble beats both earlier models on both scores. This is the form Read's original paper recommends.
Common mistakes
Expecting per-label metrics to improve. If Hamming loss or micro-F1 is your target, independent models are already theoretically adequate. Chains earn their complexity only when the set must be coherent — subset accuracy, downstream rules, user-facing tag groups.
Judging one fixed order as "the" chain result. A single order is one sample from a large space. Two orders can differ by several points of subset accuracy. Report the ensemble, or at minimum several random orders.
Forgetting exposure bias. Training uses true previous labels; prediction uses guessed ones. With weak early models the gap bites hard. Putting reliable labels first reduces it.
Chaining hundreds of labels. Cost grows linearly, and late models drag hundreds of noisy guessed features. Past a few dozen labels, prefer shared-encoder neural approaches or label trees.
Try it yourself
Build the 10-chain ensemble above and score its exact-match against the single chain. Then reverse the order to [3, 2, 1, 0] and measure how much a single chain's result swings on this dataset.
What to learn next
- Ordinal targets — when the labels are ordered instead of independent.
- Multi-label classification — the setting this lesson upgraded.
- Sequence labelling — predictions that depend on previous predictions, over text.
Researcher — Mathematics and papers.
The probabilistic view
Chains implement the exact factorisation P(Y|x) = Π_{j=1}^{L} P(y_j | x, y_1, …, y_{j−1}) — no independence assumption, unlike binary relevance which models each marginal P(y_j|x) alone. Read, Pfahringer, Holmes and Frank (2009; journal version 2011), Classifier chains for multi-label classification, Machine Learning 85, introduced the method and the ensemble-of-random-orders estimator (ECC).
Greedy prediction — committing to ŷ_j = argmax before moving on — does not recover the joint mode argmax_Y P(Y|x). Probabilistic classifier chains (Dembczyński, Cheng and Hüllermeier, 2010, ICML) perform exact inference by exploring the label tree: O(2^L) worst case, made practical by beam search or ε-approximate pruning (Kumar et al., 2013; Mena et al., 2015). The theory in Dembczyński et al. (2012) explains the experiment above: subset 0/1 loss requires joint-mode inference (chains help), Hamming loss requires only marginals (they cannot, systematically).
Cost and order
Training: L base-model fits on d + j − 1 features; prediction is sequential per chain (parallel across ensemble members only). Error propagation is the structural risk: Senge et al. (2014) analyse how the probability of a correct full set decays with chain length and early-label error rates. Order-selection heuristics exist — easiest-first, mutual-information-greedy, genetic search (GACC, Gonçalves et al., 2013) — but ECC's averaging over ~10 random orders remains the robust default and is what order="random" supports.
Relatives and current practice
Label powerset captures the full joint at combinatorial cost; RAkEL (Tsoumakas and Vlahavas, 2007) ensembles powersets over random k-subsets of labels as a middle path. Conditional random fields and structured SVMs model pairwise label potentials explicitly. Neural sequence-to-set models (Yang et al., 2018, SGM) reinvent chains with an RNN decoder emitting labels in sequence — inheriting exposure bias and order sensitivity, treated with scheduled sampling and set-level losses. At extreme scale, chains give way to label trees and shared encoders, but for tens of correlated labels with tabular or modest text features, ECC over gradient-boosted bases is still a competitive, interpretable choice — and ships in scikit-learn.
What to learn next
- Ordinal targets — when the labels are ordered instead of independent.
- Multi-label classification — the setting this lesson upgraded.
- Sequence labelling — predictions that depend on previous predictions, over text.