Imbalanced, Multi-class and Multi-label
Ordinal targets
Star ratings and severity grades have an order — treating them as unrelated classes wastes it, and a stack of yes/no questions gets it back.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An ordinal target is a label with a built-in order — and models that know the order make smaller mistakes when they are wrong.
Think of chilli ratings on a menu: mild, medium, hot, extra hot. The steps have a direction. Serving "hot" to someone who ordered "medium" is a small mistake. Serving "extra hot" to someone who ordered "mild" ruins their evening. The distance of the mistake matters.
Why it exists
Ordinary classification treats every class as an island: cat, dog, horse. Mixing up any two is equally wrong. Star ratings break that assumption. Predicting 4 stars for a 5-star review is nearly right; predicting 1 star is badly wrong. Plain classification cannot tell these mistakes apart — both count as one error.
Plain regression — covered in linear regression — fails differently. It pretends the steps are equal-sized numbers. But the gap between "mild" and "medium" is not promised to equal the gap between "hot" and "extra hot". The labels are ordered, not measured.
Ordinal targets sit exactly between: more structure than categories, less than numbers.
How it works
The oldest trick still works best as a mental model. Replace one 5-option question with four yes/no questions:
rating 0..4 becomes: is it above 0? is it above 1? is it above 2? is it above 3?
a true rating of 3 → yes yes yes no
prediction: count the yeses → threeEach question is an ordinary yes/no classifier. The order is baked in by construction: answering "above 2? yes" while "above 1? no" would be self-contradictory, and counting yeses keeps the answers consistent. Big jumps now require several questions to be wrong at once, so wild misses get rarer.
A real example you have seen
App-store ratings run 1 to 5 stars. Delivery feedback runs bad, okay, good, excellent. Disease severity grades and clothing sizes work the same way. All of these are ordinal. Any system predicting them — "you might rate this film 4 stars" — faces exactly this problem.
Remember this
- Ordinal = ordered categories: the steps have direction but not fixed sizes.
- Plain classification ignores the order; plain regression invents distances that are not there.
- The threshold trick — a stack of "is it above k?" questions — respects the order with ordinary tools.
What to learn next
- Macro, micro and weighted averaging — what standard averages hide about ordered mistakes.
- Logistic regression — the binary building block used K−1 times here.
- Linear regression — the other tempting-but-wrong default for ratings.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 and NumPy 1.26, CPU only.
Frank–Hall thresholds in fifteen lines
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(0)
X = rng.normal(size=(400, 4))
score = X @ [1.5, -1.0, 0.8, 0.5] + rng.normal(scale=2.0, size=400)
y = np.digitize(score, [-2, -0.5, 1, 2.5]) # 5 ordered ratings: 0..4
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0)
plain = LogisticRegression(max_iter=1000).fit(X_tr, y_tr)
plain_pred = plain.predict(X_te)
thresholds = [0, 1, 2, 3]
above = [LogisticRegression(max_iter=1000).fit(X_tr, y_tr > k) for k in thresholds]
p_above = np.column_stack([m.predict_proba(X_te)[:, 1] for m in above])
ord_pred = (p_above > 0.5).sum(axis=1)
for name, pred in [("plain multiclass", plain_pred), ("ordinal (F&H) ", ord_pred)]:
mae = mean_absolute_error(y_te, pred)
big = (np.abs(pred - y_te) >= 2).mean()
print(f"{name} MAE={mae:.2f} off-by-2-or-more={big:.1%}")plain multiclass MAE=0.97 off-by-2-or-more=30.0% ordinal (F&H) MAE=0.91 off-by-2-or-more=21.0%
The walkthrough
The headline is the second column. Mean absolute error improved a little, but bad misses — off by two grades or more — fell from 30% to 21%. Respecting the order pulls in the tails, which is usually what ordinal applications care about: a 5 predicted as 4 is tolerable, as 1 is a support ticket.
The data is genuinely ordinal by construction. A hidden continuous score gets cut at thresholds by np.digitize — the classic latent-variable story behind ratings. That is why y_tr > k is meaningful: the classes sort along a real axis.
(p_above > 0.5).sum(axis=1) counts cleared thresholds. Four probabilities per row; each above a half counts as one cleared step. Nothing forces the four probabilities to be monotone — model 3 could exceed model 2 — but counting sidesteps small violations gracefully.
Evaluate with distance-aware metrics. accuracy treats off-by-one and off-by-four alike. Use MAE on the grade index, the off-by-k rate, or quadratic weighted kappa — the standard on medical-grading leaderboards.
Common mistakes
One-hot encoding the target mentally. The moment you treat grades as unrelated columns, the order is gone. Ordinal handling lives in the target design, not the features.
Plain regression plus rounding as the only attempt. It sometimes works, but it assumes equal step sizes and punishes boundary cases oddly. If you do it, compare against Frank–Hall honestly before shipping.
Inconsistent hand-rolled thresholds. If you predict by first threshold failed rather than counting, non-monotone probabilities produce contradictions. Count cleared thresholds, or enforce monotonicity.
Metrics that hide the order. Reporting macro-F1 alone on an ordinal task discards exactly the structure you modelled. Pair it with MAE or weighted kappa. The averaging lesson explains what macro-F1 does and does not see.
Try it yourself
Add plain LinearRegression on y with rounding and clipping to 0–4, and put it in the comparison. Then shrink the training set to 100 rows and watch which method's off-by-2 rate degrades fastest.
What to learn next
- Macro, micro and weighted averaging — what standard averages hide about ordered mistakes.
- Logistic regression — the binary building block used K−1 times here.
- Linear regression — the other tempting-but-wrong default for ratings.
Researcher — Mathematics and papers.
Latent-variable formulation
The cumulative-link model posits a latent score z = w·x + ε and thresholds θ₁ < θ₂ < … < θ_{K−1} cutting z into K grades: y = k iff θ_k < z ≤ θ_{k+1}. With logistic ε this is proportional odds (McCullagh, 1980, Regression models for ordinal data, JRSS-B): P(y ≤ k | x) = σ(θ_k − w·x), where σ is the logistic function, w is shared across thresholds and only the intercepts θ_k differ. The shared-w assumption ("parallel lines") is testable (Brant, 1990); when violated, partial proportional-odds models free w per threshold.
Frank and Hall (2001), A simple approach to ordinal classification, ECML, is the decomposition coded above: K−1 independent binary problems P(y > k | x), recombined by P(y = k) = P(y > k−1) − P(y > k). Independence buys flexibility (any base learner) and loses guaranteed monotonicity; clipping or isotonic projection restores it.
Losses and consistency
Ordinal evaluation uses MAE over ranks, quadratic weighted kappa (Cohen, 1968), and off-by-k rates. Surrogate losses with consistency guarantees: the all-thresholds / immediate-threshold family (Rennie and Srebro, 2005), cumulative-link likelihoods, and ordinal hinge variants; Pedregosa, Bach and Gramfort (2017), On the consistency of ordinal regression methods, JMLR, establish which surrogates are Fisher-consistent for MAE-type risks. For deep models, CORAL (Cao, Mirjalili and Raschka, 2020) implements rank-consistent shared-weight thresholds in a single network head, and CORN (Shi et al., 2022) relaxes CORAL's shared-bias constraint via conditional training — both standard for tasks like diabetic-retinopathy grading and age estimation.
Complexity and practice
Frank–Hall costs K−1 base fits; cumulative-link MLE is a single convex optimisation of d + K − 1 parameters. Python tooling: statsmodels OrderedModel for proportional odds; the mord package (Pedregosa) for threshold-based scikit-learn-style estimators; coral-pytorch for deep variants. Gradient-boosted trees have no native ordinal loss in mainstream libraries, so Frank–Hall stacks of boosted binaries or regression-on-rank with tuned cut-points remain the tabular workhorses. Kappa-style metrics reward matching the marginal grade distribution, so post-hoc threshold tuning on validation data (optimising cut-points against QWK) reliably adds points on graded competitions.
What to learn next
- Macro, micro and weighted averaging — what standard averages hide about ordered mistakes.
- Logistic regression — the binary building block used K−1 times here.
- Linear regression — the other tempting-but-wrong default for ratings.