Imbalanced, Multi-class and Multi-label

Ordinal targets

Star ratings and severity grades have an order — treating them as unrelated classes wastes it, and a stack of yes/no questions gets it back.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An ordinal target is a label with a built-in order — and models that know the order make smaller mistakes when they are wrong.

Think of chilli ratings on a menu: mild, medium, hot, extra hot. The steps have a direction. Serving "hot" to someone who ordered "medium" is a small mistake. Serving "extra hot" to someone who ordered "mild" ruins their evening. The distance of the mistake matters.

Why it exists

Ordinary classification treats every class as an island: cat, dog, horse. Mixing up any two is equally wrong. Star ratings break that assumption. Predicting 4 stars for a 5-star review is nearly right; predicting 1 star is badly wrong. Plain classification cannot tell these mistakes apart — both count as one error.

Plain regression — covered in linear regression — fails differently. It pretends the steps are equal-sized numbers. But the gap between "mild" and "medium" is not promised to equal the gap between "hot" and "extra hot". The labels are ordered, not measured.

Ordinal targets sit exactly between: more structure than categories, less than numbers.

How it works

The oldest trick still works best as a mental model. Replace one 5-option question with four yes/no questions:

rating 0..4  becomes:   is it above 0?   is it above 1?   is it above 2?   is it above 3?

a true rating of 3  →       yes              yes              yes              no
prediction: count the yeses → three

Each question is an ordinary yes/no classifier. The order is baked in by construction: answering "above 2? yes" while "above 1? no" would be self-contradictory, and counting yeses keeps the answers consistent. Big jumps now require several questions to be wrong at once, so wild misses get rarer.

A real example you have seen

App-store ratings run 1 to 5 stars. Delivery feedback runs bad, okay, good, excellent. Disease severity grades and clothing sizes work the same way. All of these are ordinal. Any system predicting them — "you might rate this film 4 stars" — faces exactly this problem.

Remember this

  • Ordinal = ordered categories: the steps have direction but not fixed sizes.
  • Plain classification ignores the order; plain regression invents distances that are not there.
  • The threshold trick — a stack of "is it above k?" questions — respects the order with ordinary tools.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 and NumPy 1.26, CPU only.

Frank–Hall thresholds in fifteen lines

ordinal.py
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(0)
X = rng.normal(size=(400, 4))
score = X @ [1.5, -1.0, 0.8, 0.5] + rng.normal(scale=2.0, size=400)
y = np.digitize(score, [-2, -0.5, 1, 2.5])   # 5 ordered ratings: 0..4

X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0)

plain = LogisticRegression(max_iter=1000).fit(X_tr, y_tr)
plain_pred = plain.predict(X_te)

thresholds = [0, 1, 2, 3]
above = [LogisticRegression(max_iter=1000).fit(X_tr, y_tr > k) for k in thresholds]
p_above = np.column_stack([m.predict_proba(X_te)[:, 1] for m in above])
ord_pred = (p_above > 0.5).sum(axis=1)

for name, pred in [("plain multiclass", plain_pred), ("ordinal (F&H)   ", ord_pred)]:
    mae = mean_absolute_error(y_te, pred)
    big = (np.abs(pred - y_te) >= 2).mean()
    print(f"{name}  MAE={mae:.2f}  off-by-2-or-more={big:.1%}")
Output
plain multiclass  MAE=0.97  off-by-2-or-more=30.0%
ordinal (F&H)     MAE=0.91  off-by-2-or-more=21.0%

The walkthrough

The headline is the second column. Mean absolute error improved a little, but bad misses — off by two grades or more — fell from 30% to 21%. Respecting the order pulls in the tails, which is usually what ordinal applications care about: a 5 predicted as 4 is tolerable, as 1 is a support ticket.

The data is genuinely ordinal by construction. A hidden continuous score gets cut at thresholds by np.digitize — the classic latent-variable story behind ratings. That is why y_tr > k is meaningful: the classes sort along a real axis.

(p_above > 0.5).sum(axis=1) counts cleared thresholds. Four probabilities per row; each above a half counts as one cleared step. Nothing forces the four probabilities to be monotone — model 3 could exceed model 2 — but counting sidesteps small violations gracefully.

Evaluate with distance-aware metrics. accuracy treats off-by-one and off-by-four alike. Use MAE on the grade index, the off-by-k rate, or quadratic weighted kappa — the standard on medical-grading leaderboards.

Common mistakes

One-hot encoding the target mentally. The moment you treat grades as unrelated columns, the order is gone. Ordinal handling lives in the target design, not the features.

Plain regression plus rounding as the only attempt. It sometimes works, but it assumes equal step sizes and punishes boundary cases oddly. If you do it, compare against Frank–Hall honestly before shipping.

Inconsistent hand-rolled thresholds. If you predict by first threshold failed rather than counting, non-monotone probabilities produce contradictions. Count cleared thresholds, or enforce monotonicity.

Metrics that hide the order. Reporting macro-F1 alone on an ordinal task discards exactly the structure you modelled. Pair it with MAE or weighted kappa. The averaging lesson explains what macro-F1 does and does not see.

Try it yourself

Add plain LinearRegression on y with rounding and clipping to 0–4, and put it in the comparison. Then shrink the training set to 100 rows and watch which method's off-by-2 rate degrades fastest.

What to learn next

Researcher — Mathematics and papers.

Latent-variable formulation

The cumulative-link model posits a latent score z = w·x + ε and thresholds θ₁ < θ₂ < … < θ_{K−1} cutting z into K grades: y = k iff θ_k < z ≤ θ_{k+1}. With logistic ε this is proportional odds (McCullagh, 1980, Regression models for ordinal data, JRSS-B): P(y ≤ k | x) = σ(θ_k − w·x), where σ is the logistic function, w is shared across thresholds and only the intercepts θ_k differ. The shared-w assumption ("parallel lines") is testable (Brant, 1990); when violated, partial proportional-odds models free w per threshold.

Frank and Hall (2001), A simple approach to ordinal classification, ECML, is the decomposition coded above: K−1 independent binary problems P(y > k | x), recombined by P(y = k) = P(y > k−1) − P(y > k). Independence buys flexibility (any base learner) and loses guaranteed monotonicity; clipping or isotonic projection restores it.

Losses and consistency

Ordinal evaluation uses MAE over ranks, quadratic weighted kappa (Cohen, 1968), and off-by-k rates. Surrogate losses with consistency guarantees: the all-thresholds / immediate-threshold family (Rennie and Srebro, 2005), cumulative-link likelihoods, and ordinal hinge variants; Pedregosa, Bach and Gramfort (2017), On the consistency of ordinal regression methods, JMLR, establish which surrogates are Fisher-consistent for MAE-type risks. For deep models, CORAL (Cao, Mirjalili and Raschka, 2020) implements rank-consistent shared-weight thresholds in a single network head, and CORN (Shi et al., 2022) relaxes CORAL's shared-bias constraint via conditional training — both standard for tasks like diabetic-retinopathy grading and age estimation.

Complexity and practice

Frank–Hall costs K−1 base fits; cumulative-link MLE is a single convex optimisation of d + K − 1 parameters. Python tooling: statsmodels OrderedModel for proportional odds; the mord package (Pedregosa) for threshold-based scikit-learn-style estimators; coral-pytorch for deep variants. Gradient-boosted trees have no native ordinal loss in mainstream libraries, so Frank–Hall stacks of boosted binaries or regression-on-rank with tuned cut-points remain the tabular workhorses. Kappa-style metrics reward matching the marginal grade distribution, so post-hoc threshold tuning on validation data (optimising cut-points against QWK) reliably adds points on graded competitions.

What to learn next

What to learn next

These follow on from what you just read.

  • Hypothesis Testing and Inference

    Hypothesis testing

    Hypothesis testing asks one question — could this result be plain luck? — and gives you a number for how surprised you should be.

  • Hypothesis Testing and Inference

    T-tests

    The t-test decides whether two averages differ by more than their noise — the most-used statistical test in the world, in its three everyday forms.

  • Hypothesis Testing and Inference

    The chi-squared test

    The chi-squared test works on counts — clicks, votes, defects — and asks whether the numbers in your table drifted too far from what chance would deal.