Classification
Classification is sorting things into named groups decided in advance, and every classifier is really an argument about where to draw the boundary between them.
- 20 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Classification is sorting things into named groups that were chosen in advance.
Think about the boundary wall between two neighbours' plots of land. Everything on one side belongs to one owner, everything on the other side to the other. The wall is the entire agreement — shift it by a metre and ownership changes with it.
Now picture a mango tree growing right on the line. That tree is the argument, every single year. Classification models live and die by the same two things. Where the wall goes, and what happens to the cases sitting on top of it.
The boundary is the model
Here is the idea that ties every classifier together.
A trained classifier is a wall drawn through the space of possible inputs. Feed it something new, look at which side of the wall it lands on, and read off the group.
Two models trained on identical data can disagree completely. Not because one saw different examples, but because they are only allowed to build different shapes of wall.
That is the real difference between the methods in this section. They are all drawing a boundary. They differ in what shapes of boundary they are capable of drawing.
Three shapes of the question
Before choosing a model, work out which of these you actually have. Getting this wrong wastes weeks.
Pick one of two. Spam or not spam. Fraud or genuine. Repaid or defaulted. This is the case logistic regression handles directly.
Pick one of many. Which of ten handwritten digits is this? Which of twenty languages is this message in? Exactly one answer is right, and the others are wrong.
Pick any number of many. A holiday photo can be tagged beach and sunset and friends, all at once. A film can be a comedy and a romance. Here the tags do not compete with each other.
That third one catches people out. Sometimes several answers are true together. Build a "pick one of many" model for that job and the tags are forced to fight. The true ones lose.
The shapes of wall
A STRAIGHT WALL A STAIRCASE WALL A WALL THAT HUGS
logistic regression decision tree nearest neighbours
. . . | # # # . . . . # # # . . . . # # #
. . . | # # # . . . . # # # . . . # # # #
. . . | # # # . . . +---# # . . . # # # #
. . . | # # # . . . | . # # . . . . # # #
. . . | # # # . . . | . # # . . . . . # #
One cut, at any angle. Only horizontal and Follows the crowd,
Fast and explainable. vertical cuts, so it bending wherever the
Cannot bend. makes right angles. examples happen to sit.None of these is the best one. A straight wall is a poor fit for a round patch of land. A staircase is clumsy when a single diagonal cut would do. The wall that hugs will happily wrap itself around a mistake in your data.
The model votes, you decide
A classifier does not really hand you a name. It hands you a score for every group, and the highest score wins.
That gap matters more than it sounds. Suppose a fruit sorter scores a mystery fruit at a little over four in ten for mango. Apple and banana each score a little under three in ten. Mango wins. Mango is also far from certain.
The winning group can win while the model is deeply confused. Always look at the scores, not only at the winner.
Where you have already seen it
- Your photo gallery tagging pictures with places, pets and people.
- A bank card being declined the moment a payment looks wrong.
- Language detection in a translation app, before it translates anything.
- A hospital triage system sorting patients into urgency levels.
- The spam folder, the most successful classifier ever deployed.
The honest part
Rare groups get quietly ignored. If one group in a hundred is fraud, a model can score wonderfully by never predicting fraud at all. This trap is important enough to have its own lesson — see model evaluation.
The list of groups has to be complete, and it never is. A model trained on ten kinds of animal will confidently call an eleventh kind one of its ten. It has no way to say "this is something new". Handling that is a live research problem, not a setting you can switch on.
The wall stops fitting when the world moves. A fraud classifier trained on last year's tricks is drawing last year's wall. Fraudsters change; the wall does not, until you retrain it.
Remember this
- Every classifier is a boundary, and each method can only draw certain shapes of boundary.
- Decide first whether you need one answer from two, one from many, or any number from many.
- The winning group can win with a low score, so read the scores and not only the winner.
What to learn next
- Decision trees — how the staircase boundary gets built, one question at a time.
- Model evaluation — why the accuracy on this page's first map was meaningless.
- Train, test and validation splits — measuring a classifier on data it has not memorised.
Developer — Code and libraries.
Setup
pip install scikit-learn numpySeeing the boundary, without a plotting library
The fastest way to understand a classifier is to look at the wall it drew. Most tutorials use matplotlib for this. We will print the wall as text instead, so it works over SSH, in a notebook, and on a phone.
The task: 81 locations on a nine-by-nine map. A mobile tower sits in the middle, and its signal covers a roughly round patch.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.neighbors import KNeighborsClassifier
# 81 locations on a 9 by 9 map. Label 1 means "the phone gets a signal here".
gx, gy = np.meshgrid(np.arange(9), np.arange(9))
X = np.c_[gx.ravel(), gy.ravel()].astype(float)
# The tower sits at (4, 4) and its signal reaches a roughly round patch
y = (((X[:, 0] - 4) ** 2 + (X[:, 1] - 4) ** 2) <= 5).astype(int)
def draw(labels):
# printed with the highest row first, the way a map is drawn
for row in labels.reshape(9, 9)[::-1]:
print(" " + "".join("#" if v else "." for v in row))
print("WHAT IS ACTUALLY TRUE")
draw(y)
for name, model in [
("LOGISTIC REGRESSION", LogisticRegression()),
("DECISION TREE, depth 3", DecisionTreeClassifier(max_depth=3, random_state=0)),
("NEAREST NEIGHBOURS, k=3", KNeighborsClassifier(n_neighbors=3)),
]:
model.fit(X, y)
print()
print(f"{name} (accuracy {model.score(X, y):.3f})")
draw(model.predict(X))WHAT IS ACTUALLY TRUE ......... ......... ...###... ..#####.. ..#####.. ..#####.. ...###... ......... ......... LOGISTIC REGRESSION (accuracy 0.741) ......... ......... ......... ......... ......... ......... ......... ......... ......... DECISION TREE, depth 3 (accuracy 0.827) ......... ......... ..####### ..####### ..####### ..####### ..####### ......... ......... NEAREST NEIGHBOURS, k=3 (accuracy 0.951) ......... ......... ....###.. ..#####.. ..#####.. ...####.. ....##... ......... .........
Read those four maps slowly
This one output contains most of what this lesson is trying to teach.
Logistic regression printed an empty map. It predicted "no signal" for all 81 locations, including the 21 that have signal. It found nothing whatsoever.
Its accuracy is 0.741. That is not a bad-looking number. It is the exact trap described in model evaluation: 60 of the 81 locations genuinely have no signal, so answering "no signal" every time scores 74 percent while being useless.
This is not a bug and not a failure of tuning. No straight line can enclose a round patch. The best straight-line answer available really is "everything is outside", so that is what it correctly found.
The tree drew a box. Its wall is made only of horizontal and vertical cuts, so a circle becomes a rectangle. With max_depth=3 it has a budget of a few cuts, and it spent them on a crude box that is too wide.
Nearest neighbours drew something round-ish and lumpy. It never learned a rule at all. It looks at whichever training points are closest, so the shape follows the data, warts and all.
Three models. Same rows, same columns, same labels. Three completely different walls.
Giving logistic regression a fair chance
Logistic regression's wall is always straight in the features you hand it. That is the sentence worth memorising, because it contains the escape route.
Hand it squared and multiplied versions of the columns, and a straight wall in that new space becomes a curved wall in the original one.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures
gx, gy = np.meshgrid(np.arange(9), np.arange(9))
X = np.c_[gx.ravel(), gy.ravel()].astype(float)
y = (((X[:, 0] - 4) ** 2 + (X[:, 1] - 4) ** 2) <= 5).astype(int)
def draw(labels):
for row in labels.reshape(9, 9)[::-1]:
print(" " + "".join("#" if v else "." for v in row))
# PolynomialFeatures(2) adds x*x, y*y and x*y alongside the originals.
# A round boundary needs the squared terms, so now one is reachable.
bent = make_pipeline(PolynomialFeatures(2), LogisticRegression(C=1000)).fit(X, y)
print(f"logistic regression on bent features (accuracy {bent.score(X, y):.3f})")
draw(bent.predict(X))logistic regression on bent features (accuracy 1.000) ......... ......... ...###... ..#####.. ..#####.. ..#####.. ...###... ......... .........
Every one of the 81 locations correct. The same model class that found nothing at all now reproduces the patch exactly.
Nothing about the algorithm changed. Three extra columns did all of it. This is why feature engineering tends to beat model shopping.
One caution before you get excited. PolynomialFeatures grows fast — with 20 original columns, degree 2 gives 230. Bending the wall this way is also the easiest route to overfitting that exists.
Picking one of many
Give y three or more distinct values and scikit-learn switches to a multi-class fit on its own. predict_proba then returns one column per group.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
# Fifteen fruits already sorted by hand.
# Column 0: weight in grams. Column 1: sweetness tasted, out of 10.
# The last row of each group deliberately breaks the pattern.
fruit = np.array([
[150, 8], [170, 9], [140, 7], [165, 6], [185, 7], # mango, last one heavy
[120, 7], [130, 6], [110, 8], [145, 5], [125, 4], # banana, last one dull
[180, 4], [200, 3], [190, 5], [160, 4], [135, 5], # apple, last one light
])
kind = np.array(["mango"] * 5 + ["banana"] * 5 + ["apple"] * 5)
clf = make_pipeline(StandardScaler(), LogisticRegression()).fit(fruit, kind)
# Three unsorted fruits. The third is deliberately awkward.
mystery = np.array([[168, 9], [115, 7], [155, 6]])
print("classes, in the order the columns come out:", clf.classes_)
print()
print("probability table (rows = the three mystery fruits):")
print(clf.predict_proba(mystery).round(3))
print()
print("winner in each row:", clf.predict(mystery))
print("every row adds to :", clf.predict_proba(mystery).sum(axis=1).round(6))
print("training accuracy :", round(clf.score(fruit, kind), 3))classes, in the order the columns come out: ['apple' 'banana' 'mango'] probability table (rows = the three mystery fruits): [[0.019 0.033 0.949] [0.036 0.795 0.169] [0.296 0.278 0.425]] winner in each row: ['mango' 'banana' 'mango'] every row adds to : [1. 1. 1.] training accuracy : 0.933
Three things to take from this table.
clf.classes_ is sorted, not the order you wrote them. Alphabetical order gives apple, banana, mango. Column 0 is apple. Assuming your own ordering here is a genuinely common bug, and it produces wrong answers that look plausible.
Every row adds to one. The groups compete. Raising one group's score must lower another's, which is exactly right when only one answer can be true.
Read the third row. Mango wins with 0.425, against 0.296 and 0.278. The winner has under half the probability, and the other two together beat it. predict reports mango with no hint of any of this.
If that fruit were a loan application or a medical scan, the 0.425 is the number a human needs to see. Route low-margin cases to a person instead of accepting the winner.
Picking any number of many
For tags that do not compete, pass a two-dimensional y — one column per tag, each holding a yes or no.
import numpy as np
from sklearn.tree import DecisionTreeClassifier
# Six films. Column 0: minutes of action. Column 1: number of songs.
films = np.array([[60, 1], [55, 0], [10, 6], [5, 7], [50, 5], [45, 6]])
# Two independent yes/no tags per film: [is it action?, is it a musical?]
tags = np.array([[1, 0], [1, 0], [0, 1], [0, 1], [1, 1], [1, 1]])
tagger = DecisionTreeClassifier(random_state=0).fit(films, tags)
new_films = np.array([[58, 0], [8, 7], [48, 6]])
print("predicted [action, musical]:")
print(tagger.predict(new_films))predicted [action, musical]: [[1 0] [0 1] [1 1]]
The third film comes back as [1 1] — action and musical. A "pick one of many" model could never produce that answer, because its columns are forced to add to one.
Not every estimator accepts a 2D y. Trees, forests and nearest neighbours do. For the ones that do not, wrap them in sklearn.multioutput.MultiOutputClassifier, which fits one model per tag.
Common mistakes
Using a "pick one" model for a "pick any" job. The single most expensive error on this page. Check whether two labels can be true at once before you choose anything else.
Trusting clf.classes_ to match your input order. It is sorted. Index into it rather than hard-coding column numbers.
Judging a classifier by accuracy. The empty map above scored 0.741. Read model evaluation before you report any single number.
Splitting an imbalanced dataset without stratify. A random split can leave a rare group entirely absent from one side. Pass stratify=y — see train, test and validation splits.
Feeding text categories straight in. Most estimators need numbers for the inputs. Encode with OneHotEncoder for unordered categories. Encoding "Delhi, Mumbai, Chennai" as 0, 1, 2 tells the model that Mumbai sits between the other two, which is meaningless. Labels in y are the exception — strings are accepted there, as fruit_sorter.py shows.
Ignoring class weights. When one group is rare, most classifiers accept class_weight="balanced", which makes rare-group errors count more heavily during fitting. It is one keyword and it is frequently the largest single improvement available.
Try it yourself
In boundary_map.py, sweep the tree's depth instead of comparing three models:
for d in (1, 3, 5, None):
model = DecisionTreeClassifier(max_depth=d, random_state=0).fit(X, y)
print()
print(f"depth {d} (accuracy {model.score(X, y):.3f})")
draw(model.predict(X))Watch the staircase get finer with every step:
depth 1 (accuracy 0.741) ......... ......... ......... ......... ......... ......... ......... ......... ......... depth 3 (accuracy 0.827) ......... ......... ..####### ..####### ..####### ..####### ..####### ......... ......... depth 5 (accuracy 0.951) ......... ......... ..#####.. ..#####.. ..#####.. ..#####.. ..#####.. ......... ......... depth None (accuracy 1.000) ......... ......... ...###... ..#####.. ..#####.. ..#####.. ...###... ......... .........
Read the four maps against each other. At depth 1 every square is empty — one cut cannot enclose anything, so the tree gives up and calls everything "no signal". At depth 3 a crude box appears, too wide on the right. At depth 5 the box is about the right size but still square-cornered. At depth None the tree has reproduced the round patch exactly.
Now the important question. Depth None scores a perfect 1.000. Is that the best model?
No — and you cannot tell from this output either way. Every one of those 81 locations was used for training, so a deep enough tree can memorise all of them. A perfect training score is a warning, not a result. Overfitting and underfitting explains why, and train, test and validation splits shows the fix.
What to learn next
- Decision trees — how the staircase boundary gets built, one question at a time.
- Model evaluation — why the accuracy on this page's first map was meaningless.
- Train, test and validation splits — measuring a classifier on data it has not memorised.
Researcher — Mathematics and papers.
The decision-theoretic core
Classification is a decision problem, not an estimation problem. Separating the two stages clarifies most practical disputes.
Stage one, estimate. Model the posterior P(y = k | x).
Stage two, decide. Choose an action minimising expected loss.
For a loss matrix L(k, j) — the cost of predicting j when the truth is k — the optimal rule is:
h(x) = argmin_j SUM_k L(k, j) * P(y = k | x)L(k, j)— cost incurred by answeringjwhen the true class iskP(y = k | x)— posterior probability of classkgiven featuresx
Under 0-1 loss, L(k, j) = 1[k != j], and this collapses to argmax_k P(y = k | x), the Bayes classifier. Its risk is the Bayes error, an irreducible floor set by genuine class overlap.
The binary asymmetric case gives the threshold directly:
predict 1 iff P(y = 1 | x) > L(0,1) / ( L(0,1) + L(1,0) )L(0,1)— cost of a false positive,L(1,0)— cost of a false negative
The default threshold of 0.5 is therefore the correct choice only when the two errors cost the same. Elkan (2001) shows that for cost-sensitive problems, adjusting the threshold on a calibrated probability dominates retraining with resampled data. Rebalancing the training set and then thresholding at 0.5 is a strictly worse-conditioned way to reach the same place.
Discriminative versus generative
Discriminative: model P(y | x) directly logistic regression, SVM, trees
Generative: model P(x | y) and P(y), naive Bayes, LDA, QDA
then apply Bayes' ruleNg and Jordan (2001) give the definitive comparison. The generative model carries higher asymptotic error when its assumptions are wrong, but converges at O(log d / n) rather than O(d / n). The curves cross. Naive Bayes wins in the small-sample, high-dimension regime; logistic regression wins once n is large.
Generative models also handle missing features and novel classes more gracefully, because they carry a density over x. Discriminative models have no notion of whether an input was plausible at all.
Reductions from multi-class to binary
| Scheme | Models fitted | Cost at prediction | Notes |
|---|---|---|---|
| One-vs-rest | K | K | Scores not jointly normalised; class imbalance in every sub-problem |
| One-vs-one | K(K-1)/2 | K(K-1)/2 | Each fit uses only two classes' data, so fits are small and fast |
| Softmax | 1 | 1 | Jointly normalised; the principled default |
| ECOC | L codewords | L | Error-correcting output codes, Dietterich & Bakiri (1995) |
ECOC is under-appreciated. Assign each class a binary codeword of length L, train one binary classifier per bit, and decode by minimum Hamming distance. With a code of good minimum distance, the ensemble corrects individual classifier errors. It reduces both bias and variance relative to one-vs-rest.
Softmax is the right default for probability quality. One-vs-one is competitive when per-fit cost is superlinear in n, since each of its many fits sees only two classes' worth of data — this is why SVC uses it.
Multi-label
Let Y be a subset of 2^K rather than an element of {1..K}.
Binary relevance: K independent binary models. Ignores label correlation.
Classifier chains: model k conditions on predictions 1..k-1. Order matters.
Label powerset: treat each observed label set as one class. Up to 2^K classes.Read et al. (2011) introduced ensembles of classifier chains with randomised orderings, which remains a strong baseline. Binary relevance is optimal for Hamming loss but not for subset-0/1 loss — Dembczyński et al. (2012) prove that the loss function determines whether modelling label dependence can help at all. This is the key theoretical result in the area and it is routinely overlooked.
Evaluation needs multi-label-specific measures: Hamming loss, subset accuracy, micro- and macro-averaged F1. Note that micro-averaging weights by label frequency and macro-averaging does not, so the two can rank systems differently.
Class imbalance
The literature is larger than the effect. van den Goorbergh et al. (2022) show that resampling degrades calibration while leaving discrimination roughly unchanged, and argue against it as a default in clinical prediction.
Ordered by evidence:
- Change the threshold, using a calibrated model. Nearly always sufficient.
- Cost-sensitive weighting in the loss, via
class_weight. Equivalent to threshold shifting for many models, but interacts with regularisation. - Resampling. SMOTE (Chawla et al., 2002) synthesises minority points along segments between neighbours. Widely used; it distorts the posterior and requires recalibration afterwards.
If resampling is applied, it must live inside the cross-validation loop, applied to training folds only. Oversampling before splitting places near-duplicates on both sides and inflates every score. imblearn.pipeline.Pipeline enforces this; the scikit-learn one does not.
Open-set recognition
The closed-world assumption — every test input belongs to one of the K training classes — is false in nearly every deployment.
Scheirer et al. (2013) formalised open-set recognition and introduced open space risk. Softmax outputs are unsuitable as novelty scores because they are normalised over known classes only; an input unlike anything seen still produces one large probability.
Practical approaches: OpenMax (Bendale & Boult, 2016) fits an extreme-value model to per-class activation distances; energy-based scores (Liu et al., 2020) outperform maximum-softmax baselines; deep ensembles remain a strong uncertainty baseline (Lakshminarayanan et al., 2017). None is solved. Treat any deployed classifier as needing a separate rejection mechanism.
Calibration and thresholds under shift
Prior shift — P(y) changes while P(x | y) does not — is the common case in fraud, disease and demand. Saerens et al. (2002) give an EM procedure that corrects the posterior using only unlabelled test data:
P_new(y = k | x) ∝ P_old(y = k | x) * ( pi_new(k) / pi_old(k) )pi_old(k),pi_new(k)— class priors in the training and deployment distributions
Only the intercept requires adjustment for a logistic model, which is why prior shift is far more tractable than covariate shift.
Key references
- Dietterich, T. & Bakiri, G. (1995). Solving Multiclass Learning Problems via Error-Correcting Output Codes. JAIR 2.
- Elkan, C. (2001). The Foundations of Cost-Sensitive Learning. IJCAI.
- Ng, A. & Jordan, M. (2001). On Discriminative vs. Generative Classifiers. NeurIPS 14.
- Saerens, M., Latinne, P. & Decaestecker, C. (2002). Adjusting the Outputs of a Classifier to New a Priori Probabilities. Neural Computation 14(1).
- Chawla, N. et al. (2002). SMOTE: Synthetic Minority Over-sampling Technique. JAIR 16.
- Read, J. et al. (2011). Classifier Chains for Multi-label Classification. Machine Learning 85(3).
- Dembczyński, K. et al. (2012). On Label Dependence and Loss Minimization in Multi-label Classification. Machine Learning 88(1).
- Scheirer, W. et al. (2013). Toward Open Set Recognition. IEEE TPAMI 35(7).
- Lakshminarayanan, B., Pritzel, A. & Blundell, C. (2017). Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles. NeurIPS.
- van den Goorbergh, R. et al. (2022). The Harm of Class Imbalance Corrections for Risk Prediction Models. JAMIA 29(9).
Current state
For tabular classification, gradient-boosted trees remain the strongest default. Grinsztajn et al. (2022) analyse why neural networks still trail on tabular data: rotational invariance is the wrong inductive bias when features are individually meaningful, and trees handle irregular target functions and uninformative features better.
For perceptual inputs, the pipeline has inverted. Rather than training a classifier end to end, the standard approach is to take a pretrained representation and fit a light head on it, or to skip fitting entirely and use a zero-shot model such as CLIP (Radford et al., 2021), which classifies against arbitrary text labels supplied at prediction time. That relaxes the closed-set constraint at the interface, though not the underlying problem.
The unresolved issues are not accuracy. They are calibrated uncertainty, subgroup performance disparity, and knowing when to abstain.
What to learn next
- Decision trees — how the staircase boundary gets built, one question at a time.
- Model evaluation — why the accuracy on this page's first map was meaningless.
- Train, test and validation splits — measuring a classifier on data it has not memorised.