Supervised learning
Supervised learning is training a model on examples that already carry the correct answer, so it can answer new questions of the same kind.
- 12 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Supervised learning is teaching a model using examples that already carry the correct answer.
Think of a practice paper with the answer key printed at the back. You attempt a question, flip to the back, and compare. Where you were wrong, you correct yourself. After enough questions, you stop needing the key.
That is supervised learning, exactly. The answer key is called the label — the correct answer attached to each example. The word "supervised" points at the key. Something is standing behind the learner, marking every attempt.
Why it exists
Almost every useful prediction has a past you can learn from.
A bank has years of loan records, and it knows which loans were repaid. A hospital has scans, and it knows which patients turned out to be ill. Gmail has billions of emails that people marked as spam.
In each case the answers already exist. They were expensive to collect, and they are sitting there unused. Supervised learning is the method that turns that history into predictions about tomorrow.
The two kinds of question
Supervised learning answers two shapes of question, and telling them apart is the first skill to build.
Which bucket? — this is called classification, meaning sorting things into named groups.
- Is this email spam or not spam?
- Is this photo a cat, a dog, or a horse?
- Will this customer leave or stay?
How much? — this is called regression, meaning predicting a number on a sliding scale.
- What will this flat sell for?
- How many units will we ship next month?
- How many minutes until the delivery arrives?
A quick test. If the answer could sit between two of your options, it is regression. A price of 61.5 lakh is meaningful. A "half cat" is not.
How it works
TRAINING — the model sees the answers
Example Label
hours studied, attendance
------------------------- ------
3 hrs, 60 out of 100 → failed ─┐
7 hrs, 85 out of 100 → passed ├→ Learning → Trained
2 hrs, 50 out of 100 → failed │ algorithm model
9 hrs, 95 out of 100 → passed ─┘
PREDICTION — the answer is hidden
5 hrs, 70 out of 100 → Trained model → "passed"
↑ (not very sure)
a student the model has never seenThe loop inside the box is always the same. The model guesses, the guess is compared against the label, and the model is nudged toward being less wrong. Repeat until it stops improving.
Where you have already seen it
- Gmail's spam folder. Trained on mail that humans marked as spam.
- Your bank blocking a card. Trained on past transactions known to be fraud.
- Insurance premium quotes. Trained on past claims and what they cost.
- Weather apps predicting rain. Trained on decades of recorded weather.
- Delivery apps predicting arrival time. Trained on millions of completed trips.
The honest catch
Supervised learning needs labels, and labels are expensive.
Say you want a model that spots lung disease in a scan. Someone must first look at thousands of scans and mark each one. That someone is a doctor. Their time costs real money, and they disagree with each other more often than you would expect.
This is the reason so much of machine learning research is about needing fewer labels. Labels, not algorithms, are usually the bottleneck.
There is a second catch. The model learns whatever is in the labels, including the unfairness. If past hiring decisions favoured one group, a model trained on them learns to favour that group. It has no idea it is doing something wrong. It is copying the answer key it was given.
Remember this
- Supervised learning needs examples where the correct answer is already known.
- Classification puts things in buckets. Regression predicts a number.
- The model can only be as good, and as fair, as the labels it was given.
What to learn next
- Linear regression — the simplest supervised model that predicts a number.
- Unsupervised learning — what to do when you have no labels at all.
- Model evaluation — checking whether the trained model is actually good.
Developer — Code and libraries.
Both tasks use the same three-line pattern in scikit-learn. Learn it once, apply it everywhere.
model = SomeEstimator() # choose the shape of the rule
model.fit(X, y) # learn from features X and labels y
model.predict(X_new) # answer new questionsX is always a 2D structure — rows are examples, columns are features. y is always 1D, one label per row.
Setup
pip install scikit-learnClassification — which bucket?
from sklearn.linear_model import LogisticRegression
# Inputs: [hours studied per day, attendance percent]
X = [[1, 45], [2, 50], [3, 60], [4, 62], [6, 78], [7, 85], [8, 90], [9, 95]]
# Labels: 1 means the student passed, 0 means the student failed
y = [0, 0, 0, 0, 1, 1, 1, 1]
model = LogisticRegression(max_iter=1000)
model.fit(X, y)
# Three students the model has never seen: strong, weak, and borderline
new_students = [[8, 88], [1, 30], [5, 70]]
print("predicted label:", model.predict(new_students))
print("chance of passing:", model.predict_proba(new_students)[:, 1].round(2))predicted label: [1 0 0] chance of passing: [1. 0. 0.46]
Read the third number carefully. The borderline student gets 0.46, which is close to a coin toss. The model is telling you it does not know. That confession is more valuable than the prediction.
Now read the first two numbers with suspicion. A confidence of 1.00 from eight rows of training data is not a triumph — it is a warning. Those values are rounded from something like 0.998. They are that extreme because the two groups in this tiny dataset never overlap. Real data overlaps, and honest models rarely report certainty.
predict()gives the label.predict_proba()gives the probability of each class.[:, 1]selects the second column, which is the probability of class1. Column0holds the probability of class0, and the two always add to one.max_iter=1000raises the optimiser's iteration budget. With the default of 100 this cleanly-separated data triggers aConvergenceWarning.
Regression — how much?
from sklearn.linear_model import LinearRegression
# Input: hours studied per day. Label: marks out of 100.
X = [[1], [2], [3], [4], [5], [6]]
y = [35, 42, 52, 57, 68, 74]
model = LinearRegression().fit(X, y)
print("marks predicted for 4.5 hours:", round(model.predict([[4.5]])[0], 1))
print("marks gained per extra hour:", round(model.coef_[0], 2))marks predicted for 4.5 hours: 62.6 marks gained per extra hour: 7.94
Same three lines. Different estimator, different question. The output is now a number on a continuous scale, not a bucket.
coef_ holds the learned relationship — roughly 7.94 extra marks for each extra study hour. Note carefully that this describes a pattern in the data, not a promise about a real student.
Common mistakes
Shapes are wrong. X must be 2D, y must be 1D. Passing X = [1, 2, 3] raises ValueError: Expected 2D array, got 1D array instead. Reshape with [[1], [2], [3]] or np.array(X).reshape(-1, 1).
Using a classifier for a numeric target. LogisticRegression on marks out of 100 treats every distinct mark as a separate class. It loses the fact that 61 is nearer to 62 than to 20. It will "work" and be nonsense. Ask which shape of question you have first.
Scoring on the training data. Call model.score(X, y) on the same rows used for fit() and you measure memorisation. It says nothing about tomorrow. Split the data first — see train, test and validation splits.
Leaking the answer into the features. Suppose a column was recorded after the outcome was already known. The model leans on it and scores near-perfectly in testing. Then it collapses in production, because that column does not exist at prediction time. This one is genuinely hard to spot, and it has embarrassed experienced teams.
Try it yourself
In pass_or_fail.py, add one student who studied 8 hours, attended 90 percent, and still failed. Append [8, 90] to X, and 0 to y.
Rerun it. The confidence on the strong student will drop below 1.00. One contradictory example is enough to teach the model doubt — which is exactly what you want it to learn.
What to learn next
- Linear regression — open up what
LinearRegressionactually did. - Train, test and validation splits — how to score a model honestly.
- Model evaluation — why accuracy alone is a poor measure.
Researcher — Mathematics and papers.
Formal statement
Supervised learning assumes a sample drawn independently from a fixed joint distribution.
Given S = {(x_1, y_1), ..., (x_n, y_n)} drawn i.i.d. from D over X x Y
Find h : X -> Y minimising R(h) = E_{(x,y) ~ D} [ L(h(x), y) ]x_i— feature vector for examplei, typically inR^dy_i— label;{1..K}for K-class classification,Rfor regressionD— the unknown joint distribution over features and labelsL— a loss function scoring the mismatch between prediction and labelR(h)— the true risk, the quantity we want small
The i.i.d. assumption is doing enormous work here. It fails under distribution shift, time-series dependence, and any sampling process correlated with the label. Most production failures of supervised learning are failures of this assumption, not of the optimiser.
Loss functions
The choice of L defines what "correct" means, and different losses induce genuinely different optimal predictors.
0-1 loss L(y_hat, y) = 1 if y_hat != y else 0
Squared error L(y_hat, y) = (y_hat - y)^2
Absolute error L(y_hat, y) = |y_hat - y|
Binary cross-ent. L(p, y) = -( y*log(p) + (1-y)*log(1-p) )
Hinge (SVM) L(f, y) = max(0, 1 - y*f), y in {-1, +1}y_hat— the model's predictiony— the true labelp— predicted probability of the positive class, in the open interval (0, 1)f— the raw real-valued model score before any threshold
The minimiser of expected squared error is the conditional mean E[y | x]. The minimiser of expected absolute error is the conditional median. This is why absolute error is more robust to outliers. It is not a heuristic preference — it is a different estimand.
0-1 loss is discontinuous and non-convex, so it is not directly optimised. Cross-entropy and hinge loss are convex surrogates, meaning tractable stand-ins. Bartlett, Jordan & McAuliffe (2006) established the classification-calibration conditions under which minimising a surrogate also minimises 0-1 risk.
The Bayes optimal predictor
The best achievable predictor under 0-1 loss is:
h*(x) = argmax_{k} P(y = k | x)Its risk R(h*) is the Bayes error, the irreducible error arising from genuine overlap between classes. No model, no amount of data, and no architecture can go below it.
This matters when reading benchmark claims. Suppose a paper reports 92 percent accuracy. That is excellent against a Bayes error of 7 percent, and mediocre against 1 percent. The Bayes error is unknown for real datasets. That is precisely why human-level performance is often used as a rough proxy.
Consistency and sample complexity
An estimator is consistent if R(h_hat_n) -> R(h*) in probability as n -> infinity. Consistency is an asymptotic guarantee, and asymptotic guarantees say nothing about the sample you actually have.
For a finite hypothesis class H in the realisable case (some h in H has zero risk), PAC theory gives:
n >= ( log|H| + log(1/delta) ) / epsilonn— number of labelled examples required|H|— size of the hypothesis classepsilon— target excess riskdelta— allowed failure probability
In the agnostic case the dependence worsens to 1/epsilon^2. Replacing log|H| with the VC dimension extends this to infinite classes.
Label noise
Real labels are wrong some of the time. Under symmetric label noise with flip rate eta < 0.5, Natarajan et al. (2013) showed that loss correction can recover an unbiased risk estimate. Under asymmetric or instance-dependent noise, guarantees are considerably weaker.
Empirically, deep networks fit clean patterns before they memorise noisy labels (Arpit et al., 2017). That ordering is why early stopping is a surprisingly effective noise defence.
Reducing the label bottleneck
Labels dominate cost, so several families of methods attack that directly.
| Approach | Idea | Representative work |
|---|---|---|
| Semi-supervised | Use unlabelled data alongside a small labelled set | Chapelle et al. (2006); FixMatch, Sohn et al. (2020) |
| Active learning | Choose which points to label next | Settles (2009) survey |
| Weak supervision | Combine noisy labelling heuristics | Snorkel, Ratner et al. (2017) |
| Transfer learning | Pretrain on a large task, fine-tune on a small one | Yosinski et al. (2014) |
| Self-supervised | Invent labels from the data itself | SimCLR, Chen et al. (2020) |
Transfer and self-supervised pretraining are the dominant practical answers today. They largely explain why the field moved from training per-task models to fine-tuning foundation models.
Key references
- Vapnik, V. (1998). Statistical Learning Theory. Wiley.
- Hastie, T., Tibshirani, R. & Friedman, J. (2009). The Elements of Statistical Learning, 2nd ed. Springer. Free PDF from the authors.
- Bartlett, P., Jordan, M. & McAuliffe, J. (2006). Convexity, Classification, and Risk Bounds. JASA 101(473).
- Natarajan, N. et al. (2013). Learning with Noisy Labels. NeurIPS.
- Arpit, D. et al. (2017). A Closer Look at Memorization in Deep Networks. ICML.
What to learn next
- Model evaluation — estimating
R(h)and quantifying the estimate's variance. - Overfitting and underfitting — decomposing the gap between empirical and true risk.
- Linear regression — the closed-form case where all of this is exactly solvable.