Loss functions
A loss function is the marking scheme a model is trained against, and picking the wrong one quietly teaches the model to be good at the wrong thing.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A loss function is the rule that decides how wrong an answer was.
Think about how your school papers were marked. Imagine a teacher who cut ten marks for untidy handwriting and one mark for a wrong answer. Within a month the whole class would have beautiful handwriting and still not understand the subject.
The marking scheme did not describe what you learned. It decided what you learned.
A loss function is the marking scheme for a model. It is one number, and a bigger number means a worse answer. Training is the process of pushing that number down.
Why it exists
A model cannot improve from the word "wrong". It needs a size.
Being off by two rupees and being off by two lakh rupees are both "wrong". One of them should worry you far more. The loss function turns that difference into a number the machine can act on.
The training loop then does the rest. It measures the loss, works out which knob to turn, and turns it a little.
So the loss sits at the centre of everything. Every weight in the model moves because of it.
The two big families
Almost every loss you meet belongs to one of two groups.
"How far off were you?" — used when the answer is a number. Predicting rent, temperature, delivery time, a person's age. You compare your number to the real number and measure the gap.
"How surprised should you be?" — used when the answer is a choice from a list. Spam or not spam. Which of ten digits. Which of a thousand words comes next. Here you compare a confidence against the truth, and confident wrong answers are punished hardest.
That second family has a name you will see everywhere: cross-entropy. It grows sharply when the model was sure and wrong.
How it works
input → model → guess: 8.2
|
real answer: 9.0 |
\ /
\ /
→ loss function ← gap of 0.8 → loss = 0.64
|
↓
"turn the knobs this way"The loss is computed for a whole batch of examples at once and averaged. That average is the single number the model is trying to shrink.
The part that catches everyone
Two rulers can disagree about which model is better.
Say a rent model is off by 2 thousand on four flats, and off by 40 thousand on one unusual bungalow. Measure the gap directly and the four small misses matter. Square the gap first and the single bungalow drowns out everything else.
Squaring is the common default. It means the model will bend itself out of shape to fix one weird example. The ordinary ones get slightly worse in exchange.
Neither ruler is correct. It depends on whether that bungalow is a real case you must handle, or a typing error in your spreadsheet. This is a decision you make, not one the maths makes for you.
Where you have already seen this
- A food delivery app showing "arriving in 32 minutes". Being 5 minutes late is annoying, being 45 minutes late loses the customer. The marking scheme has to reflect that.
- Your keyboard suggesting the next word. It is scored on how surprised it was by the word you actually typed.
- A bank flagging a payment as fraud. Missing a real fraud costs far more than a false alarm, and the loss is weighted to say so.
What is honestly hard here
The loss is not the thing you care about. It is a stand-in for the thing you care about.
You care about "does the doctor trust this scan report". You cannot write that as a number, so you train against something you can compute instead. The gap between those two is where most real failures live.
Read that twice. It confuses almost everyone at first, and it never fully goes away. A model with a lovely low loss can still be useless, and the loss will not tell you.
Remember this
- The loss function is the marking scheme: one number, lower is better.
- Numbers get "how far off" losses. Choices get "how surprised" losses.
- Choosing the loss is a design decision, and it decides what your model becomes good at.
What to learn next
- Gradient descent — what the model does with the number the loss gives it.
- PyTorch basics — the built-in loss classes, in a real training loop.
- Model evaluation — the metrics you report, as opposed to the loss you train on.
Developer — Code and libraries.
Setup
pip install numpyNumPy alone covers everything important. One later example uses PyTorch to show a mistake that only appears in a framework, and it is marked so you can skip it.
Three rulers, one set of errors
Same predictions, same truth, three different loss functions. Watch what a single outlier does to each.
import numpy as np
# Five houses. What we predicted, and what the price really was (in lakhs).
pred = np.array([50.0, 62.0, 45.0, 70.0, 55.0])
truth = np.array([52.0, 60.0, 47.0, 68.0, 95.0]) # the last one is a mansion we missed badly
error = pred - truth
def mse(e): # squares the miss, so a big miss dominates
return np.mean(e ** 2)
def mae(e): # plain distance, every rupee counts the same
return np.mean(np.abs(e))
def huber(e, delta=5.0): # squared when close, straight-line when far
small = np.abs(e) <= delta
return np.mean(np.where(small, 0.5 * e ** 2, delta * (np.abs(e) - 0.5 * delta)))
print("errors: ", error)
print("MSE :", round(mse(error), 2))
print("MAE :", round(mae(error), 2))
print("Huber:", round(huber(error), 2))
# Now delete the mansion and score the same four remaining houses.
e4 = error[:4]
print()
print("without the outlier")
print("MSE :", round(mse(e4), 2))
print("MAE :", round(mae(e4), 2))
print("Huber:", round(huber(e4), 2))errors: [ -2. 2. -2. 2. -40.] MSE : 323.2 MAE : 9.6 Huber: 39.1 without the outlier MSE : 4.0 MAE : 2.0 Huber: 2.0
Read those numbers carefully
One bad example multiplied MSE by about 80. It multiplied MAE by under 5.
MSE is mean squared error: average of the squared gaps. Because it squares, an error of 40 counts 400 times as much as an error of 2. During training, that means almost the entire gradient comes from your worst examples.
MAE is mean absolute error: average of the plain gaps. Every rupee of error counts the same, so a stray outlier cannot hijack training. The cost is that its gradient has the same size no matter how close you are, which makes the final approach to a good answer jittery.
Huber is the compromise. Inside delta it behaves like MSE, so the gradient shrinks as you get close. Outside delta it behaves like MAE, so one wild example cannot dominate. delta is yours to set, and it should be roughly the error size above which you stop believing the data point.
Why squared error is the wrong ruler for yes/no answers
Now the second family. The model outputs a confidence between 0 and 1, and the truth is 1.
import numpy as np
def bce(p, y):
p = np.clip(p, 1e-12, 1 - 1e-12) # log(0) is -inf and poisons everything after it
return -(y * np.log(p) + (1 - y) * np.log(1 - p))
truth = 1.0 # the real answer is "yes, this is spam"
for guess in [0.99, 0.90, 0.60, 0.50, 0.10, 0.01]:
squared = (guess - truth) ** 2
print(f"model said {guess:.2f} -> squared error {squared:.4f} cross-entropy {bce(guess, truth):.4f}")model said 0.99 -> squared error 0.0001 cross-entropy 0.0101 model said 0.90 -> squared error 0.0100 cross-entropy 0.1054 model said 0.60 -> squared error 0.1600 cross-entropy 0.5108 model said 0.50 -> squared error 0.2500 cross-entropy 0.6931 model said 0.10 -> squared error 0.8100 cross-entropy 2.3026 model said 0.01 -> squared error 0.9801 cross-entropy 4.6052
Squared error can never punish a wrong answer by more than 1, no matter how confident the model was. Cross-entropy has no ceiling. Being 99% sure and wrong costs about 4.6; being 50/50 costs about 0.69.
That unbounded punishment is the point. A classifier that is confidently wrong is worse than one that is unsure, and the loss has to say so.
There is a second reason, and it is about gradients. Squared error on top of a sigmoid produces a gradient that shrinks to nothing exactly when the model is most confidently wrong, so training stalls. Cross-entropy cancels that term out. The researcher tab on this page shows the cancellation in full.
The mistake every PyTorch beginner makes
This one needs pip install torch. Skip it if you would rather not download a large package today — the lesson survives without running it.
nn.CrossEntropyLoss and F.cross_entropy apply softmax for you, internally. Applying it yourself first is silent: no error, no warning, a model that trains badly.
import torch
import torch.nn.functional as F
logits = torch.tensor([[2.0, 0.5, -1.0]]) # raw scores for 3 classes, straight out of the last layer
target = torch.tensor([0]) # the correct class is class 0
right = F.cross_entropy(logits, target)
probs = F.softmax(logits, dim=1) # the mistake: softmax applied before the loss
wrong = F.cross_entropy(probs, target)
print("probabilities:", [round(p, 4) for p in probs[0].tolist()])
print("correct (logits in):", round(right.item(), 4))
print("wrong (softmax in) :", round(wrong.item(), 4))probabilities: [0.7856, 0.1753, 0.0391] correct (logits in): 0.2413 wrong (softmax in) : 0.7017
Both numbers look like a plausible loss. Only one of them is the loss you meant. The double-softmax squashes the scores towards each other, gradients get tiny, and your model learns at a crawl for no visible reason.
Rule to memorise: pass raw scores (called logits, the unsquashed numbers from the final layer) into CrossEntropyLoss and BCEWithLogitsLoss. Apply softmax or sigmoid only when you want to show a probability to a human.
A short menu
| Your output is | Use | Notes |
|---|---|---|
| One number | MSE (nn.MSELoss) | Default. Outlier-sensitive. |
| One number, messy data | Huber (nn.SmoothL1Loss) | Set beta near your noise level. |
| Yes / no | nn.BCEWithLogitsLoss | Feed logits. Use pos_weight for imbalance. |
| One of N classes | nn.CrossEntropyLoss | Feed logits. Targets are class indices, not one-hot. |
| Several labels at once | nn.BCEWithLogitsLoss | One independent yes/no per label. |
| Rank / similarity | Triplet, InfoNCE | Compares pairs, not single predictions. |
Common mistakes
Applying softmax before CrossEntropyLoss. Covered above. There is no error message, which is what makes it expensive.
One-hot targets for CrossEntropyLoss. It wants class indices such as tensor([0, 2, 1]), not tensor([[1,0,0],[0,0,1],[0,1,0]]). Modern PyTorch accepts float probability targets too, which makes the shape error even less likely to be caught.
Treating accuracy as a loss. Accuracy is a count, so its gradient is zero almost everywhere. Nudging a weight a little never changes the count, so there is nothing to descend. Train on cross-entropy, then report accuracy.
Ignoring class imbalance. With 99% negatives, a model that answers "no" every time reaches a low loss and is worthless. Use pos_weight, class weights, or resampling, and evaluate with precision and recall from model evaluation.
Comparing loss values across different losses. An MSE of 0.4 and a cross-entropy of 0.4 have no relationship. Loss values are comparable only against themselves, on the same data, with the same loss.
Try it yourself
Take rulers.py and change the mansion's true price from 95 to 200. Predict which of the three numbers moves most before you run it.
Then set delta=100 in huber and watch it become MSE. Set delta=0.5 and watch it become MAE. That single knob spans both families.
What to learn next
- Gradient descent — what the model does with the number the loss gives it.
- PyTorch basics — the built-in loss classes, in a real training loop.
- Model evaluation — the metrics you report, as opposed to the loss you train on.
Researcher — Mathematics and papers.
The regression losses
Written for a batch of n examples, with prediction p_i and target y_i, and e_i = p_i - y_i.
MSE = (1/n) Σ e_i² gradient wrt p_i: 2 e_i / n
MAE = (1/n) Σ |e_i| gradient wrt p_i: sign(e_i) / n
Huber = (1/n) Σ 0.5 e_i² if |e_i| ≤ δ
δ(|e_i| − 0.5 δ) otherwisen is the batch size. δ is the crossover threshold, in the same units as the target. sign(e) is +1, -1 or 0.
The estimator each loss implies is the useful way to remember them. Minimising MSE over a constant predictor returns the mean of the targets. Minimising MAE returns the median. Minimising the pinball loss at quantile τ returns the τ-quantile, which is what quantile regression uses to produce prediction intervals.
Huber's gradient is bounded by δ/n, which is what makes it robust: a single corrupted target can move the parameters by a bounded amount, whereas under MSE its influence is unbounded. Huber (1964), Robust estimation of a location parameter, is the original.
Losses are negative log-likelihoods
Almost every standard loss is −log p(y | x, θ) under some assumed noise model. This is not a coincidence and it is the fastest way to derive a loss for a new problem.
Gaussian noise, fixed variance → −log p ∝ (p − y)² → MSE
Laplace noise → −log p ∝ |p − y| → MAE
Bernoulli outcome → −log p = BCE
Categorical outcome → −log p = cross-entropy
Poisson counts → −log p = p − y log p → Poisson NLLThe practical payoff: if your targets are counts, MSE assumes Gaussian noise on counts, which is wrong at low rates. Use the Poisson loss and the model becomes better calibrated at the small values that dominate count data.
Cross-entropy and the softmax gradient
For K classes with logits z ∈ R^K and a one-hot target y:
softmax: q_k = exp(z_k) / Σ_j exp(z_j)
loss: L = − Σ_k y_k log q_k
gradient: ∂L/∂z_k = q_k − y_kz_k is the raw score for class k, q_k the predicted probability, y_k the one-hot target. The gradient is the entire reason this pairing is standard: it is prediction − truth, one subtraction, with no derivative of the squashing function left in it.
Contrast with MSE on a sigmoid output. There the gradient carries a factor σ'(z) = σ(z)(1 − σ(z)), which approaches zero as |z| grows. A confidently wrong prediction therefore produces a vanishing gradient and the model stalls exactly where correction is most needed.
Numerically, never compute log(softmax(z)) directly. Use the log-sum-exp identity with the max subtracted:
log q_k = z_k − m − log Σ_j exp(z_j − m), m = max_j z_jm cancels algebraically and bounds every exponent at zero, so nothing overflows. F.cross_entropy and BCEWithLogitsLoss implement this fused form; a hand-written softmax followed by a log does not.
Cross-entropy, entropy and KL
H(y, q) = H(y) + D_KL(y ‖ q)H(y) is the entropy of the target distribution and D_KL the Kullback–Leibler divergence. With hard one-hot targets H(y) = 0, so minimising cross-entropy is exactly minimising KL divergence to the label. With soft targets — label smoothing, or a teacher model in knowledge distillation — H(y) is a non-zero constant with respect to θ, so the two objectives still share a gradient but no longer share a value.
Imbalance and hard-example weighting
Class weighting multiplies each class's term by w_c, typically w_c ∝ 1/n_c or its square root. This changes the implied prior; the model's output probabilities are no longer calibrated to the training distribution, and thresholds must be re-tuned on a validation set.
Focal loss (Lin et al., 2017) down-weights examples the model already gets right:
FL = − α_t (1 − q_t)^γ log q_tq_t is the predicted probability of the true class, γ ≥ 0 the focusing parameter (2.0 is the paper's default), α_t an optional per-class weight. At γ = 0 it reduces to weighted cross-entropy. Designed for dense object detection where background boxes outnumber objects by roughly 1000:1.
Label smoothing (Szegedy et al., 2016) replaces the one-hot target with (1 − ε) y + ε/K. It reliably improves top-1 accuracy and reduces overconfidence. Müller et al. (2019), When does label smoothing help?, show it also collapses within-class logit spread, which measurably degrades knowledge distillation from a smoothed teacher.
Proper scoring rules and calibration
A scoring rule is proper when the expected score is optimised by reporting the true probability. Log loss and Brier score are proper; accuracy and F1 are not, which is why optimising them directly encourages degenerate thresholding rather than honest probability estimates.
Minimising a proper scoring rule does not by itself deliver calibration in practice. Guo et al. (2017), On calibration of modern neural networks, document that deep networks trained to convergence on cross-entropy are systematically overconfident, and that temperature scaling — a single scalar divisor on the logits, fit on a validation set — removes most of the miscalibration at no cost to accuracy.
The surrogate gap
Training minimises a differentiable surrogate; deployment cares about a non-differentiable objective. The gap is a modelling assumption, not an implementation detail, and it is where most production disappointment originates.
Three concrete mismatches worth naming:
- Metric mismatch. Cross-entropy optimises average log-likelihood; the product cares about recall at a fixed false-positive rate. These have different optima.
- Distribution mismatch. The loss is averaged over the training distribution, and deployment draws from another one.
- Cost mismatch. Errors have asymmetric real-world costs that a symmetric loss never encodes. Encoding them explicitly, through class weights or a cost matrix, is nearly always better than discovering them in production.
Papers
- Huber, Robust Estimation of a Location Parameter, Annals of Mathematical Statistics, 1964.
- Szegedy et al., Rethinking the Inception Architecture for Computer Vision, 2015 — label smoothing — arxiv.org/abs/1512.00567
- Guo et al., On Calibration of Modern Neural Networks, 2017 — arxiv.org/abs/1706.04599
- Lin et al., Focal Loss for Dense Object Detection, 2017 — arxiv.org/abs/1708.02002
- Müller, Kornblith and Hinton, When Does Label Smoothing Help?, 2019 — arxiv.org/abs/1906.02629
- Gneiting and Raftery, Strictly Proper Scoring Rules, Prediction, and Estimation, JASA, 2007.
What to learn next
- Gradient descent — what the model does with the number the loss gives it.
- PyTorch basics — the built-in loss classes, in a real training loop.
- Model evaluation — the metrics you report, as opposed to the loss you train on.