When validation loss rises but accuracy improves
Loss punishes confidence and accuracy does not, so a model growing more sure of itself can get more answers right while its loss curve turns upward — and knowing which curve to obey is a decision you make before training.
- 13 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Loss reads confidence, accuracy reads only right-or-wrong — so a model growing more sure of itself can get more answers right while its loss gets worse.
Picture the weather presenter on the evening news. On Monday she says "60% chance of rain" and it stays dry. Nobody minds. On Tuesday she says "certain rain, carry an umbrella" and it stays dry. Now people are annoyed.
Both days she was wrong once. The second wrongness cost far more, because of the confidence attached to it. That extra cost is exactly what loss counts and accuracy ignores.
So when your two curves disagree, neither one is broken. They are answering different questions, and you have to decide which question you are being paid to answer.
Why this had to be invented
Accuracy is a counting measure: right or wrong, one mark each. It cannot tell a barely-correct guess from a rock-solid one.
That blindness is a real problem. A screening model that says "70% chance of disease" beats one that says only "disease". The number changes what the doctor does next. So training uses a score that reads the confidence, not only the verdict. That score is the loss.
Loss rewards confident-and-right and punishes confident-and-wrong, harshly. As training continues, models tend to get more confident about everything. Most of that confidence lands on answers they get right, which helps. A little lands on answers they get wrong, and each one of those is expensive.
When the expensive misses outweigh the cheap gains, loss turns upward — while the count of correct answers is still climbing.
How it works
model early in training model later in training
(cautious) (confident)
right -> "probably yes" (0.60) "definitely yes" (0.99)
small cost tiny cost <- loss improves
wrong -> "probably yes" (0.60) "definitely yes" (0.97)
small cost HUGE cost <- loss worsens
ACCURACY sees: one right, one wrong -> unchanged, or better
LOSS sees: one tiny saving, one enormous penalty -> worseOne confidently wrong answer can outweigh twenty confidently right ones. That imbalance is the whole phenomenon.
A real example you have seen
Your phone's face unlock is the same trade in miniature. Waving it at your face and having it open is the accuracy question. Whether it opens instantly or hesitates for a beat is the confidence question.
A model can get more decisive and still let in the occasional stranger with total conviction. The count of successful unlocks looks fine. The one confident mistake is the one that matters.
Remember this
- Accuracy counts verdicts. Loss reads confidence. They can move in opposite directions honestly.
- Rising loss with rising accuracy usually means the model is becoming overconfident, not worse.
- Decide before training which curve your stopping rule obeys, and write it down.
What to learn next
- Platt scaling — the standard repair for a model that is right but too sure.
- Reliability diagrams and calibration error — how to see overconfidence directly instead of inferring it from curves.
- When your validation curve is too noisy to trust — before you act on a 0.003 accuracy improvement.
Developer — Code and libraries.
Setup
pip install torch numpyVerified with torch 2.5.1 (CPU), numpy 1.26.4, Python 3.10. Both scripts finish in under twenty seconds, generate their data in memory, and download nothing.
First, the mechanism in six rows
Before the training loop, look at the arithmetic on a validation set of six examples. Two snapshots of the same model: one cautious, one confident.
import numpy as np
truth = np.array([1, 1, 1, 1, 0, 0])
early = np.array([0.60, 0.60, 0.60, 0.45, 0.40, 0.60]) # cautious snapshot
late = np.array([0.99, 0.99, 0.99, 0.55, 0.40, 0.97]) # confident snapshot
def accuracy(p): return ((p > 0.5).astype(int) == truth).mean()
def log_loss(p): return -np.mean(truth*np.log(p) + (1-truth)*np.log(1-p))
print("snapshot accuracy log loss")
for name, p in [("cautious", early), ("confident", late)]:
print(f"{name:11s} {accuracy(p):7.3f} {log_loss(p):8.3f}")
print("\nwhere the loss went")
for t, a, b in zip(truth, early, late):
la = -(t*np.log(a) + (1-t)*np.log(1-a))
lb = -(t*np.log(b) + (1-t)*np.log(1-b))
mark = " <-- confidently wrong" if lb > 2 else ""
print(f" truth {t} p {a:.2f} -> {b:.2f} loss {la:.3f} -> {lb:.3f}{mark}")snapshot accuracy log loss cautious 0.667 0.626 confident 0.833 0.774 where the loss went truth 1 p 0.60 -> 0.99 loss 0.511 -> 0.010 truth 1 p 0.60 -> 0.99 loss 0.511 -> 0.010 truth 1 p 0.60 -> 0.99 loss 0.511 -> 0.010 truth 1 p 0.45 -> 0.55 loss 0.799 -> 0.598 truth 0 p 0.40 -> 0.40 loss 0.511 -> 0.511 truth 0 p 0.60 -> 0.97 loss 0.916 -> 3.507 <-- confidently wrong
Accuracy rose from 0.667 to 0.833. Loss rose from 0.626 to 0.774. Both readings are correct. Four of the six rows improved their loss and a fifth did not move, while the sixth alone cost 2.59 — more than all the savings combined.
Now watch it happen during real training
Same synthetic task as how to read a loss curve, tracked at higher resolution around the turn.
import torch, torch.nn as nn
def make_data(n, seed):
g = torch.Generator().manual_seed(seed)
X = torch.randn(n, 8, generator=g)
y = (X[:, 0] * X[:, 1] + 0.4 * X[:, 2] + 0.3 * torch.randn(n, generator=g) > 0).float()
return X, y[:, None]
Xtr, ytr = make_data(600, 0)
Xva, yva = make_data(2000, 999)
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(8, 64), nn.ReLU(), nn.Linear(64, 64), nn.ReLU(), nn.Linear(64, 1))
opt, loss_fn = torch.optim.Adam(net.parameters(), lr=0.003), nn.BCEWithLogitsLoss()
print("epoch val_loss val_acc mean|logit|")
for e in range(111):
net.train(); opt.zero_grad(); loss_fn(net(Xtr), ytr).backward(); opt.step()
net.eval()
with torch.no_grad():
logits = net(Xva)
loss = loss_fn(logits, yva).item()
acc = ((logits > 0).float() == yva).float().mean().item()
conf = logits.abs().mean().item() # how far from the fence the model stands
if e >= 40 and e % 10 == 0:
print(f"{e:5d} {loss:8.4f} {acc:7.4f} {conf:11.2f}")epoch val_loss val_acc mean|logit| 40 0.4023 0.8170 1.54 50 0.3537 0.8375 2.41 60 0.3488 0.8440 3.24 70 0.3606 0.8475 3.85 80 0.3771 0.8415 4.26 90 0.3958 0.8435 4.63 100 0.4232 0.8370 5.05 110 0.4591 0.8370 5.54
Between epoch 60 and epoch 70, validation loss rises (0.3488 to 0.3606) and validation accuracy also rises (0.8440 to 0.8475). Stopping on loss picks epoch 60. Stopping on accuracy picks epoch 70. Both are defensible, and the two checkpoints are different models.
The walkthrough
mean|logit| is the confidence meter, and it never stops climbing. A logit is the model's raw pre-probability output; a logit of 0 means "no idea", and a large magnitude means "very sure". It grows 1.54 → 5.54 while the model's actual skill flattens after epoch 70. This single column explains the whole divergence and costs one line to log — add it to every classifier you train.
Loss turns before accuracy does. Loss bottoms near epoch 57 in this run, accuracy peaks near epoch 70. Loss is the more sensitive instrument, because it notices confidence drift while accuracy is still rounding everything to right-or-wrong.
Accuracy is noisier than it looks. On 2,000 validation rows, a 0.0035 accuracy change is 7 examples. The rise from 0.8440 to 0.8475 is inside the noise band — read it with the error bars from noisy validation curves before you build a decision on it.
Nothing here is a bug. No leakage, no broken metric, no wrong label. This is what a correctly-implemented classifier does when trained past the point where it has learnt the pattern. Why your loss is not going down covers the case where something is actually broken.
What to do about it
Three responses, in the order most teams should try them.
Pick the curve that matches the decision, in advance. If a human reads the probability — a doctor, a credit officer, a routing rule with a threshold — stop on loss, because the probability itself is the product. If a fixed rule consumes a hard yes/no, stop on the metric that rule uses. Choose once, before seeing results, and record it with the experiment config.
Save both checkpoints. Tracking best-by-loss and best-by-accuracy separately costs a few megabytes and removes the argument entirely — see early stopping and best model.
Fix the confidence instead of the training. Overconfidence is repairable after the fact. Platt scaling and isotonic regression fit a small correction on validation data. That often recovers the early checkpoint's loss while keeping the late one's accuracy. This is usually the best of the three.
Common mistakes
Calling it overfitting and stopping there. Classic overfitting drags accuracy down with loss. When accuracy is rising, "overfitting" is the wrong word and it points at the wrong repair — a smaller model does not fix miscalibration.
Switching the stopping metric after seeing the curves. Picking whichever metric flatters the run is test-set overfitting wearing a different hat. Decide first.
Averaging loss over an epoch while measuring accuracy at the end. Then the two curves are measured at different moments and disagree for a boring bookkeeping reason. Score both from the same forward pass, as the script above does.
Reporting only accuracy for a probability product. If downstream code thresholds or ranks the probability, accuracy is not the metric of record. Choosing a threshold from costs is the conversation to have instead.
Try it yourself
Add a temperature knob: divide the validation logits by T before computing the loss, and search T over [1.0, 1.5, 2.0, 3.0] at epoch 110. Find the T that makes epoch 110's loss match epoch 60's. You have implemented temperature scaling by hand — and you will notice accuracy never moves, because dividing by a positive number cannot change the sign of a logit.
What to learn next
- Platt scaling — the standard repair for a model that is right but too sure.
- Reliability diagrams and calibration error — how to see overconfidence directly instead of inferring it from curves.
- When your validation curve is too noisy to trust — before you act on a 0.003 accuracy improvement.
Researcher — Mathematics and papers.
Proper scoring rules and the decomposition that explains the split
Log loss is a strictly proper scoring rule: its expectation is uniquely minimised by reporting the true conditional probability $P(y=1 \mid x)$. Accuracy is not proper — it is minimised by any predictor that lands on the correct side of the decision boundary, so an entire family of miscalibrated models achieves the optimum. The two curves are therefore optimising over different equivalence classes of models, and there is no reason for their argmins to coincide.
The mechanism is visible in the Brier decomposition (Murphy, 1973), which splits a proper score into
$$ \text{score} = \text{uncertainty} - \text{resolution} + \text{reliability} $$
where uncertainty depends only on the base rate, resolution rewards separating examples into groups with different outcome frequencies, and reliability penalises predicted probabilities that do not match observed frequencies. Late-training overconfidence increases resolution slightly while increasing reliability error a great deal — score worsens, discrimination (and hence accuracy and AUC) does not. Proper scoring rules and reliability diagrams develop both halves.
Why modern networks drift toward overconfidence
Guo et al. (2017), On Calibration of Modern Neural Networks (ICML), is the reference result: compared with the shallower networks of the 1990s, contemporary architectures achieve better accuracy and substantially worse calibration, with expected calibration error growing with depth, width and reduced weight decay. Their observed training dynamic is exactly this lesson's table — NLL overfits while 0/1 error continues to improve, and the network is, in their framing, "overfitting to the loss without overfitting to the error".
The underlying pressure is structural. With separable or near-separable training data and an unbounded softmax, cross-entropy has no finite minimiser; gradient descent on logistic loss drives the weight norm to infinity in the max-margin direction (Soudry et al., 2018, The Implicit Bias of Gradient Descent on Separable Data, JMLR). Growing $|\theta|$ scales all logits, which is precisely the mean|logit| column climbing while the decision boundary barely moves. Confidence inflation is therefore the default asymptotic behaviour, not a pathology to be surprised by.
Their prescribed remedy, temperature scaling, learns a single $T > 0$ and reports $\sigma(z/T)$ where $z$ is the logit vector; because $T$ is positive it is a monotone transform, so argmax predictions and hence accuracy, AUC and every rank-based metric are provably unchanged, while NLL and ECE typically fall to near their achievable minimum. One parameter fitted on a held-out split recovers most of the loss gap that the stopping-rule argument was fighting over.
Related dissociations worth recognising
The loss/accuracy split is one member of a family, and misdiagnosis is common.
- Loss down, accuracy flat: confidence improving on already-correct examples. Harmless, and invisible to a 0/1 metric.
- Loss up, AUC up: ranking improving while calibration decays — the same phenomenon read through a rank metric, which temperature scaling also leaves untouched. See ROC versus precision-recall curves.
- Loss down, downstream metric down: the surrogate and the objective genuinely diverge, common under class imbalance where macro versus micro averaging changes the ranking of checkpoints outright.
- Both worsen after a gap: ordinary overfitting, and the only member of the family where a capacity or regularisation change is the right response.
Label smoothing (Szegedy et al., 2016; analysed in Müller, Kornblith and Hinton, 2019, When Does Label Smoothing Help?, NeurIPS) caps attainable confidence by training against targets $1-\varepsilon$ and $\varepsilon/(K-1)$ for $K$ classes, flattening the divergence at source — at the documented cost of collapsing within-class logit structure, which degrades knowledge distillation from the smoothed teacher. Focal loss, mixup, ensembling and deep ensembles all shift the trade-off; Ovadia et al. (2019), Can You Trust Your Model's Uncertainty? (NeurIPS), benchmark them under distribution shift and find that no method preserves calibration once the test distribution moves, which is the honest caveat on all of the above.
What to learn next
- Platt scaling — the standard repair for a model that is right but too sure.
- Reliability diagrams and calibration error — how to see overconfidence directly instead of inferring it from curves.
- When your validation curve is too noisy to trust — before you act on a 0.003 accuracy improvement.