Loss function
In one sentence A loss function scores how wrong a prediction is, giving training one number to push down.
Updated
A loss function is the formula that scores how wrong a prediction was, turning the model's mistake into one number that training can push down.
A teacher marking papers does not say "not bad". They deduct specific marks, and the total tells the student how far they are from full marks. The loss is that total. Everything the model learns comes from trying to make this one number smaller, which means the loss function is where you state what "good" actually means for your problem.
Choosing it is not a formality. It is the most direct way you tell the model what to care about.
The ones you will meet
| Task | Loss | What it punishes |
|---|---|---|
| Predicting a number | Mean squared error | Large misses very heavily |
| Predicting a number, with outliers | Mean absolute error / Huber | Misses evenly, ignores extremes less |
| Picking one of several classes | Cross-entropy | Confident wrong answers hardest |
| Yes or no | Binary cross-entropy | Same, for two outcomes |
Cross-entropy behaves in a way worth feeling. If the true answer is "cat" and the model said cat with probability 0.9, the loss is small. At 0.1, the loss is large. At 0.01 it is larger still, and it climbs towards infinity as confidence in a wrong answer approaches certainty. Being confidently wrong is punished far more than being unsure.
true label: cat
model says cat 0.90 → loss 0.11
model says cat 0.50 → loss 0.69
model says cat 0.10 → loss 2.30
model says cat 0.01 → loss 4.61One practical rule saves a lot of pain in PyTorch: pass raw scores, not probabilities. nn.CrossEntropyLoss applies softmax internally and nn.BCEWithLogitsLoss applies sigmoid internally, both in a numerically stable way. Applying softmax yourself first is a common cause of a loss that turns into NaN.
Where to go next
- Full lesson: Loss functions
- Related terms: gradient-descent, backpropagation, activation-function, overfitting