How Models Are Actually Trained
Cross-entropy and perplexity
Perplexity turns a training loss number into a plain statement — how many equally likely options the model was choosing between at each step.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Perplexity says how confused the model was, measured as "how many options did it seem to be picking between".
The analogy you have already lived
You are watching a quiz show and trying to answer along. On an easy question you narrow it down to two options before the host finishes. On a hard one you are stuck between all four.
Someone could describe your whole evening with one number. On average, you were choosing between about three options per question. That number is your perplexity.
A perplexity of 1 means you knew every answer with certainty. A perplexity of 4 on a four-option quiz means you were guessing throughout.
Why it exists
Training reports a loss — a score of how wrong the model was, where lower is better. Loss is fine for the optimiser, but it is a hard number for a human to feel.
Nobody has an instinct for whether a loss of 2.9 is good. Everybody has an instinct for "the model was picking between about eighteen words".
Perplexity is the loss translated into that second sentence. It is the same information wearing friendlier clothes.
How it works
model's confidence loss perplexity
------------------ ---- ----------
sure, and right tiny near 1
leaning the right way small 2 to 5
no idea, spread evenly medium the vocabulary size
sure, and WRONG huge enormousTwo things are worth burning into memory.
A model that has learned nothing scores its whole vocabulary size. Say it knows fifty thousand words and spreads its guess evenly. Then it is choosing between fifty thousand options. That is the starting line.
Being confidently wrong is punished far harder than being unsure. A shrug costs a little. A wrong answer stated with total certainty costs a great deal. This is deliberate, and it is why models learn to hedge.
Where you have already seen this
- Research papers reporting "perplexity 3.2 on WikiText" to compare models.
- Model cards on Hugging Face listing perplexity per quantised version.
- Spam filters and old speech recognisers, which used the same measure.
What is honestly hard here
Perplexity numbers from two different papers are often not comparable, and people compare them anyway.
The number depends on how the text was cut into pieces. It also depends on which text was used, and on the window size. Change any of those and the number moves, with no change to the model.
Treat perplexity as a measurement of one model on one dataset with one setup. It is a thermometer, not a leaderboard.
Remember this
- Perplexity is "how many options the model seemed to be choosing between".
- A clueless model scores its whole vocabulary size. Lower is better.
- Numbers from different papers, datasets or tokenisers are not comparable.
What to learn next
- Packing documents into fixed-length batches — the batching that these losses are averaged over.
- Reading a loss curve — what the shape of the curve is telling you.
- Measuring what quantisation costs you — where perplexity misleads badly.
Developer — Code and libraries.
Setup
pip install torchThe second example adds pip install transformers and downloads GPT-2, about 550 MB.
The definition, made concrete
import math
import torch
import torch.nn.functional as F
V = 8 # a toy vocabulary of eight tokens
cases = {
"knows nothing (uniform)": torch.full((V,), 1.0 / V),
"leans the right way": torch.tensor([.40, .20, .10, .10, .08, .06, .04, .02]),
"confident and correct": torch.tensor([.93, .02, .02, .01, .01, .005, .003, .002]),
"confident and WRONG": torch.tensor([.02, .93, .02, .01, .01, .005, .003, .002]),
}
target = torch.tensor([0]) # token 0 is what actually came next
print(f"{'model':<24} {'p(correct)':>10} {'loss':>8} {'perplexity':>11}")
for name, probs in cases.items():
loss = F.nll_loss(probs.log().unsqueeze(0), target)
print(f"{name:<24} {probs[0]:>10.3f} {loss:>8.3f} {loss.exp():>11.2f}")
print(f"\nvocabulary size = {V}, so ln(V) = {math.log(V):.3f} and exp(ln(V)) = {V}")
# Perplexity is an average over MANY tokens, not one.
seq_probs = torch.tensor([0.9, 0.9, 0.9, 0.9, 0.001]) # four easy tokens, one shock
nll = -seq_probs.log()
print("\nper-token loss:", [f"{v:.2f}" for v in nll.tolist()])
print(f"mean loss = {nll.mean():.4f}")
print(f"perplexity = {nll.mean().exp():.2f}")
print(f"perplexity of the first four alone = {(-seq_probs[:4].log()).mean().exp():.2f}")model p(correct) loss perplexity knows nothing (uniform) 0.125 2.079 8.00 leans the right way 0.400 0.916 2.50 confident and correct 0.930 0.073 1.08 confident and WRONG 0.020 3.912 50.00 vocabulary size = 8, so ln(V) = 2.079 and exp(ln(V)) = 8 per-token loss: ['0.11', '0.11', '0.11', '0.11', '6.91'] mean loss = 1.4658 perplexity = 4.33 perplexity of the first four alone = 1.11
Reading that output
The uniform model scores exactly 8.00 on an 8-token vocabulary. That is the whole intuition in one line: perplexity is an effective number of choices.
Confidently wrong scores 50, on a vocabulary of 8. Perplexity is not bounded by the vocabulary size. A model can be far worse than clueless, and this asymmetry is the single most important property of the measure.
One bad token dragged perplexity from 1.11 to 4.33. Four tokens at 90% confidence, one at 0.1%, and the average is wrecked. Perplexity is a geometric mean over tokens, so rare disasters dominate it. When your evaluation perplexity jumps, look for a handful of catastrophic tokens before you suspect a general regression.
Real perplexity on a real model
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2").eval()
def token_losses(text):
ids = tok(text, return_tensors="pt").input_ids
with torch.no_grad():
logits = model(ids).logits
# position i predicts token i+1, so drop the last logit and the first label
losses = F.cross_entropy(logits[0, :-1], ids[0, 1:], reduction="none")
return ids[0], losses
for text in ["The capital of France is Paris",
"The capital of France is Kolkata",
"France of capital the is Paris"]:
ids, losses = token_losses(text)
ppl = losses.mean().exp().item()
print(f"\n{text!r} {len(ids)} tokens, perplexity {ppl:.1f}")
for i, loss in enumerate(losses.tolist()):
print(f" predict {tok.decode(ids[i + 1])!r:<10} loss {loss:5.2f}")'The capital of France is Paris' 6 tokens, perplexity 72.9
predict ' capital' loss 9.07
predict ' of' loss 1.81
predict ' France' loss 5.03
predict ' is' loss 2.11
predict ' Paris' loss 3.43
'The capital of France is Kolkata' 8 tokens, perplexity 73.8
predict ' capital' loss 9.07
predict ' of' loss 1.81
predict ' France' loss 5.03
predict ' is' loss 2.11
predict ' K' loss 8.68
predict 'olk' loss 3.41
predict 'ata' loss 0.00
'France of capital the is Paris' 6 tokens, perplexity 3193.3
predict ' of' loss 4.13
predict ' capital' loss 10.55
predict ' the' loss 6.10
predict ' is' loss 10.12
predict ' Paris' loss 9.45Written against transformers 5.6.2 and PyTorch 2.5.1, CPU, greedy scoring with no sampling — so these numbers reproduce exactly. Transformers may print a loss_type=None informational line for GPT-2; it is harmless.
This output contains four lessons
The model prefers the truth, and the evidence is per token. Paris costs 3.43, K (starting Kolkata) costs 8.68. That gap is the model's factual knowledge, visible directly.
Sentence-level perplexity hid that completely: 72.9 against 73.8. The false sentence scored almost identically, because four shared tokens dominated the average and the extra tokens diluted it. Never use whole-sentence perplexity to test whether a model knows a fact. Score the answer tokens only.
'ata' cost 0.00. After ' K' and 'olk', nothing else can follow. Multi-token words hand the model free, meaningless wins. A tokenizer that splits a word into five pieces makes your perplexity look better without the model improving at all.
The first token is always expensive. ' capital' cost 9.07 after only 'The'. Short texts are dominated by this cold-start penalty, which is why perplexity is reported over long documents.
Common mistakes
Forgetting the shift. Comparing logits[i] to ids[i] instead of ids[i+1] gives a suspiciously low number. Using model(ids, labels=ids) does the shift for you and is the safer route.
Averaging perplexities. Perplexity is exponentiated, so averaging two perplexities is wrong. Average the losses, then exponentiate once at the end.
Ignoring padding. With batches, padding tokens must be labelled -100 so cross_entropy skips them. Without that, you are measuring how well the model predicts padding, and perplexity looks great. See padding, truncation and attention masks.
Comparing across tokenisers. A model with a 32k vocabulary and one with a 256k vocabulary cut the same paragraph into different numbers of tokens. Their per-token perplexities are different units. Normalise by characters or bytes — bits per byte — if you must compare.
Evaluating on text the model trained on. Perplexity on training data is near-meaningless, and web-scale training corpora contain most public benchmark text. See benchmark contamination.
Try it yourself
Score "The capital of France is" followed by each of Paris, Lyon, Delhi, Bangalore. Print the loss of the answer tokens only, ignoring the shared prefix. That is how a factual-knowledge probe is actually built.
What to learn next
- Packing documents into fixed-length batches — the batching that these losses are averaged over.
- Reading a loss curve — what the shape of the curve is telling you.
- Measuring what quantisation costs you — where perplexity misleads badly.
Researcher — Mathematics and papers.
Definition
For a tokenised sequence $x_{1:T}$, perplexity is the exponentiated mean negative log-likelihood:
$$ \mathrm{PPL}(x_{1:T}) = \exp!\left(-\frac{1}{T}\sum_{t=1}^{T}\log p_\theta(x_t \mid x_{<t})\right) $$
$T$ is the token count, $p_\theta$ the model's conditional distribution. Since the exponent is exactly the cross-entropy loss $\mathcal{L}$ in nats, $\mathrm{PPL} = e^{\mathcal{L}}$.
Equivalently, perplexity is the reciprocal geometric mean of the assigned token probabilities:
$$ \mathrm{PPL} = \left(\prod_{t=1}^{T} p_\theta(x_t \mid x_{<t})\right)^{-1/T} $$
The geometric mean is why a single near-zero probability dominates the score, as the toy example above demonstrated.
Relation to entropy and cross-entropy
Under the true distribution $q$:
$$ \mathbb{E}q[\mathcal{L}] = H(q, p\theta) = H(q) + D_{\mathrm{KL}}(q | p_\theta) $$
So $\mathrm{PPL} \geq e^{H(q)}$, with equality only for a perfect model. Perplexity has a floor set by the data, not by the model. Reported perplexity below the entropy of the test set is evidence of contamination, not of a better model.
Bits per byte
Perplexity is not comparable across tokenisers. Bits per byte is:
$$ \mathrm{BPB} = \frac{\mathcal{L}{\text{total, nats}}}{\ln(2)\cdot n{\text{bytes}}} $$
$\mathcal{L}{\text{total, nats}}$ is the summed negative log-likelihood over the document, and $n{\text{bytes}}$ its length in UTF-8 bytes. Because the denominator is a property of the text rather than of the tokeniser, BPB compares any two models on the same corpus. The Pile (Gao et al., 2020) popularised it for exactly this reason, and it is the correct metric whenever vocabulary sizes differ.
The conversion, when the tokeniser is held fixed, is
$$ \mathrm{BPB} = \frac{\log_2 \mathrm{PPL} \cdot n_{\text{tokens}}}{n_{\text{bytes}}} $$
The strided-window protocol
A model with context $C$ cannot score a document longer than $C$ in one pass. Three protocols exist and they do not agree:
- Disjoint chunks. Split into non-overlapping blocks of $C$. Cheap, and pessimistic — every block's first tokens are scored with almost no context.
- Fully sliding window, stride 1. Score each token with the maximum available context. Optimal, and costs one forward pass per token.
- Strided window, stride $s < C$. Advance by $s$, score only the final $s$ tokens of each window. The standard compromise; the HuggingFace perplexity guide uses it.
Reported perplexity varies by tens of percent across these. Any comparison must state the stride. This is the largest single source of irreproducible perplexity numbers in the literature.
What perplexity does and does not predict
Perplexity correlates strongly with downstream capability within an architecture family trained on similar data, which is why it anchors scaling laws. It correlates far more weakly across families.
Three documented failure modes:
- Post-training breaks the link. Instruction tuning and RLHF typically raise perplexity on generic web text while sharply improving usefulness. Perplexity is a pretraining metric; it is close to useless for judging an aligned chat model.
- Emergence is invisible. Wei et al., 2022 reported capabilities appearing abruptly with scale while loss fell smoothly. Schaeffer et al., 2023 (Are Emergent Abilities of Large Language Models a Mirage?) argued convincingly that the abruptness is an artefact of discontinuous metrics such as exact match, not of the underlying continuous improvement. Both papers agree perplexity alone will not show you the transition.
- Quantisation hides damage. A 4-bit model can move perplexity by under 1% while losing several points on reasoning benchmarks. See measuring what quantisation costs you.
Related measures
| Measure | Definition | Use |
|---|---|---|
| Perplexity | $e^{\mathcal{L}}$ | fixed tokeniser, pretraining progress |
| Bits per byte | $\mathcal{L}{\text{tot}} / (\ln 2 \cdot n{\text{bytes}})$ | cross-tokeniser comparison |
| Bits per character | as BPB, over characters | older character-LM literature |
| Token accuracy | fraction where $\arg\max$ is correct | insensitive to calibration; weak |
Perplexity is a proper scoring rule in the strict sense: it is uniquely minimised by reporting the true conditional distribution. That is what token accuracy loses — see proper scoring rules.
Papers and sources
- Jelinek et al., Perplexity — a measure of the difficulty of speech recognition tasks, 1977 — the original definition.
- Brown et al., An Estimate of an Upper Bound for the Entropy of English, 1992 — Computational Linguistics 18(1).
- Gao et al., The Pile, 2020 — arxiv.org/abs/2101.00027
- Wei et al., Emergent Abilities of Large Language Models, 2022 — arxiv.org/abs/2206.07682
- Schaeffer et al., Are Emergent Abilities of Large Language Models a Mirage?, 2023 — arxiv.org/abs/2304.15004
- HuggingFace, Perplexity of fixed-length models — huggingface.co/docs/transformers/perplexity
What to learn next
- Packing documents into fixed-length batches — the batching that these losses are averaged over.
- Reading a loss curve — what the shape of the curve is telling you.
- Measuring what quantisation costs you — where perplexity misleads badly.