How Models Are Actually Trained

Cross-entropy and perplexity

Perplexity turns a training loss number into a plain statement — how many equally likely options the model was choosing between at each step.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. Where you have already seen this
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Perplexity says how confused the model was, measured as "how many options did it seem to be picking between".

The analogy you have already lived

You are watching a quiz show and trying to answer along. On an easy question you narrow it down to two options before the host finishes. On a hard one you are stuck between all four.

Someone could describe your whole evening with one number. On average, you were choosing between about three options per question. That number is your perplexity.

A perplexity of 1 means you knew every answer with certainty. A perplexity of 4 on a four-option quiz means you were guessing throughout.

Why it exists

Training reports a loss — a score of how wrong the model was, where lower is better. Loss is fine for the optimiser, but it is a hard number for a human to feel.

Nobody has an instinct for whether a loss of 2.9 is good. Everybody has an instinct for "the model was picking between about eighteen words".

Perplexity is the loss translated into that second sentence. It is the same information wearing friendlier clothes.

How it works

   model's confidence          loss        perplexity
   ------------------          ----        ----------
   sure, and right             tiny        near 1
   leaning the right way       small       2 to 5
   no idea, spread evenly      medium      the vocabulary size
   sure, and WRONG             huge        enormous

Two things are worth burning into memory.

A model that has learned nothing scores its whole vocabulary size. Say it knows fifty thousand words and spreads its guess evenly. Then it is choosing between fifty thousand options. That is the starting line.

Being confidently wrong is punished far harder than being unsure. A shrug costs a little. A wrong answer stated with total certainty costs a great deal. This is deliberate, and it is why models learn to hedge.

Where you have already seen this

  • Research papers reporting "perplexity 3.2 on WikiText" to compare models.
  • Model cards on Hugging Face listing perplexity per quantised version.
  • Spam filters and old speech recognisers, which used the same measure.

What is honestly hard here

Perplexity numbers from two different papers are often not comparable, and people compare them anyway.

The number depends on how the text was cut into pieces. It also depends on which text was used, and on the window size. Change any of those and the number moves, with no change to the model.

Treat perplexity as a measurement of one model on one dataset with one setup. It is a thermometer, not a leaderboard.

Remember this

  • Perplexity is "how many options the model seemed to be choosing between".
  • A clueless model scores its whole vocabulary size. Lower is better.
  • Numbers from different papers, datasets or tokenisers are not comparable.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

The second example adds pip install transformers and downloads GPT-2, about 550 MB.

The definition, made concrete

perplexity_basics.py
import math
import torch
import torch.nn.functional as F

V = 8  # a toy vocabulary of eight tokens

cases = {
    "knows nothing (uniform)": torch.full((V,), 1.0 / V),
    "leans the right way":     torch.tensor([.40, .20, .10, .10, .08, .06, .04, .02]),
    "confident and correct":   torch.tensor([.93, .02, .02, .01, .01, .005, .003, .002]),
    "confident and WRONG":     torch.tensor([.02, .93, .02, .01, .01, .005, .003, .002]),
}

target = torch.tensor([0])  # token 0 is what actually came next

print(f"{'model':<24} {'p(correct)':>10} {'loss':>8} {'perplexity':>11}")
for name, probs in cases.items():
    loss = F.nll_loss(probs.log().unsqueeze(0), target)
    print(f"{name:<24} {probs[0]:>10.3f} {loss:>8.3f} {loss.exp():>11.2f}")

print(f"\nvocabulary size = {V}, so ln(V) = {math.log(V):.3f} and exp(ln(V)) = {V}")

# Perplexity is an average over MANY tokens, not one.
seq_probs = torch.tensor([0.9, 0.9, 0.9, 0.9, 0.001])   # four easy tokens, one shock
nll = -seq_probs.log()
print("\nper-token loss:", [f"{v:.2f}" for v in nll.tolist()])
print(f"mean loss  = {nll.mean():.4f}")
print(f"perplexity = {nll.mean().exp():.2f}")
print(f"perplexity of the first four alone = {(-seq_probs[:4].log()).mean().exp():.2f}")
Output
model                    p(correct)     loss  perplexity
knows nothing (uniform)       0.125    2.079        8.00
leans the right way           0.400    0.916        2.50
confident and correct         0.930    0.073        1.08
confident and WRONG           0.020    3.912       50.00

vocabulary size = 8, so ln(V) = 2.079 and exp(ln(V)) = 8

per-token loss: ['0.11', '0.11', '0.11', '0.11', '6.91']
mean loss  = 1.4658
perplexity = 4.33
perplexity of the first four alone = 1.11

Reading that output

The uniform model scores exactly 8.00 on an 8-token vocabulary. That is the whole intuition in one line: perplexity is an effective number of choices.

Confidently wrong scores 50, on a vocabulary of 8. Perplexity is not bounded by the vocabulary size. A model can be far worse than clueless, and this asymmetry is the single most important property of the measure.

One bad token dragged perplexity from 1.11 to 4.33. Four tokens at 90% confidence, one at 0.1%, and the average is wrecked. Perplexity is a geometric mean over tokens, so rare disasters dominate it. When your evaluation perplexity jumps, look for a handful of catastrophic tokens before you suspect a general regression.

Real perplexity on a real model

gpt2_perplexity.py
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModelForCausalLM

tok = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2").eval()


def token_losses(text):
    ids = tok(text, return_tensors="pt").input_ids
    with torch.no_grad():
        logits = model(ids).logits
    # position i predicts token i+1, so drop the last logit and the first label
    losses = F.cross_entropy(logits[0, :-1], ids[0, 1:], reduction="none")
    return ids[0], losses


for text in ["The capital of France is Paris",
             "The capital of France is Kolkata",
             "France of capital the is Paris"]:
    ids, losses = token_losses(text)
    ppl = losses.mean().exp().item()
    print(f"\n{text!r}   {len(ids)} tokens, perplexity {ppl:.1f}")
    for i, loss in enumerate(losses.tolist()):
        print(f"    predict {tok.decode(ids[i + 1])!r:<10} loss {loss:5.2f}")
Output
'The capital of France is Paris'   6 tokens, perplexity 72.9
    predict ' capital' loss  9.07
    predict ' of'      loss  1.81
    predict ' France'  loss  5.03
    predict ' is'      loss  2.11
    predict ' Paris'   loss  3.43

'The capital of France is Kolkata'   8 tokens, perplexity 73.8
    predict ' capital' loss  9.07
    predict ' of'      loss  1.81
    predict ' France'  loss  5.03
    predict ' is'      loss  2.11
    predict ' K'       loss  8.68
    predict 'olk'      loss  3.41
    predict 'ata'      loss  0.00

'France of capital the is Paris'   6 tokens, perplexity 3193.3
    predict ' of'      loss  4.13
    predict ' capital' loss 10.55
    predict ' the'     loss  6.10
    predict ' is'      loss 10.12
    predict ' Paris'   loss  9.45

Written against transformers 5.6.2 and PyTorch 2.5.1, CPU, greedy scoring with no sampling — so these numbers reproduce exactly. Transformers may print a loss_type=None informational line for GPT-2; it is harmless.

This output contains four lessons

The model prefers the truth, and the evidence is per token. Paris costs 3.43, K (starting Kolkata) costs 8.68. That gap is the model's factual knowledge, visible directly.

Sentence-level perplexity hid that completely: 72.9 against 73.8. The false sentence scored almost identically, because four shared tokens dominated the average and the extra tokens diluted it. Never use whole-sentence perplexity to test whether a model knows a fact. Score the answer tokens only.

'ata' cost 0.00. After ' K' and 'olk', nothing else can follow. Multi-token words hand the model free, meaningless wins. A tokenizer that splits a word into five pieces makes your perplexity look better without the model improving at all.

The first token is always expensive. ' capital' cost 9.07 after only 'The'. Short texts are dominated by this cold-start penalty, which is why perplexity is reported over long documents.

Common mistakes

Forgetting the shift. Comparing logits[i] to ids[i] instead of ids[i+1] gives a suspiciously low number. Using model(ids, labels=ids) does the shift for you and is the safer route.

Averaging perplexities. Perplexity is exponentiated, so averaging two perplexities is wrong. Average the losses, then exponentiate once at the end.

Ignoring padding. With batches, padding tokens must be labelled -100 so cross_entropy skips them. Without that, you are measuring how well the model predicts padding, and perplexity looks great. See padding, truncation and attention masks.

Comparing across tokenisers. A model with a 32k vocabulary and one with a 256k vocabulary cut the same paragraph into different numbers of tokens. Their per-token perplexities are different units. Normalise by characters or bytes — bits per byte — if you must compare.

Evaluating on text the model trained on. Perplexity on training data is near-meaningless, and web-scale training corpora contain most public benchmark text. See benchmark contamination.

Try it yourself

Score "The capital of France is" followed by each of Paris, Lyon, Delhi, Bangalore. Print the loss of the answer tokens only, ignoring the shared prefix. That is how a factual-knowledge probe is actually built.

What to learn next

Researcher — Mathematics and papers.

Definition

For a tokenised sequence $x_{1:T}$, perplexity is the exponentiated mean negative log-likelihood:

$$ \mathrm{PPL}(x_{1:T}) = \exp!\left(-\frac{1}{T}\sum_{t=1}^{T}\log p_\theta(x_t \mid x_{<t})\right) $$

$T$ is the token count, $p_\theta$ the model's conditional distribution. Since the exponent is exactly the cross-entropy loss $\mathcal{L}$ in nats, $\mathrm{PPL} = e^{\mathcal{L}}$.

Equivalently, perplexity is the reciprocal geometric mean of the assigned token probabilities:

$$ \mathrm{PPL} = \left(\prod_{t=1}^{T} p_\theta(x_t \mid x_{<t})\right)^{-1/T} $$

The geometric mean is why a single near-zero probability dominates the score, as the toy example above demonstrated.

Relation to entropy and cross-entropy

Under the true distribution $q$:

$$ \mathbb{E}q[\mathcal{L}] = H(q, p\theta) = H(q) + D_{\mathrm{KL}}(q | p_\theta) $$

So $\mathrm{PPL} \geq e^{H(q)}$, with equality only for a perfect model. Perplexity has a floor set by the data, not by the model. Reported perplexity below the entropy of the test set is evidence of contamination, not of a better model.

Bits per byte

Perplexity is not comparable across tokenisers. Bits per byte is:

$$ \mathrm{BPB} = \frac{\mathcal{L}{\text{total, nats}}}{\ln(2)\cdot n{\text{bytes}}} $$

$\mathcal{L}{\text{total, nats}}$ is the summed negative log-likelihood over the document, and $n{\text{bytes}}$ its length in UTF-8 bytes. Because the denominator is a property of the text rather than of the tokeniser, BPB compares any two models on the same corpus. The Pile (Gao et al., 2020) popularised it for exactly this reason, and it is the correct metric whenever vocabulary sizes differ.

The conversion, when the tokeniser is held fixed, is

$$ \mathrm{BPB} = \frac{\log_2 \mathrm{PPL} \cdot n_{\text{tokens}}}{n_{\text{bytes}}} $$

The strided-window protocol

A model with context $C$ cannot score a document longer than $C$ in one pass. Three protocols exist and they do not agree:

  1. Disjoint chunks. Split into non-overlapping blocks of $C$. Cheap, and pessimistic — every block's first tokens are scored with almost no context.
  2. Fully sliding window, stride 1. Score each token with the maximum available context. Optimal, and costs one forward pass per token.
  3. Strided window, stride $s < C$. Advance by $s$, score only the final $s$ tokens of each window. The standard compromise; the HuggingFace perplexity guide uses it.

Reported perplexity varies by tens of percent across these. Any comparison must state the stride. This is the largest single source of irreproducible perplexity numbers in the literature.

What perplexity does and does not predict

Perplexity correlates strongly with downstream capability within an architecture family trained on similar data, which is why it anchors scaling laws. It correlates far more weakly across families.

Three documented failure modes:

  • Post-training breaks the link. Instruction tuning and RLHF typically raise perplexity on generic web text while sharply improving usefulness. Perplexity is a pretraining metric; it is close to useless for judging an aligned chat model.
  • Emergence is invisible. Wei et al., 2022 reported capabilities appearing abruptly with scale while loss fell smoothly. Schaeffer et al., 2023 (Are Emergent Abilities of Large Language Models a Mirage?) argued convincingly that the abruptness is an artefact of discontinuous metrics such as exact match, not of the underlying continuous improvement. Both papers agree perplexity alone will not show you the transition.
  • Quantisation hides damage. A 4-bit model can move perplexity by under 1% while losing several points on reasoning benchmarks. See measuring what quantisation costs you.
MeasureDefinitionUse
Perplexity$e^{\mathcal{L}}$fixed tokeniser, pretraining progress
Bits per byte$\mathcal{L}{\text{tot}} / (\ln 2 \cdot n{\text{bytes}})$cross-tokeniser comparison
Bits per characteras BPB, over charactersolder character-LM literature
Token accuracyfraction where $\arg\max$ is correctinsensitive to calibration; weak

Perplexity is a proper scoring rule in the strict sense: it is uniquely minimised by reporting the true conditional distribution. That is what token accuracy loses — see proper scoring rules.

Papers and sources

  • Jelinek et al., Perplexity — a measure of the difficulty of speech recognition tasks, 1977 — the original definition.
  • Brown et al., An Estimate of an Upper Bound for the Entropy of English, 1992 — Computational Linguistics 18(1).
  • Gao et al., The Pile, 2020 — arxiv.org/abs/2101.00027
  • Wei et al., Emergent Abilities of Large Language Models, 2022 — arxiv.org/abs/2206.07682
  • Schaeffer et al., Are Emergent Abilities of Large Language Models a Mirage?, 2023 — arxiv.org/abs/2304.15004
  • HuggingFace, Perplexity of fixed-length models — huggingface.co/docs/transformers/perplexity

What to learn next