AI glossary

Perplexity

In one sentence Perplexity measures how surprised a language model is by text — lower means the model predicts that text better.

By Updated

Perplexity measures how well a language model predicts a piece of text: the lower the number, the less the text surprises the model.

The intuition is a guessing game. Read someone a sentence, pausing before each word, and ask them to guess it. A fluent English speaker guessing a newspaper sentence is rarely shocked — most words were among their top few guesses. Someone with weak English is constantly surprised. Perplexity is that surprise, averaged over a whole text and put on a standard scale.

The scale has a neat reading: a perplexity of 20 means the model is, on average, as uncertain as if it were choosing between 20 equally likely tokens at each step. Perplexity 1 would be perfect prediction; a model choosing at random from a 50,000-token vocabulary would score 50,000. Mathematically it is the exponential of the average cross-entropy loss — the same loss-function that training pushes down, which is why training curves and perplexity move together.

Where you meet it: comparing base models during pretraining, checking how much quantization damaged a model, and detecting domain mismatch (a model with low perplexity on news may have high perplexity on legal text). Two cautions: scores are only comparable between models sharing a tokenizer, and low perplexity means fluent prediction, not truthfulness or usefulness — chat quality is judged by evals instead.

Where to go next