Reading logprobs
A logprob is the model telling you how likely it thought each word was, and reading them is the cheapest confidence signal you can get out of a language model.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A logprob tells you how likely the model thought a word was, at the moment it wrote it.
Think about a weather forecast that says seventy percent chance of rain. That number is more useful than a flat "it will rain". You take an umbrella either way, but you now know how much to trust the forecast.
Every word a model writes comes with a number like that. Most tools throw it away before you ever see it.
Turning it back on is the cheapest way to find out where a model was guessing.
What the number actually is
At every step, the model gives every word it knows a probability. Those probabilities add up to one across the whole vocabulary.
A logprob is that probability written on a different scale, because raw probabilities get impractically small. Zero means completely certain. Negative numbers mean less certain, and the further from zero, the less certain.
You do not need to do the conversion by hand. Every library does it for you. What matters is the direction: closer to zero means the model was more sure.
Why the scale is not percentages
Multiply thirty probabilities together and the result is so small that a computer rounds it to zero. That would make a long sentence's overall likelihood unusable.
On the log scale, multiplying becomes adding. Adding thirty numbers is safe. That is the only reason for the strange units.
What you can do with them
Find where the model was guessing. Run through the answer and look for the word with the lowest number. That word is where a confident-sounding sentence stops being supported.
Catch made-up specifics. A model recalling a fact it knows well produces high numbers. A model inventing a plausible name or date usually does not. This is a genuine signal, and it is imperfect.
Compare candidate answers. Generate three answers, total up their numbers, and prefer the one the model itself found most natural.
Measure how surprised the model is by your text. Feed the model a document it has never seen and add up the numbers. That total is the basis of a score called perplexity.
How it looks
"The capital of France is Paris."
token how sure the model was
capital very unsure <- it had no idea this word was coming
of very sure
France unsure
is fairly sure
Paris unsure <- it did not actually know this
. sureThe last line is the interesting one. A small model can produce the correct word without being confident about it. Correctness and confidence are separate things.
Where you have already seen this
- A speech-to-text app showing uncertain words in a lighter colour.
- Autocomplete offering three suggestions, most likely first.
- A translation tool flagging a sentence it is unsure about.
- Spam filters scoring how unusual an email's wording is.
What is honestly hard here
A confident model is not a correct model. Models are routinely confident and wrong, and the failure is silent.
So logprobs give you a signal, not a verdict. Low confidence is a useful warning. High confidence is not a guarantee, and treating it as one has caused real harm in real deployments.
Remember this
- A logprob says how likely the model thought each word was.
- Closer to zero means more confident; far from zero means guessing.
- Low confidence is a real warning; high confidence is not proof.
What to learn next
- Speculative decoding — an algorithm built entirely on comparing two models' probabilities.
- Hallucination — the failure logprobs partly, and only partly, detect.
- Model evaluation — where confidence scores belong in a measurement plan.
Developer — Code and libraries.
Setup
pip install torch transformersThis one uses a real model. distilgpt2 is about 340 MB and runs on a CPU in a second or two. It downloads the first time and prints a progress bar that is not shown in the output below.
Getting the numbers out
import torch, torch.nn.functional as F
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("distilgpt2")
model = AutoModelForCausalLM.from_pretrained("distilgpt2").eval()
def next_token_table(prompt, k=5):
ids = tok(prompt, return_tensors="pt").input_ids
with torch.no_grad():
logits = model(ids).logits[0, -1] # the LAST position predicts the next token
logprobs = F.log_softmax(logits, dim=-1) # log of the probability, in nats
top = torch.topk(logprobs, k)
print(f"prompt: {prompt!r}")
for lp, i in zip(top.values, top.indices):
print(f" {tok.decode(i)!r:<14} logprob {lp:>7.3f} probability {lp.exp():>6.2%}")
print()
next_token_table("The capital of France is")
next_token_table("The capital of Burkina Faso is")
def score(text):
"""Total and per-token log-probability of a whole string."""
ids = tok(text, return_tensors="pt").input_ids
with torch.no_grad():
logits = model(ids).logits
lp = F.log_softmax(logits[0, :-1], dim=-1) # position i predicts token i+1
chosen = lp.gather(1, ids[0, 1:, None]).squeeze(1)
return ids[0], chosen
for text in ("The capital of France is Paris.",
"The capital of France is Madrid."):
ids, chosen = score(text)
total = chosen.sum().item()
ppl = torch.exp(-chosen.mean()).item()
print(f"{text!r}")
print(f" total logprob {total:>8.3f} mean {chosen.mean():>7.3f} perplexity {ppl:>8.2f}")
for i, lp in enumerate(chosen):
print(f" {tok.decode(ids[i + 1])!r:<10}{lp:>8.3f}")
print()prompt: 'The capital of France is'
' the' logprob -2.139 probability 11.77%
' home' logprob -2.498 probability 8.23%
' now' logprob -2.998 probability 4.99%
' a' logprob -3.009 probability 4.93%
' being' logprob -3.335 probability 3.56%
prompt: 'The capital of Burkina Faso is'
' the' logprob -2.593 probability 7.48%
' home' logprob -2.617 probability 7.30%
' now' logprob -3.086 probability 4.57%
' a' logprob -3.194 probability 4.10%
' being' logprob -3.198 probability 4.08%
'The capital of France is Paris.'
total logprob -30.065 mean -5.011 perplexity 150.02
' capital' -14.300
' of' -0.951
' France' -4.620
' is' -3.009
' Paris' -5.891
'.' -1.293
'The capital of France is Madrid.'
total logprob -31.838 mean -5.306 perplexity 201.62
' capital' -14.300
' of' -0.951
' France' -4.620
' is' -3.009
' Madrid' -7.926
'.' -1.032Run against transformers 5.6.2 and torch 2.5.1 on CPU. Greedy scoring of a fixed string is deterministic, so these numbers reproduce. The last decimal can differ on other hardware, for the reason covered in temperature zero is still not deterministic.
Reading the output, which is more interesting than it looks
distilgpt2 does not know that the capital of France is Paris. Its top five continuations are the, home, now, a, being. Its favourite is the, at under twelve percent. Small models are grammar engines with patchy facts, and this is what that looks like numerically.
The Burkina Faso prompt gives almost the same list. Same five tokens, same order, slightly lower confidence. The model has no country-specific knowledge in play at all — it is completing the pattern "The capital of X is ...". A model that does not know need not look different from one that does. That is why an eyeball check of the top token is a weak test.
Paris scored -5.891 and Madrid scored -7.926. Paris is about seven and a half times more likely. Both are unlikely in absolute terms — Paris is at roughly 0.28 percent. The right answer won, and it won without the model being remotely confident. Confidence and correctness came apart in a single example.
capital scored -14.300, by far the worst token in both sentences. Nothing in "The" predicts "capital". This is the most useful habit to build: the lowest-scoring tokens are usually places where information arrived from outside the model rather than places where the model was wrong.
Perplexity is 150.02 for the Paris sentence. Perplexity is the exponential of the average negative logprob. It reads as "about as unsure as choosing uniformly among 150 options each step". Lower is less surprised. The Madrid sentence scores 201.62, correctly ranking the two.
The first token has no score. Six tokens are listed for a seven-token sentence. Nothing predicts The, so it has no probability under a causal model. Off-by-one errors here are the most common bug in home-made scoring code, and they inflate scores in a way that is hard to notice.
Getting logprobs from an API
Most providers expose them, with a cap on how many alternatives per position:
resp = client.chat.completions.create(
model="...", messages=[...],
logprobs=True, top_logprobs=5,
)
for tok in resp.choices[0].logprobs.content:
print(tok.token, tok.logprob, [(a.token, a.logprob) for a in tok.top_logprobs])No output block: the values depend on the provider and the model, and inventing them would teach you to expect something that will not happen. The shape of the response is what matters, and it is stable across OpenAI-compatible servers including vLLM.
Two limits worth knowing. Reasoning models often do not return logprobs for hidden reasoning tokens. And top_logprobs is capped, typically at 20, so you see the head of the distribution and never the tail.
Three things logprobs are genuinely good for
Extraction confidence. For a field the model filled in, take the minimum token logprob across that field's tokens. Route anything below a threshold to a human. This is the highest-value use of logprobs in production, and it needs no extra model call.
Choosing between fixed options. For a classification with labels yes and no, compare the logprob of each label token directly instead of generating and string-matching. Cheaper, and it gives you a graded score rather than a bare answer.
Detecting a model out of its depth. A rolling average of per-token logprob across a conversation drifting downward means the model is on unfamiliar ground. Useful as a monitoring signal, not as a hard gate.
Common mistakes
Comparing totals across different lengths. Longer text always has a lower total. Compare mean logprob per token, or perplexity, never the sum.
Comparing perplexity across tokenizers. A model that splits text into more tokens gets a different perplexity on the same string. Cross-model perplexity comparisons are only meaningful with an identical tokenizer.
Treating high confidence as truth. Models are frequently confident and wrong, and the confident errors are the ones that reach users. Logprobs measure fluency-of-fit, not factual support. For that you want RAG and a grounding check.
Reading logprobs after truncation. With top-p or top-k in play, the reported logprobs may come from the filtered distribution rather than the raw one, depending on the server. If you are using logprobs as a confidence signal, request them at temperature zero with no truncation.
Forgetting the off-by-one. logits[i] predicts token[i+1]. Every scoring bug in this area is this line.
Try it yourself
Score "The capital of France is Delhi." and compare against Paris and Madrid. Then score all three under a larger model such as gpt2-medium and watch the gap widen enormously. The size of that gap is a rough measure of how much the model actually knows, as opposed to how well it forms sentences.
What to learn next
- Speculative decoding — an algorithm built entirely on comparing two models' probabilities.
- Hallucination — the failure logprobs partly, and only partly, detect.
- Model evaluation — where confidence scores belong in a measurement plan.
Researcher — Mathematics and papers.
Definitions
For a causal model with parameters $\theta$, the log-probability of a sequence factorises:
$$ \log P_\theta(y) = \sum_{t=1}^{T} \log P_\theta(y_t \mid y_{<t}) $$
Each term comes from a log-softmax over logits $z^{(t)} \in \mathbb{R}^{|V|}$:
$$ \log P_\theta(y_t \mid y_{<t}) = z^{(t)}_{y_t} - \log \sum_{v \in V} \exp z^{(t)}_{v} $$
Compute this with log_softmax, never as log(softmax(z)). The fused form subtracts the maximum before exponentiating and is stable at logit magnitudes where the naive form returns $-\infty$.
Perplexity over $T$ tokens:
$$ \mathrm{PPL}(y) = \exp!\left( -\frac{1}{T} \sum_{t=1}^{T} \log P_\theta(y_t \mid y_{<t}) \right) $$
which is the exponential of the cross-entropy in nats, interpretable as an effective branching factor.
Perplexity is tokenizer-dependent, and this is not a technicality. The same string under two tokenizations has different $T$ and different per-token terms. Comparing perplexity across model families is meaningless unless the tokenizer is shared. Bits-per-byte, $\frac{-\log_2 P(y)}{|\text{bytes}(y)|}$, is the tokenizer-invariant alternative and is what careful scaling-law work reports.
Calibration
A model is calibrated if, among tokens it assigns probability $p$, a fraction $p$ are correct. Measured by expected calibration error over confidence bins.
Three findings that are stable enough to rely on:
- Base language models are reasonably well calibrated on multiple-choice tasks at scale (Kadavath et al., Language Models (Mostly) Know What They Know, 2022, arxiv.org/abs/2207.05221).
- RLHF degrades calibration substantially. The OpenAI GPT-4 technical report shows a well-calibrated pre-trained model and a badly calibrated post-trained one on the same benchmark. Alignment training sharpens the distribution toward confident-sounding output regardless of correctness.
- Verbalised confidence — asking the model to state a percentage — is worse calibrated than the logprobs, and is also easier to collect, which is why it is used anyway.
The practical reading: logprobs from an instruction-tuned chat model are a ranking signal you should calibrate yourself on held-out data, not a probability you can take at face value.
Sequence-level confidence
Several aggregations, with different failure modes:
| Aggregation | Formula | Behaviour |
|---|---|---|
| Sum | $\sum_t \log p_t$ | Penalises length; use only at fixed length |
| Mean | $\frac{1}{T}\sum_t \log p_t$ | Length-normalised; a single bad token is diluted |
| Minimum | $\min_t \log p_t$ | Finds the weakest link; sensitive to formatting tokens |
| Entropy | $-\sum_v p_v \log p_v$ per step | Uncertainty over the whole distribution, not only the chosen token |
Entropy deserves more use than it gets. Logprob of the chosen token conflates two situations: a peaked distribution where the model is sure, and a flat one where it happened to pick a token with moderate probability. Entropy separates them. Predictive-entropy methods for hallucination detection are built on that separation (Farquhar et al., Detecting Hallucinations in Large Language Models Using Semantic Entropy, Nature 2024).
Uses in the literature
- Membership inference and contamination detection. Anomalously high logprobs on a benchmark's test set relative to comparable held-out text is evidence of training-set contamination. Min-K% Prob (Shi et al., 2023, arxiv.org/abs/2310.16789) uses the mean of the $k\%$ lowest-probability tokens, on the reasoning that memorised text lacks unusually surprising tokens.
- Machine-generated text detection. DetectGPT (Mitchell et al., 2023, arxiv.org/abs/2301.11305) starts from one observation. Model-generated text sits near a local maximum of the model's log-probability, so small perturbations lower the score more than for human text.
- Reranking. Sample $n$ candidates, rerank by mean logprob or by a separate reward model. Mean logprob is a weak reranker on its own — it prefers bland text, for the reason discussed in beam search.
- Speculative decoding. The acceptance rule compares draft and target probabilities directly, which is the next lesson.
Papers
- Kadavath et al., Language Models (Mostly) Know What They Know, 2022 — arxiv.org/abs/2207.05221
- Mitchell et al., DetectGPT, 2023 — arxiv.org/abs/2301.11305
- Shi et al., Detecting Pretraining Data from Large Language Models, 2023 — arxiv.org/abs/2310.16789
- Farquhar et al., Detecting Hallucinations in Large Language Models Using Semantic Entropy, Nature, 2024
What to learn next
- Speculative decoding — an algorithm built entirely on comparing two models' probabilities.
- Hallucination — the failure logprobs partly, and only partly, detect.
- Model evaluation — where confidence scores belong in a measurement plan.