Repetition and frequency penalties
Three different ways to make a model stop repeating itself, two of which work differently from how almost everyone assumes, and one of which quietly punishes your own prompt.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A repetition penalty lowers the score of words the model has already used. That makes it less likely to say them again.
Think of a friend who tells the same story at every family gathering. The first time it is charming. By the fourth telling, everyone winces.
Now imagine charging a small fine for each retelling. Nothing is banned. The story is still allowed. It has become slightly less attractive to tell, and other stories get a turn.
That is exactly what these penalties do. They do not forbid a word. They tax it.
Why models repeat at all
A language model writes by looking at what it has already written. So its own output becomes the strongest hint about what comes next.
Write "the office is closed on Sunday" and that sentence is now the most recent, most relevant text in view. The most fitting next words often look like the ones already there.
That creates a loop. Say something, be nudged toward saying it again, say it again, be nudged harder. Left alone, models fall into these loops and stay there.
You have seen this. A chatbot repeating a phrase forever, or a model that ends every paragraph with the same closing sentence.
The three penalties, and how they differ
There are three in common use and they are not the same thing.
Repetition penalty divides the word's score, so it scales with how strong the score already was. Used or unused is all it checks — using a word five times costs the same as using it once.
Frequency penalty subtracts a fixed amount for each use. Say a word five times and you pay five times.
Presence penalty subtracts a fixed amount once, the first time a word is used at all. It pushes for new vocabulary rather than against repeated vocabulary.
How it works
the model's scores after "the cat the the"
word score times used after a repetition penalty
the 2.00 3 1.67 (taxed)
cat 1.40 1 1.17 (taxed)
sat 0.90 0 0.90 (untouched)
on 0.30 0 0.30 (untouched)"the" and "cat" both drop. "sat" and "on" do not move, so they gain ground without being helped.
Notice "the" was used three times and "cat" once, and both were taxed the same way. That is the repetition penalty's blind spot, and it is what the frequency penalty fixes.
The part everybody gets wrong
These penalties punish every word the model has seen, and that includes your prompt.
Paste in a document and ask for a summary. Every important word in that document has now been used. A strong penalty makes the model avoid those exact words in the summary, so it reaches for vaguer alternatives.
The summary gets worse and nobody can see why. Good libraries let you exclude the prompt from the count. Many do not.
Where you have already seen this
- A chatbot stuck repeating one sentence until you stop it.
- API settings named
frequency_penaltyandpresence_penaltyon a slider from zero to two. - A local model tool with a
repeat_penaltybox, usually set near one point one. - A summary that carefully avoids using the document's own words, and reads oddly for it.
Remember this
- Penalties tax words the model has already used, they do not ban them.
- Repetition penalty scales the score; frequency counts uses; presence fires once.
- All three usually punish your prompt too, which quietly damages summaries.
What to learn next
- Beam search, and why chat models dropped it — the decoding method whose main flaw was exactly this repetition.
- Constrained decoding and grammars — the right tool when the output has a required shape.
- Prompt engineering — often a better cure for repetition than any sampling parameter.
Developer — Code and libraries.
Setup
pip install numpyThree penalties, one table, and the sign trap that makes the first one behave unlike anything you would guess.
All three side by side
import numpy as np
VOCAB = ["the", "cat", "sat", "on", "mat", "and", "very", "happy"]
def repetition_penalty(logits, seen, penalty):
"""CTRL-style. Divide positive logits, MULTIPLY negative ones."""
out = logits.copy()
for t in set(seen):
out[t] = out[t] / penalty if out[t] > 0 else out[t] * penalty
return out
def frequency_presence(logits, seen, freq=0.0, presence=0.0):
"""OpenAI-style. Subtract a fixed amount, counted or not counted."""
out = logits.copy()
counts = {t: seen.count(t) for t in set(seen)}
for t, n in counts.items():
out[t] -= freq * n + presence
return out
logits = np.array([2.0, 1.4, 0.9, 0.3, -0.5, -1.2, -2.0, -3.0])
seen = [0, 1, 0, 0] # "the cat the the"
print("token logit rep 1.2 freq 0.5 presence 0.5 seen")
after_rep = repetition_penalty(logits, seen, 1.2)
after_freq = frequency_presence(logits, seen, freq=0.5)
after_pres = frequency_presence(logits, seen, presence=0.5)
for i, w in enumerate(VOCAB):
n = seen.count(i)
print(f"{w:<10}{logits[i]:>6.2f}{after_rep[i]:>10.3f}{after_freq[i]:>11.2f}"
f"{after_pres[i]:>15.2f}{n:>9}")
print("\nThe trap, spelled out:")
print(f" 'mat' started at {logits[4]:>5.2f}. It was never used, so nothing happens.")
neg = repetition_penalty(np.array([-0.5]), [0], 1.2)[0]
print(f" a token at -0.50 that HAS been used becomes {neg:.2f} - pushed further down.")
pos = repetition_penalty(np.array([2.0]), [0], 1.2)[0]
print(f" a token at +2.00 that HAS been used becomes {pos:.2f} - also pushed down.")
print(" Same penalty, opposite arithmetic, because a plain divide would RAISE a")
print(" negative logit. That is the whole reason for the sign test.")
print("\nWhat a penalty of 1.2 costs in probability terms:")
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
for p in (1.0, 1.05, 1.1, 1.2, 1.5, 2.0):
q = softmax(repetition_penalty(logits, seen, p))
print(f" penalty {p:<5} P(the)={q[0]:.3f} P(cat)={q[1]:.3f} P(sat)={q[2]:.3f}")token logit rep 1.2 freq 0.5 presence 0.5 seen the 2.00 1.667 0.50 1.50 3 cat 1.40 1.167 0.90 0.90 1 sat 0.90 0.900 0.90 0.90 0 on 0.30 0.300 0.30 0.30 0 mat -0.50 -0.500 -0.50 -0.50 0 and -1.20 -1.200 -1.20 -1.20 0 very -2.00 -2.000 -2.00 -2.00 0 happy -3.00 -3.000 -3.00 -3.00 0 The trap, spelled out: 'mat' started at -0.50. It was never used, so nothing happens. a token at -0.50 that HAS been used becomes -0.60 - pushed further down. a token at +2.00 that HAS been used becomes 1.67 - also pushed down. Same penalty, opposite arithmetic, because a plain divide would RAISE a negative logit. That is the whole reason for the sign test. What a penalty of 1.2 costs in probability terms: penalty 1.0 P(the)=0.452 P(cat)=0.248 P(sat)=0.150 penalty 1.05 P(the)=0.436 P(cat)=0.246 P(sat)=0.160 penalty 1.1 P(the)=0.421 P(cat)=0.244 P(sat)=0.168 penalty 1.2 P(the)=0.395 P(cat)=0.240 P(sat)=0.183 penalty 1.5 P(the)=0.338 P(cat)=0.226 P(sat)=0.219 penalty 2.0 P(the)=0.282 P(cat)=0.209 P(sat)=0.255
Reading the output
Look at the the row across the three columns. Repetition penalty took it from 2.00 to 1.667, a cut of about seventeen percent of its value. Frequency penalty took it to 0.50, because it was used three times and each use costs 0.5. Presence penalty took it to 1.50, one flat charge. Three penalties, three completely different results on the same token.
cat and the are treated identically by the repetition penalty. One use versus three, same rule. If your complaint is a word appearing over and over, the frequency penalty targets that directly and the repetition penalty does not.
The sign test is not a style choice. A logit of -0.50 divided by 1.2 gives -0.42, which is higher. A plain divide would reward exactly the tokens you meant to punish. HuggingFace, vLLM and llama.cpp all implement score / penalty if score > 0 else score * penalty for this reason. Every reimplementation that skips the branch is broken, and the breakage is invisible until the model starts looping on rare tokens.
A penalty of 1.2 moved P(the) from 0.452 to 0.395. That is a modest nudge, not a ban. To actually break a hard loop you need something stronger, which is why no_repeat_ngram_size exists. And notice P(sat) climbing from 0.150 to 0.255 as the penalty rises — the untouched tokens gain because the survivors are renormalised.
Which one to reach for
| Symptom | Reach for |
|---|---|
| The same phrase repeated verbatim, forever | no_repeat_ngram_size — a hard block, not a tax |
| One word appearing far too often | frequency penalty, 0.1 to 0.5 |
| Vocabulary feels narrow and samey | presence penalty, 0.1 to 0.5 |
| General mild drift toward repetition | repetition penalty, 1.05 to 1.15 |
| Summaries avoiding the source's own words | turn the penalty down, or exclude the prompt |
no_repeat_ngram_size=3 forbids any three-token sequence from appearing twice. It ends loops absolutely, and it also forbids legitimate repeats — a person's name, a technical term, for i in range. On code it is actively harmful.
In the real libraries
# HuggingFace: penalises the prompt too, unless you tell it not to
out = model.generate(**inputs, do_sample=True, repetition_penalty=1.1,
no_repeat_ngram_size=0)
# vLLM: all three available, and its repetition_penalty excludes nothing by default
from vllm import SamplingParams
params = SamplingParams(repetition_penalty=1.1, frequency_penalty=0.2,
presence_penalty=0.0)transformers 5.x exposes a prompt_ignore_length on RepetitionPenaltyLogitsProcessor so the prompt can be excluded. The plain repetition_penalty= argument to generate does not set it. If prompt contamination matters to you, build the processor yourself and pass it in logits_processor=.
Common mistakes
Setting repetition_penalty=2.0 because 1.1 did not work. Above roughly 1.2 the text degrades fast, in a characteristic way: the model starts using unusual synonyms and its grammar frays. If 1.15 has not fixed the problem, the problem is something else.
Applying penalties to structured output. JSON needs ", {, : and , repeatedly. Penalising them produces malformed output. Use constrained decoding instead, and set every penalty to zero.
Penalising chat and code. Code repeats identifiers constantly. Chat repeats the user's own words on purpose. Both are cases where the default should be no penalty at all.
Mixing repetition and frequency penalties without checking. They stack multiplicatively and additively at the same time, which is difficult to reason about. Pick one.
Assuming a penalty fixes a looping model. Persistent loops usually mean the sampling temperature is near zero, or the prompt is degenerate, or the context has been truncated mid-sentence. A penalty masks that, and masking it is worse than seeing it.
Try it yourself
Set seen = [0] * 20 and print the repetition-penalty column. It is unchanged from seen = [0], because the rule uses set(seen). Then do the same with the frequency penalty and watch it collapse. That contrast is the single most useful thing to know about these two parameters.
What to learn next
- Beam search, and why chat models dropped it — the decoding method whose main flaw was exactly this repetition.
- Constrained decoding and grammars — the right tool when the output has a required shape.
- Prompt engineering — often a better cure for repetition than any sampling parameter.
Researcher — Mathematics and papers.
Repetition penalty
Introduced in Keskar et al., CTRL: A Conditional Transformer Language Model for Controllable Generation, 2019 (arxiv.org/abs/1909.05858), section 4.1. For logit $z_i$ and the set $G$ of previously generated token ids:
$$ z_i' = \begin{cases} z_i / \theta & i \in G,\ z_i > 0 \ z_i \cdot \theta & i \in G,\ z_i \le 0 \ z_i & i \notin G \end{cases} $$
with $\theta \ge 1$. CTRL reported $\theta \approx 1.2$ as a reasonable balance between truthful continuation and lack of repetition.
The transformation is not shift-invariant, and that is its defining pathology. Logits are only defined up to an additive constant. A model whose logits are all shifted by $+5$ is the same model, and this penalty produces a different result on it. Two checkpoints that are functionally identical can require different $\theta$. It is worth understanding that this parameter is not measuring what its name suggests.
Frequency and presence penalties
From the OpenAI API, and additive rather than multiplicative:
$$ z_i' = z_i - \alpha_{\text{freq}} \cdot c_i - \alpha_{\text{pres}} \cdot \mathbb{1}[c_i > 0] $$
where $c_i$ is the number of times token $i$ has appeared. Both parameters range over $[-2, 2]$ in that API, and negative values encourage repetition, which is occasionally useful for forced formats.
These are shift-invariant, since subtracting a constant from one logit changes the softmax in a well-defined way regardless of the overall offset. They are the better-behaved formulation, and the fact that the older, worse-behaved one is more widely deployed is an accident of history.
Hard constraints
no_repeat_ngram_size = n sets $z_i = -\infty$ for any token that would complete an $n$-gram already present in the sequence. It is usually credited to Paulus et al., 2017 in abstractive summarisation; the implementation everyone copies is Fairseq's, which is what the HuggingFace processor cites.
This is a blunt instrument with a real cost: it makes some correct outputs unreachable. On a document mentioning "Reserve Bank of India" twice, $n = 3$ makes the second mention impossible to produce. Any text with legitimate repeated phrasing — legal, medical, code — is damaged by it in a way that no penalty tuning recovers.
Why repetition happens
Two accounts, and they are not exclusive.
Self-reinforcement. Xu et al., 2022 (arxiv.org/abs/2206.02369) measure that the more times a sentence already appears in the context, the higher the probability of generating it again. Sentences that start with a higher probability show a stronger version of the effect. The loop is an attractor, not a coincidence.
Likelihood as the wrong objective. Welleck et al., Neural Text Degeneration with Unlikelihood Training, 2019 (arxiv.org/abs/1908.04319) blame the training objective. Maximum-likelihood training assigns too much mass to repetitive continuations, because nothing in the loss penalises them. Their fix is an auxiliary training loss rather than a decoding-time patch, and it works better than a penalty does — at the cost of retraining.
The practical implication is that a repetition penalty is a symptom treatment. Well-post-trained instruction models repeat far less, and typically need no penalty at all. Reaching for a large penalty is often evidence about the checkpoint rather than about the decoding.
Interaction with truncation
Penalties are applied to logits before truncation in every mainstream implementation. This ordering has consequences that are easy to miss.
A penalty can push a token below a top-p or min-p threshold, removing it from the candidate set entirely rather than only reducing its probability. Under a tight nucleus, a moderate penalty behaves like a hard ban. This is why penalty and truncation settings cannot be tuned independently, and why a penalty that seems harmless at $p = 0.95$ can be destructive at $p = 0.8$.
Papers
- Keskar et al., CTRL, 2019 — arxiv.org/abs/1909.05858
- Welleck et al., Neural Text Generation with Unlikelihood Training, 2019 — arxiv.org/abs/1908.04319
- Holtzman et al., The Curious Case of Neural Text Degeneration, ICLR 2020 — arxiv.org/abs/1904.09751
- Xu et al., Learning to Break the Loop: Analyzing and Mitigating Repetitions, 2022 — arxiv.org/abs/2206.02369
What to learn next
- Beam search, and why chat models dropped it — the decoding method whose main flaw was exactly this repetition.
- Constrained decoding and grammars — the right tool when the output has a required shape.
- Prompt engineering — often a better cure for repetition than any sampling parameter.