How Text Is Generated

Repetition and frequency penalties

Three different ways to make a model stop repeating itself, two of which work differently from how almost everyone assumes, and one of which quietly punishes your own prompt.

On this page 7
  1. Why models repeat at all
  2. The three penalties, and how they differ
  3. How it works
  4. The part everybody gets wrong
  5. Where you have already seen this
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A repetition penalty lowers the score of words the model has already used. That makes it less likely to say them again.

Think of a friend who tells the same story at every family gathering. The first time it is charming. By the fourth telling, everyone winces.

Now imagine charging a small fine for each retelling. Nothing is banned. The story is still allowed. It has become slightly less attractive to tell, and other stories get a turn.

That is exactly what these penalties do. They do not forbid a word. They tax it.

Why models repeat at all

A language model writes by looking at what it has already written. So its own output becomes the strongest hint about what comes next.

Write "the office is closed on Sunday" and that sentence is now the most recent, most relevant text in view. The most fitting next words often look like the ones already there.

That creates a loop. Say something, be nudged toward saying it again, say it again, be nudged harder. Left alone, models fall into these loops and stay there.

You have seen this. A chatbot repeating a phrase forever, or a model that ends every paragraph with the same closing sentence.

The three penalties, and how they differ

There are three in common use and they are not the same thing.

Repetition penalty divides the word's score, so it scales with how strong the score already was. Used or unused is all it checks — using a word five times costs the same as using it once.

Frequency penalty subtracts a fixed amount for each use. Say a word five times and you pay five times.

Presence penalty subtracts a fixed amount once, the first time a word is used at all. It pushes for new vocabulary rather than against repeated vocabulary.

How it works

   the model's scores after "the cat the the"

   word    score    times used     after a repetition penalty
   the      2.00        3                 1.67   (taxed)
   cat      1.40        1                 1.17   (taxed)
   sat      0.90        0                 0.90   (untouched)
   on       0.30        0                 0.30   (untouched)

"the" and "cat" both drop. "sat" and "on" do not move, so they gain ground without being helped.

Notice "the" was used three times and "cat" once, and both were taxed the same way. That is the repetition penalty's blind spot, and it is what the frequency penalty fixes.

The part everybody gets wrong

These penalties punish every word the model has seen, and that includes your prompt.

Paste in a document and ask for a summary. Every important word in that document has now been used. A strong penalty makes the model avoid those exact words in the summary, so it reaches for vaguer alternatives.

The summary gets worse and nobody can see why. Good libraries let you exclude the prompt from the count. Many do not.

Where you have already seen this

  • A chatbot stuck repeating one sentence until you stop it.
  • API settings named frequency_penalty and presence_penalty on a slider from zero to two.
  • A local model tool with a repeat_penalty box, usually set near one point one.
  • A summary that carefully avoids using the document's own words, and reads oddly for it.

Remember this

  • Penalties tax words the model has already used, they do not ban them.
  • Repetition penalty scales the score; frequency counts uses; presence fires once.
  • All three usually punish your prompt too, which quietly damages summaries.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Three penalties, one table, and the sign trap that makes the first one behave unlike anything you would guess.

All three side by side

penalties.py
import numpy as np

VOCAB = ["the", "cat", "sat", "on", "mat", "and", "very", "happy"]


def repetition_penalty(logits, seen, penalty):
    """CTRL-style. Divide positive logits, MULTIPLY negative ones."""
    out = logits.copy()
    for t in set(seen):
        out[t] = out[t] / penalty if out[t] > 0 else out[t] * penalty
    return out


def frequency_presence(logits, seen, freq=0.0, presence=0.0):
    """OpenAI-style. Subtract a fixed amount, counted or not counted."""
    out = logits.copy()
    counts = {t: seen.count(t) for t in set(seen)}
    for t, n in counts.items():
        out[t] -= freq * n + presence
    return out


logits = np.array([2.0, 1.4, 0.9, 0.3, -0.5, -1.2, -2.0, -3.0])
seen = [0, 1, 0, 0]                       # "the cat the the"

print("token      logit   rep 1.2   freq 0.5   presence 0.5     seen")
after_rep = repetition_penalty(logits, seen, 1.2)
after_freq = frequency_presence(logits, seen, freq=0.5)
after_pres = frequency_presence(logits, seen, presence=0.5)
for i, w in enumerate(VOCAB):
    n = seen.count(i)
    print(f"{w:<10}{logits[i]:>6.2f}{after_rep[i]:>10.3f}{after_freq[i]:>11.2f}"
          f"{after_pres[i]:>15.2f}{n:>9}")

print("\nThe trap, spelled out:")
print(f"  'mat'  started at {logits[4]:>5.2f}. It was never used, so nothing happens.")
neg = repetition_penalty(np.array([-0.5]), [0], 1.2)[0]
print(f"  a token at -0.50 that HAS been used becomes {neg:.2f} - pushed further down.")
pos = repetition_penalty(np.array([2.0]), [0], 1.2)[0]
print(f"  a token at +2.00 that HAS been used becomes {pos:.2f} - also pushed down.")
print("  Same penalty, opposite arithmetic, because a plain divide would RAISE a")
print("  negative logit. That is the whole reason for the sign test.")

print("\nWhat a penalty of 1.2 costs in probability terms:")


def softmax(z):
    e = np.exp(z - z.max())
    return e / e.sum()


for p in (1.0, 1.05, 1.1, 1.2, 1.5, 2.0):
    q = softmax(repetition_penalty(logits, seen, p))
    print(f"  penalty {p:<5} P(the)={q[0]:.3f}  P(cat)={q[1]:.3f}  P(sat)={q[2]:.3f}")
Output
token      logit   rep 1.2   freq 0.5   presence 0.5     seen
the         2.00     1.667       0.50           1.50        3
cat         1.40     1.167       0.90           0.90        1
sat         0.90     0.900       0.90           0.90        0
on          0.30     0.300       0.30           0.30        0
mat        -0.50    -0.500      -0.50          -0.50        0
and        -1.20    -1.200      -1.20          -1.20        0
very       -2.00    -2.000      -2.00          -2.00        0
happy      -3.00    -3.000      -3.00          -3.00        0

The trap, spelled out:
  'mat'  started at -0.50. It was never used, so nothing happens.
  a token at -0.50 that HAS been used becomes -0.60 - pushed further down.
  a token at +2.00 that HAS been used becomes 1.67 - also pushed down.
  Same penalty, opposite arithmetic, because a plain divide would RAISE a
  negative logit. That is the whole reason for the sign test.

What a penalty of 1.2 costs in probability terms:
  penalty 1.0   P(the)=0.452  P(cat)=0.248  P(sat)=0.150
  penalty 1.05  P(the)=0.436  P(cat)=0.246  P(sat)=0.160
  penalty 1.1   P(the)=0.421  P(cat)=0.244  P(sat)=0.168
  penalty 1.2   P(the)=0.395  P(cat)=0.240  P(sat)=0.183
  penalty 1.5   P(the)=0.338  P(cat)=0.226  P(sat)=0.219
  penalty 2.0   P(the)=0.282  P(cat)=0.209  P(sat)=0.255

Reading the output

Look at the the row across the three columns. Repetition penalty took it from 2.00 to 1.667, a cut of about seventeen percent of its value. Frequency penalty took it to 0.50, because it was used three times and each use costs 0.5. Presence penalty took it to 1.50, one flat charge. Three penalties, three completely different results on the same token.

cat and the are treated identically by the repetition penalty. One use versus three, same rule. If your complaint is a word appearing over and over, the frequency penalty targets that directly and the repetition penalty does not.

The sign test is not a style choice. A logit of -0.50 divided by 1.2 gives -0.42, which is higher. A plain divide would reward exactly the tokens you meant to punish. HuggingFace, vLLM and llama.cpp all implement score / penalty if score > 0 else score * penalty for this reason. Every reimplementation that skips the branch is broken, and the breakage is invisible until the model starts looping on rare tokens.

A penalty of 1.2 moved P(the) from 0.452 to 0.395. That is a modest nudge, not a ban. To actually break a hard loop you need something stronger, which is why no_repeat_ngram_size exists. And notice P(sat) climbing from 0.150 to 0.255 as the penalty rises — the untouched tokens gain because the survivors are renormalised.

Which one to reach for

SymptomReach for
The same phrase repeated verbatim, foreverno_repeat_ngram_size — a hard block, not a tax
One word appearing far too oftenfrequency penalty, 0.1 to 0.5
Vocabulary feels narrow and sameypresence penalty, 0.1 to 0.5
General mild drift toward repetitionrepetition penalty, 1.05 to 1.15
Summaries avoiding the source's own wordsturn the penalty down, or exclude the prompt

no_repeat_ngram_size=3 forbids any three-token sequence from appearing twice. It ends loops absolutely, and it also forbids legitimate repeats — a person's name, a technical term, for i in range. On code it is actively harmful.

In the real libraries

python
# HuggingFace: penalises the prompt too, unless you tell it not to
out = model.generate(**inputs, do_sample=True, repetition_penalty=1.1,
                     no_repeat_ngram_size=0)

# vLLM: all three available, and its repetition_penalty excludes nothing by default
from vllm import SamplingParams
params = SamplingParams(repetition_penalty=1.1, frequency_penalty=0.2,
                        presence_penalty=0.0)

transformers 5.x exposes a prompt_ignore_length on RepetitionPenaltyLogitsProcessor so the prompt can be excluded. The plain repetition_penalty= argument to generate does not set it. If prompt contamination matters to you, build the processor yourself and pass it in logits_processor=.

Common mistakes

Setting repetition_penalty=2.0 because 1.1 did not work. Above roughly 1.2 the text degrades fast, in a characteristic way: the model starts using unusual synonyms and its grammar frays. If 1.15 has not fixed the problem, the problem is something else.

Applying penalties to structured output. JSON needs ", {, : and , repeatedly. Penalising them produces malformed output. Use constrained decoding instead, and set every penalty to zero.

Penalising chat and code. Code repeats identifiers constantly. Chat repeats the user's own words on purpose. Both are cases where the default should be no penalty at all.

Mixing repetition and frequency penalties without checking. They stack multiplicatively and additively at the same time, which is difficult to reason about. Pick one.

Assuming a penalty fixes a looping model. Persistent loops usually mean the sampling temperature is near zero, or the prompt is degenerate, or the context has been truncated mid-sentence. A penalty masks that, and masking it is worse than seeing it.

Try it yourself

Set seen = [0] * 20 and print the repetition-penalty column. It is unchanged from seen = [0], because the rule uses set(seen). Then do the same with the frequency penalty and watch it collapse. That contrast is the single most useful thing to know about these two parameters.

What to learn next

Researcher — Mathematics and papers.

Repetition penalty

Introduced in Keskar et al., CTRL: A Conditional Transformer Language Model for Controllable Generation, 2019 (arxiv.org/abs/1909.05858), section 4.1. For logit $z_i$ and the set $G$ of previously generated token ids:

$$ z_i' = \begin{cases} z_i / \theta & i \in G,\ z_i > 0 \ z_i \cdot \theta & i \in G,\ z_i \le 0 \ z_i & i \notin G \end{cases} $$

with $\theta \ge 1$. CTRL reported $\theta \approx 1.2$ as a reasonable balance between truthful continuation and lack of repetition.

The transformation is not shift-invariant, and that is its defining pathology. Logits are only defined up to an additive constant. A model whose logits are all shifted by $+5$ is the same model, and this penalty produces a different result on it. Two checkpoints that are functionally identical can require different $\theta$. It is worth understanding that this parameter is not measuring what its name suggests.

Frequency and presence penalties

From the OpenAI API, and additive rather than multiplicative:

$$ z_i' = z_i - \alpha_{\text{freq}} \cdot c_i - \alpha_{\text{pres}} \cdot \mathbb{1}[c_i > 0] $$

where $c_i$ is the number of times token $i$ has appeared. Both parameters range over $[-2, 2]$ in that API, and negative values encourage repetition, which is occasionally useful for forced formats.

These are shift-invariant, since subtracting a constant from one logit changes the softmax in a well-defined way regardless of the overall offset. They are the better-behaved formulation, and the fact that the older, worse-behaved one is more widely deployed is an accident of history.

Hard constraints

no_repeat_ngram_size = n sets $z_i = -\infty$ for any token that would complete an $n$-gram already present in the sequence. It is usually credited to Paulus et al., 2017 in abstractive summarisation; the implementation everyone copies is Fairseq's, which is what the HuggingFace processor cites.

This is a blunt instrument with a real cost: it makes some correct outputs unreachable. On a document mentioning "Reserve Bank of India" twice, $n = 3$ makes the second mention impossible to produce. Any text with legitimate repeated phrasing — legal, medical, code — is damaged by it in a way that no penalty tuning recovers.

Why repetition happens

Two accounts, and they are not exclusive.

Self-reinforcement. Xu et al., 2022 (arxiv.org/abs/2206.02369) measure that the more times a sentence already appears in the context, the higher the probability of generating it again. Sentences that start with a higher probability show a stronger version of the effect. The loop is an attractor, not a coincidence.

Likelihood as the wrong objective. Welleck et al., Neural Text Degeneration with Unlikelihood Training, 2019 (arxiv.org/abs/1908.04319) blame the training objective. Maximum-likelihood training assigns too much mass to repetitive continuations, because nothing in the loss penalises them. Their fix is an auxiliary training loss rather than a decoding-time patch, and it works better than a penalty does — at the cost of retraining.

The practical implication is that a repetition penalty is a symptom treatment. Well-post-trained instruction models repeat far less, and typically need no penalty at all. Reaching for a large penalty is often evidence about the checkpoint rather than about the decoding.

Interaction with truncation

Penalties are applied to logits before truncation in every mainstream implementation. This ordering has consequences that are easy to miss.

A penalty can push a token below a top-p or min-p threshold, removing it from the candidate set entirely rather than only reducing its probability. Under a tight nucleus, a moderate penalty behaves like a hard ban. This is why penalty and truncation settings cannot be tuned independently, and why a penalty that seems harmless at $p = 0.95$ can be destructive at $p = 0.8$.

Papers

What to learn next