Summarisation

Controlling summary length

A summariser does not know how long a summary should be unless you tell it, using generation settings like max_length and min_length.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A summariser has no built-in sense of "how long a summary should be" unless you tell it.

Picture asking a colleague to summarise a report. "Give me the two-line version" and "walk me through it properly" get very different answers, from the exact same report.

A summarisation model needs the same instruction. Left alone, it picks a length based on patterns in its training data. That may not match what you need right now.

Why it exists

Different situations need different lengths. A push notification needs one line. A briefing document needs a paragraph. An executive summary needs half a page.

Without a length control, you get whatever length the model feels like producing. That might be too short to be useful, or too long to read on a phone screen.

Length controls give you a dial. Turn it toward "short" for a headline. Turn it toward "long" for real detail. Same document, same model, different outcome.

How it works

Same article, three settings:

max_length=15   ->  "Central bank held rates steady."
max_length=35   ->  "Central bank held rates steady, citing cooling
                     inflation. Growth forecast raised."
max_length=60   ->  "Central bank held rates steady, citing cooling
                     inflation but persistent food-price risk. Growth
                     forecast raised. Markets reacted calmly."

The model is not writing three unrelated summaries. It is compressing the same understanding of the article into progressively smaller boxes.

Where you have already seen it

  • Push notification previews vs. the full article summary in a news app. Same underlying model, different length setting.
  • "Show more" buttons on AI-generated summaries. Often a second, longer generation, not only text hidden with CSS.
  • Email clients that show a one-line preview, then a longer AI summary on open. Two lengths, one summarising system.
  • Voice assistants giving a short spoken answer, with a longer written version available. Length tuned to the medium.

Remember this

  • A summariser will not guess your ideal length; you have to specify it.
  • Length controls trade detail for brevity — there is no setting that gives you both.
  • The same model can serve a headline and a full briefing, with only the length setting changed.

What to learn next

Developer — Code and libraries.

max_length and min_length are the two settings that matter most. Below, the same article is summarised three times, only changing those two numbers.

Setup

bash
pip install transformers torch

Three lengths, one article

length_control.py
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tokenizer = AutoTokenizer.from_pretrained("t5-small")
model = AutoModelForSeq2SeqLM.from_pretrained("t5-small")

text = (
    "summarize: The Reserve Bank of India kept its main interest rate unchanged "
    "at its meeting on Wednesday, matching what most economists had expected. "
    "The central bank said inflation had cooled in recent months but warned that "
    "food prices remained a risk. It also raised its growth forecast for the "
    "current financial year, citing stronger consumer spending and factory output. "
    "Markets reacted calmly to the announcement, with the rupee holding steady "
    "against the dollar."
)

inputs = tokenizer(text, return_tensors="pt", truncation=True)

for label, max_len, min_len in [("short", 15, 5), ("medium", 35, 15), ("long", 60, 40)]:
    ids = model.generate(**inputs, max_length=max_len, min_length=min_len, num_beams=4)
    summary = tokenizer.decode(ids[0], skip_special_tokens=True)
    print(f"{label:6s} (max_length={max_len:2d}): {summary}")
Output
short  (max_length=15): the central bank said inflation had cooled in recent months. but
medium (max_length=35): the central bank said inflation had cooled in recent months but warned food prices remained a risk. it also raised its growth forecast for the current financial year 
long   (max_length=60): the central bank said inflation had cooled in recent months but warned food prices remained a risk. it also raised its growth forecast for the current financial year. markets reacted calmly to the announcement, with the rupee holding steady against the dollar.

Line by line

max_length and min_length count tokens, not words or characters. A max_length of 15 does not mean 15 words. Most English words are one or two tokens, so 15 tokens lands closer to 10-12 words.

Look closely at the "short" output. It ends mid-sentence: "...cooled in recent months. but". The model hit its 15-token ceiling mid-thought and generation stopped there, with an awkward dangling word.

This is a real, common failure, not a one-off bug. A hard max_length cutoff does not know or care about sentence boundaries. It stops counting tokens and halts, wherever that lands.

min_length forces the model to keep going. Without it, a model can produce a very short, unhelpfully vague summary. min_length blocks the end-of-sequence token from firing too early, at the cost of sometimes padding with lower-value content.

Common mistakes

Treating max_length as a word count. As shown above, it is a token count. For a target word count, estimate roughly words * 1.3 tokens for English, then verify against the actual tokenizer.

Setting max_length too tight and shipping the broken output anyway. The cut-off "short" summary above is genuinely unusable as-is. Either raise the limit, or post-process to trim back to the last full sentence.

Confusing max_length with max_new_tokens. max_length in generate() can count the input length too, depending on the model and settings. max_new_tokens counts only newly generated tokens, and is usually the clearer, safer choice.

Assuming a longer summary is always more faithful. A longer summary has more room to state something the source never said, not less. Length and faithfulness are separate concerns; see faithfulness in summarisation.

Try it yourself

Fix the "short" summary's mid-sentence cutoff. Split the output on ".", drop the final fragment if it lacks a full stop, then rejoin. Compare readability before and after.

What to learn next

Researcher — Mathematics and papers.

Length as a constraint on generation, not an input feature

Autoregressive generation, defined in abstractive summarisation, produces one token at a time, each conditioned on the sequence so far. Length control is applied at the decoding layer, not as part of the model's learned distribution over y directly.

max_length caps total sequence length. Generation is truncated once this many tokens have been emitted, whether or not the end-of-sequence token has fired. min_length masks the end-of-sequence token's logit to negative infinity until that many tokens have been produced, forcing continuation.

The length-normalisation interaction

Beam search, covered in beam search, and why chat models dropped it, scores a hypothesis by summed log-probability. Unnormalised, this structurally favours shorter sequences: every additional token multiplies the sequence probability by a value under one.

length_penalty in generate() corrects this by dividing the score by length^alpha before comparing beams:

text
score(y) = ( sum over t of log P(y_t | y_<t, x) ) / length(y)^alpha
  • alpha > 1 favours longer sequences more strongly.
  • alpha < 1 favours shorter ones.
  • alpha = 1 is standard length normalisation, dividing by length directly.

This interacts with min_length in a way worth naming precisely. min_length is a hard floor enforced during decoding. length_penalty only reshapes preference among already-valid candidate lengths. They solve different problems and are usually set together.

Length as a learned, controllable signal

Rather than only constraining decoding after the fact, some systems train length control directly into the model. Fan, Grangier & Auli (2018) prepend a length token to the input. This gives the model a length target during training. It then learns to plan a summary that reaches roughly that length on its own. It is not cut off arbitrarily later.

This produces noticeably more coherent short summaries than post-hoc truncation. The model learned to compress its content to fit the target, instead of stopping mid-sentence when a counter runs out.

Evaluation-side effects of length

ROUGE is sensitive to length in a way worth naming. Recall-oriented ROUGE variants reward longer summaries almost mechanically, since a longer summary has more chances to overlap with the reference. Comparing ROUGE scores across systems with different output lengths, without controlling for length, is a common evaluation error.

Key references

  • Fan, A., Grangier, D. & Auli, M. (2018). Controllable Abstractive Summarization. arXiv:1711.05217
  • Kryscinski, W. et al. (2019). Neural Text Summarization: A Critical Evaluation. arXiv:1908.08960 — includes a detailed analysis of length-versus-quality trade-offs.

Current state and open problems

LLMs prompted with an explicit length instruction, such as "in two sentences," follow it inconsistently. Very short or very long targets fare worst. This remains an open weakness even in frontier models. Instruction-following for length is measurably worse than instruction-following for content or tone.

No widely deployed system guarantees exact length. None also guarantee the sentence-boundary coherence that trained-in length control (Fan et al., 2018) achieves. Production systems mostly pick one imperfect strategy. Either a hard token cutoff with cleanup, or a soft length instruction with no guarantee at all.

What to learn next