Summarisation

Fine-tuning your own summariser

Fine-tuning adjusts a pretrained summariser's weights on your own examples, so it learns your house style instead of a generic one.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Fine-tuning adjusts a pretrained summariser's own weights, using your examples, so it learns your specific style.

Picture hiring a fully trained chef. They already know how to cook. You do not teach them cooking from scratch. You show them a few dishes, plated your restaurant's way, and they adapt.

Fine-tuning a summariser works the same way. The model already knows language and already knows how to compress text. You show it examples in your style, and its internal weights shift toward that style.

Why it exists

A general summariser, like the one used in earlier lessons, writes in whatever style its training data taught it. That might not be your style at all.

A news company might want headlines. A legal team might want summaries that never omit dates. A support team might want a strict "issue, cause, fix" format, every time.

Prompting can push a model toward a format, but the pull is often weak and inconsistent. Fine-tuning changes the model itself, making the desired style its new default, not only a request it might follow.

How it works

Pretrained model (knows language broadly)
              |
              v
   Show it examples: (document, your-style summary)
              |
              v
   Adjust weights slightly, so it predicts your summaries better
              |
              v
Fine-tuned model (knows language + your house style)

Each training example nudges the model's internal numbers a small amount. Repeat this over many examples, many times, and the nudges accumulate into a real shift in behaviour.

Where you have already seen it

  • A news app whose summaries always read like headlines. That consistency rarely happens by accident; it is usually a fine-tuned model.
  • A company's internal support-ticket summariser. Trained on that company's own resolved tickets, in that company's own format.
  • Legal-tech tools that summarise contracts in a fixed template. Every summary follows the same structure, because the model was trained on examples in that structure.
  • Voice assistants that summarise your calendar in a consistent tone. That house style was taught, not guessed.

Remember this

  • Fine-tuning adjusts a pretrained model's weights using your own labelled examples.
  • It teaches style and format consistency that prompting alone struggles to guarantee.
  • It needs real examples of the style you want, and more of them than you might expect.

What to learn next

Developer — Code and libraries.

Below, t5-small is trained on only five examples, each pairing a news sentence with a headline-style summary. Five examples is nowhere near enough for production use. It is enough to watch training actually happen, in under a minute on CPU.

Setup

bash
pip install transformers torch

A tiny fine-tuning run

finetune.py
import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tokenizer = AutoTokenizer.from_pretrained("t5-small")
model = AutoModelForSeq2SeqLM.from_pretrained("t5-small")

# House style: always a one-line headline starting with a verb.
examples = [
    ("summarize: The city council approved a new metro line connecting the airport "
     "to the tech park, with construction starting next year.",
     "Approves new airport-to-tech-park metro line."),
    ("summarize: Heavy rain flooded three neighbourhoods overnight, forcing dozens "
     "of families to move to a relief shelter set up by the district office.",
     "Floods force dozens of families into relief shelter."),
    ("summarize: The state exam board pushed back the results date by two weeks "
     "after a printing error was found in one subject's answer key.",
     "Delays exam results by two weeks over printing error."),
    ("summarize: A local startup raised funding to expand its solar-powered "
     "irrigation pumps to five more districts over the next year.",
     "Raises funding to expand solar irrigation to five districts."),
    ("summarize: The transport department launched a mobile app that lets bus "
     "riders track arrival times in real time across the city.",
     "Launches app for real-time bus tracking."),
]

optimizer = torch.optim.AdamW(model.parameters(), lr=5e-5)
model.train()
for step in range(60):
    source, target = examples[step % len(examples)]
    inputs = tokenizer(source, return_tensors="pt", truncation=True)
    labels = tokenizer(target, return_tensors="pt", truncation=True).input_ids
    loss = model(**inputs, labels=labels).loss
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()
    if step % 10 == 0:
        print(f"step {step:2d}  loss {loss.item():.3f}")

model.eval()

def summarize(text):
    inputs = tokenizer(text, return_tensors="pt", truncation=True)
    ids = model.generate(**inputs, max_new_tokens=16, num_beams=4)
    return tokenizer.decode(ids[0], skip_special_tokens=True)

print("\nOn a TRAINING sentence (seen 12 times):")
print(" ", summarize(examples[0][0]))

print("\nOn a NEW sentence the model never saw:")
test_text = ("summarize: The water board announced a two-day supply cut next week "
             "to repair a burst pipe near the old bus depot.")
print(" ", summarize(test_text))
Output
step  0  loss 3.358
step 10  loss 3.061
step 20  loss 2.009
step 30  loss 1.935
step 40  loss 1.474
step 50  loss 1.337

On a TRAINING sentence (seen 12 times):
  The city approves a new metro line connecting the airport to the tech park

On a NEW sentence the model never saw:
  The water board announced a two-day supply cut next week to repair

Loss values will vary a little between runs. Weight updates involve floating-point arithmetic that is not perfectly reproducible across machines. The overall downward trend should hold every time.

Line by line

The loss is falling. From 3.358 down to 1.337 over 60 steps. The model is measurably adjusting toward the training examples, exactly what fine-tuning is supposed to do.

But read the two test outputs carefully. On a sentence it saw 12 times, the model gets closer to the target headline style. It is still not an exact match. On a brand-new sentence, it barely changed at all. It reproduced the source almost verbatim, then got cut off.

This is the honest, important result. Five examples, repeated many times, teach the model to partially memorise those five examples. They do not teach it a transferable style it can apply to new text. That takes far more varied examples, not more repetitions of the same few.

Common mistakes

Watching only the loss curve, and declaring success. A falling loss proves the model is fitting the training data. It does not prove the model generalises to new input, as this lesson's own output shows directly.

Using too few, too similar examples. Five sentences, all news-style, all similar length, give the model very little variety to generalise from. Real fine-tuning datasets usually need hundreds to thousands of diverse examples.

Forgetting model.eval() before generating. Some layers, like dropout, behave differently during training versus inference. Skipping .eval() can quietly produce noisier, less consistent output.

Not holding out a validation set. This toy example tests on a training example and a held-out one side by side, deliberately. Production fine-tuning always needs a separate validation set, never scored during training, to catch this exact memorisation-versus-generalisation gap.

Try it yourself

Add three more examples in the same style, then rerun with range(120) instead of range(60). Watch whether the new-sentence output moves any closer to a proper headline, or stays stuck.

What to learn next

Researcher — Mathematics and papers.

Fine-tuning as continued optimisation

Fine-tuning continues gradient descent on a pretrained model's existing parameters theta_0, using a new, typically much smaller dataset:

text
theta* = argmin over theta of  sum over (x_i, y_i) in D_finetune of  L(f_theta(x_i), y_i)
  • theta_0 is initialised from pretraining, not randomly. This is what separates fine-tuning from training from scratch.
  • L is the same token-level cross-entropy loss used in pretraining.
  • D_finetune is small relative to pretraining data, often by several orders of magnitude.

Why five examples are not enough

The developer block's memorisation-without-generalisation result has a standard explanation. With |D_finetune| = 5, the model has far more parameters than training signal. Nothing in the objective tells apart two very different solutions: learning the general mapping, or memorising these five pairs directly. Both minimise training loss equally well.

Generalisation needs enough diversity in D_finetune. Only then does memorising individual examples stop being a lower-loss shortcut than learning the shared pattern. There is no fixed threshold. It depends on task complexity, and on how different the target style is from the base model's pretrained behaviour.

Catastrophic forgetting

Full fine-tuning updates every parameter, including ones responsible for general language ability learned during pretraining. Aggressive fine-tuning on a narrow dataset can degrade tasks the base model previously handled well. This is called catastrophic forgetting (McCloskey & Cohen, 1989; revisited for transformers by Kirkpatrick et al., 2017).

Mitigations include a low learning rate and few epochs. Mixing a little general-purpose data into the fine-tuning set also helps. It stops the model drifting too far from its pretrained behaviour.

Parameter-efficient alternatives

Full fine-tuning of a large model is expensive: every parameter needs a gradient and an optimiser state in memory. LoRA (Hu et al., 2021) instead freezes the pretrained weights. It learns a small low-rank update alongside them instead. That cuts trainable parameters by orders of magnitude, while keeping most of full fine-tuning's benefit.

For summarisation specifically, LoRA fine-tuning of encoder-decoder models is common practice. The needed style-adaptation signal is typically low-rank: a shift in tone or structure, not a new capability.

Data requirements in practice

Published results on dialogue and news summarisation fine-tuning use training sets in the thousands to tens of thousands of examples. SAMSum (Gliwa et al., 2019) and XSum (Narayan et al., 2018) are two such datasets. Few-shot fine-tuning, with dozens to low hundreds of examples, can still shift style noticeably, particularly with parameter-efficient methods. It rarely matches full-dataset quality on held-out generalisation.

Key references

  • McCloskey, M. & Cohen, N. (1989). Catastrophic Interference in Connectionist Networks. Psychology of Learning and Motivation.
  • Hu, E. et al. (2021). LoRA. arXiv:2106.09685
  • Narayan, S., Cohen, S. & Lapata, M. (2018). XSum. arXiv:1808.08745
  • Gliwa, B. et al. (2019). SAMSum Corpus. arXiv:1911.12237

Current state and open problems

Instruction-tuned LLMs have shifted much of this work toward prompting. Few-shot examples in the prompt itself often replace gradient-based fine-tuning now. This trade-off is discussed directly in prompt an LLM or fine-tune a small model?

The open problem is data efficiency. No reliable method predicts, ahead of time, how many examples a task needs before fine-tuning generalises rather than memorises. This lesson's own experiment sat on the wrong side of that line.

What to learn next