The BERT Family

T5 and the text-to-text idea

T5 treats every NLP task as the same problem, text in and text out, so one model and one training method can translate, summarise and classify, only by changing the instruction written into the input.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

T5 treats every language task the same way: text goes in, text comes out. Translation, summarisation and grammar checking are all different instructions typed into the same input format.

Think about a single multi-purpose kitchen appliance — the kind that grinds, blends and juices. It uses the same motor and the same slot you drop ingredients into. You do not need a separate machine for each job. You choose the attachment, drop food in the same slot, and get a different result depending on what you asked for.

T5 works the same way for language tasks. Translation, summarisation, grammar checking — instead of a separate model for each, T5 does all of them with one model and one training method. Put text in, with a short instruction attached, and get text back out.

Why it exists

Before T5, different NLP tasks tended to need genuinely different model setups. A translation model expected source and target language pairs. A classification model expected a text and a fixed set of output labels. A summarisation model expected long input, short output. Each shape of input and output usually meant a differently-structured model.

T5's authors asked a simple question. What if every one of these tasks could be reframed as "text in, text out"? A translation task becomes "translate English to German: [sentence]" in, and the German sentence out. A classification task becomes "cola sentence: [sentence]" in, and the word "acceptable" or "unacceptable" out — text, not a probability over fixed classes. Reframed this way, one architecture and one loss function can handle all of them.

How it works

  Task 1:  IN:  "translate English to German: The weather is nice today."
           OUT: "Das Wetter ist heute schön."

  Task 2:  IN:  "summarize: The Chennai Metro Rail network has expanded
                 rapidly since 2015, adding new corridors..."
           OUT: "the network has expanded rapidly since 2015..."

  Task 3:  IN:  "cola sentence: The bus have arrived."
           OUT: "acceptable"    (or "unacceptable")

  Same model. Same input format: [task instruction] + [text].
  Same output format: plain text, generated one word at a time.

Nothing about the model's architecture changes between these three tasks. Only the words typed into the input change. The model reads that instruction the same way it reads everything else: as text.

Where you have already seen it

  • Modern instruction-following chatbots inherit this exact idea directly. Typing an instruction in plain English and getting a task-appropriate response back is the text-to-text idea, generalised further and scaled up massively.
  • Multi-task NLP products, where one deployed model handles translation, summarisation and rewriting requests without switching between separate specialised systems.
  • Grammar and style checkers that phrase their internal task as a text-to-text rewrite: broken sentence in, corrected sentence out.

Remember this

  • T5 reframes every NLP task as text in, text out, using a short instruction prefix to say which task is meant.
  • One architecture, one training method and one loss function handle every task this way.
  • This "everything is text" idea is a direct conceptual ancestor of how modern instruction-following chatbots work.

What to learn next

Developer — Code and libraries.

T5-small runs comfortably on CPU. It downloads once (about 240 MB) and is cached for later runs.

Setup

bash
pip install transformers torch

Minimal runnable code

t5_multitask.py
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tok = AutoTokenizer.from_pretrained("t5-small")
model = AutoModelForSeq2SeqLM.from_pretrained("t5-small")

tasks = [
    "translate English to German: The weather is nice today.",
    "summarize: The Chennai Metro Rail network has expanded rapidly since 2015, "
    "adding new corridors that connect the airport, the IT corridor and the old city, "
    "cutting commute times for hundreds of thousands of daily riders.",
    "cola sentence: The bus have arrived.",
]

for task in tasks:
    ids = tok(task, return_tensors="pt").input_ids
    out = model.generate(ids, max_new_tokens=30)
    print(f"IN:  {task}")
    print(f"OUT: {tok.decode(out[0], skip_special_tokens=True)}")
    print()
Output
IN:  translate English to German: The weather is nice today.
OUT: Das Wetter ist heute schön.

IN:  summarize: The Chennai Metro Rail network has expanded rapidly since 2015, adding new corridors that connect the airport, the IT corridor and the old city, cutting commute times for hundreds of thousands of daily riders.
OUT: the network has expanded rapidly since 2015. the new corridors connect the airport, the IT corridor and the old city.

IN:  cola sentence: The bus have arrived.
OUT: acceptable

Three completely different tasks, one model, one generate() call each — no task-specific output layer, no switching models.

Line by line

The task prefix — "translate English to German: ", "summarize: ", "cola sentence: " — is not special syntax the model parses. It is plain text T5 was trained to recognise as a pattern during pretraining and fine-tuning. Change the wording and the model may not respond the same way — the exact prefixes shown here match what t5-small was actually trained on.

AutoModelForSeq2SeqLM, not AutoModelForCausalLM or AutoModelForMaskedLM, is the class for encoder-decoder generation — see Encoder, decoder and encoder-decoder models for why the class differs by architecture family.

tok.decode(out[0], skip_special_tokens=True) strips T5's internal special tokens (like its end-of-sequence marker) from the output text, leaving only the readable result.

Reading the CoLA result honestly

The third example asked T5 to judge whether "The bus have arrived." is grammatically acceptable — CoLA is the Corpus of Linguistic Acceptability, a standard benchmark task. "The bus have arrived" has a subject-verb agreement error ("bus" is singular, "have" is plural), so the grammatically correct judgement is "unacceptable". T5-small said "acceptable" — it got this one wrong.

This is worth sitting with rather than glossing over. t5-small is the smallest checkpoint in the T5 family, and grammatical acceptability judgement is a genuinely hard task for a small model. A model producing fluent, confident-looking output is not the same thing as a model being correct — exactly the same lesson from the ModernBERT lesson's "capital of India" example two lessons back.

Common mistakes

Assuming task prefixes are universal across every text-to-text model. T5's specific prefixes were learned during its own training; a different text-to-text model, or a later T5 checkpoint fine-tuned differently, may expect different wording, or none at all.

Treating generated text output as a reliable classification signal without validation. Because T5 outputs "acceptable" or "unacceptable" as free text rather than a fixed-class probability, a bug in decoding, or the model outputting something entirely unexpected, is a real possibility that a proper classification head's fixed output space would rule out by construction.

Not checking max_new_tokens is generous enough for the task. A summary or translation that gets cut off mid-sentence usually means this limit was set too low for the input, not that the model failed at the task.

Try it yourself

Try the CoLA prefix on a sentence that is actually grammatically correct, such as "cola sentence: The bus has arrived.", and check whether the model's judgement changes. Then try a sentence with a more obvious grammar error, such as "cola sentence: Bus the arrived have.", to see whether a clearer violation gets caught more reliably.

What to learn next

Researcher — Mathematics and papers.

Formal framing

T5 (Raffel et al., 2020) trains a single conditional generation objective across every task:

text
P(y | x) = product over t=1..|y| of P(y_t | y_{<t}, x)
  • x is the input sequence, always formatted as "[task prefix]: [input text]".
  • y is the target sequence — a translation, a summary, a label word, or any other task's expected output, always represented as text.
  • The loss is standard cross-entropy over the decoder's output distribution, identical in form regardless of which task generated the example, since every task shares one output format.

This unifies what would otherwise be several distinct loss functions — sequence classification cross-entropy, span-extraction loss, translation cross-entropy — into a single training signal, which is the paper's central architectural claim: task diversity does not require architectural diversity if every task's input and output can be represented as text.

The Colossal Clean Crawled Corpus (C4) and the empirical study

T5's paper is as much a large-scale empirical study as an architecture proposal. Raffel et al. systematically ablate architecture choice (encoder-decoder versus decoder-only, at matched parameter count), pretraining objective (several span-corruption variants), pretraining dataset (C4 versus alternatives), and model and data scale, reporting downstream performance for each variant on the GLUE, SuperGLUE, SQuAD and summarisation benchmarks. The finding most directly relevant to this section: at matched parameter count and compute budget, an encoder-decoder architecture with a denoising (span-corruption) pretraining objective outperformed decoder-only alternatives on the benchmark suite they tested, though this finding predates the scale and training-data regimes of current frontier LLMs, and later work has shown decoder-only architectures scale at least as well given sufficient data and compute.

Span corruption pretraining

Unlike BERT's single-token masking (see Masked language modelling), T5's pretraining objective corrupts contiguous spans of the input, replacing each corrupted span with a single sentinel token, and trains the decoder to generate the original corrupted spans, each preceded by its sentinel:

text
Input:   "Thank you <X> me to your party <Y> week."
Target:  "<X> for inviting <Y> last <Z>"
  • <X>, <Y>, <Z> are sentinel tokens marking corrupted spans.
  • The decoder only has to generate the missing spans, not reproduce the entire input — a considerably shorter, more efficient target sequence per training example than reconstructing the full input would require.

Relative position bias

T5 uses a learned, bucketed relative position bias added directly to attention scores, rather than absolute position embeddings — the ancestor of the log-bucketing scheme referenced in DeBERTa's researcher block. Distance buckets grow logarithmically for larger offsets, following the same reasoning DeBERTa's relative-position bucketing later reused: exact distance matters less once two tokens are already far apart.

Key references

  • Raffel, C. et al. (2020). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683
  • Warstadt, A., Singh, A. & Bowman, S. (2019). Neural Network Acceptability Judgments. Transactions of the ACL — introduces the CoLA benchmark used in the developer block's example.
  • Raffel, C. et al. release the C4 (Colossal Clean Crawled Corpus) dataset alongside the T5 paper, a heavily filtered Common Crawl snapshot that became a widely reused pretraining corpus well beyond T5 itself.

Current state and open problems

The text-to-text unification T5 demonstrated at the architecture-and-training level is now visible, in a more general and less rigidly-prefixed form, in how instruction-tuned decoder-only LLMs are used: a natural-language instruction plus input, producing a natural-language output, covering an even broader range of tasks than T5's fixed set of prefixes ever targeted. Whether this represents T5's core idea "winning" in a different architectural package, or a genuinely different mechanism (in-context instruction-following learned from a much larger and more diverse instruction-tuning dataset) reaching a superficially similar interface, is a genuinely debated framing rather than a settled one. T5's specific encoder-decoder architecture, meanwhile, remains actively used where its efficiency argument holds — bounded-length, well-defined transformation tasks (translation, summarisation of a known input) — rather than open-ended dialogue, where decoder-only models currently dominate.

What to learn next