Encoder, decoder and encoder-decoder models
Whether a transformer is an encoder, a decoder, or both decides what job it can do at all — BERT cannot write a sentence, GPT cannot look ahead, and T5 was built to do both.
- 11 min read
- 3 reading levels
- Published
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An encoder reads a whole passage at once to understand it. A decoder writes text one word at a time, only ever looking backward. Some models do both, one after the other.
Think about the difference between reading a finished letter and writing a reply. Reading the letter, your eyes can jump anywhere. You can glance back at the first line while reading the last one, and back again. Writing your reply is different. You commit to each word in order. You cannot un-write a word you have already sent, or peek ahead at a sentence you have not written yet.
BERT is a reader. It looks at a whole finished sentence at once, in both directions, to understand it. GPT is a writer. It produces one word, then the next, committed to order, never allowed to see a future word it has not generated yet. This single difference decides what job each model can do.
Why it exists
Reading and writing are genuinely different jobs. Building a transformer around one of them constrains what it is allowed to do.
An encoder like BERT lets every word attend to every other word, forward and backward, freely. Full context is available, because the whole input is already there before the model starts working. This makes it excellent at understanding a piece of text. But it has no mechanism for producing new text one word after another, because nothing in its design ever commits to an order.
A decoder like GPT enforces a rule called causal masking. Word i can only look at words 1 through i, never anything after it. This is what makes generation possible — the model always knows exactly what it has "said" so far, and produces the next word consistently with that. But it means a decoder can never look ahead the way an encoder does, even during training.
Some tasks genuinely need both: reading one piece of text fully, then producing a different piece of text word by word. Translation and summarisation are like this. That is what an encoder-decoder model is built for, and the next lesson covers it, T5, in detail.
How it works
ENCODER (BERT):
Input: The chai was hot .
Every word sees every other word, both directions:
"chai" can look at "hot", and "hot" can look at "chai" -- freely.
DECODER (GPT):
Input so far: The chai was
Generating: -> "hot"
"was" can look at "The" and "chai", never at "hot" -- it hasn't
generated "hot" yet, and is not allowed to peek ahead.
ENCODER-DECODER (T5):
Encoder reads the FULL input first, both directions, same as BERT.
Decoder then generates output one word at a time, same as GPT,
while also being allowed to look back at everything the encoder read.Ask BERT to write a sentence and it cannot. It has no generation mechanism at all. Ask GPT to deeply cross-reference the end of a passage while processing the beginning, and it cannot, by design, look ahead. Each architecture is genuinely suited to a different job.
Where you have already seen it
- Search engines use encoder models to understand your query and rank documents against it — no generation needed, only understanding.
- Chatbots use decoder models to write their replies one token at a time, the causal, one-direction-only process described above.
- Translation and summarisation tools often use encoder-decoder models. An encoder reads the full source text, then a decoder generates the target text — the next lesson covers exactly this.
Remember this
- An encoder (BERT) reads a whole passage at once, both directions — excellent for understanding, unable to generate text.
- A decoder (GPT) writes one word at a time, only ever looking backward — this restriction is exactly what makes generation possible.
- An encoder-decoder model (T5) combines both: read everything, then generate a response word by word.
What to learn next
- T5 and the text-to-text idea — the encoder-decoder model this lesson sets up.
- BERT — the encoder model whose limitations motivated this comparison.
- Attention — the causal masking mechanism that makes decoders fundamentally different from encoders.
Developer — Code and libraries.
The architectural difference between encoder, decoder and encoder-decoder models is visible directly in transformers, without downloading full model weights — a config file alone reveals which category a model belongs to.
Setup
pip install transformers torchMinimal runnable code: reading the architecture from config alone
from transformers import AutoConfig
for name in ["bert-base-uncased", "gpt2", "t5-small"]:
cfg = AutoConfig.from_pretrained(name)
print(f"{name:20} is_decoder={getattr(cfg,'is_decoder',None)!s:6} "
f"is_encoder_decoder={cfg.is_encoder_decoder!s:6} arch={cfg.architectures}")bert-base-uncased is_decoder=False is_encoder_decoder=False arch=['BertForMaskedLM'] gpt2 is_decoder=None is_encoder_decoder=False arch=['GPT2LMHeadModel'] t5-small is_decoder=False is_encoder_decoder=True arch=['T5ForConditionalGeneration']
BERT: neither flag set, an encoder. GPT-2: is_encoder_decoder=False and no is_decoder field at all — its config never needed the flag, since GPT-2 is inherently decoder-only, nothing else to distinguish it from. T5: is_encoder_decoder=True, confirmed directly from its own config, not inferred.
The difference is not only a flag — it changes what the model can do
from transformers import AutoTokenizer, AutoModelForMaskedLM
tok = AutoTokenizer.from_pretrained("distilbert-base-uncased")
model = AutoModelForMaskedLM.from_pretrained("distilbert-base-uncased")
ids = tok("The best filter coffee in Chennai is", return_tensors="pt").input_ids
try:
model.generate(ids, max_new_tokens=12)
except Exception as e:
print(f"{type(e).__name__}: {e}")AttributeError: 'DistilBertForMaskedLM' object has no attribute 'generate'
This is not a missing feature someone forgot to add. BERT-family models have no .generate() method at all, because generation requires a mechanism — causal masking plus autoregressive decoding — that an encoder architecture does not have.
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("distilgpt2")
model = AutoModelForCausalLM.from_pretrained("distilgpt2")
enc = tok("The best filter coffee in Chennai is", return_tensors="pt")
out = model.generate(**enc, max_new_tokens=12, do_sample=False, pad_token_id=tok.eos_token_id)
print(tok.decode(out[0]))The best filter coffee in Chennai is the Kolkata-based Kolkata-based K
distilgpt2 generates without error, since decoder architectures are built for exactly this. Read the output honestly, though: it is repetitive and factually meaningless — "Kolkata-based Kolkata-based K" is a small, undertrained model falling into a repetition loop under greedy decoding, a well-known failure mode covered in Repetition and frequency penalties. The point of this example is that generation runs, not that a 82-million-parameter distilled model produces good text.
Line by line
getattr(cfg, 'is_decoder', None) handles GPT-2's config not defining the attribute at all, rather than crashing — AutoConfig objects vary in which fields they define, since not every architecture needs every flag.
pad_token_id=tok.eos_token_id avoids a real, common warning: GPT-2 has no dedicated padding token by convention, since it was never trained with batched, padded input the way BERT was — reusing the end-of-sequence token as the pad token is the standard workaround.
do_sample=False requests greedy decoding — always picking the single highest-probability next token — which is what makes this output exactly reproducible across runs, unlike sampling-based generation.
Common mistakes
Trying to call .generate() on a BERT-family model and expecting a bug report to fix it. As demonstrated, this is not a bug — it is a structural property of the architecture. Use a decoder or encoder-decoder model for any task requiring open-ended text generation.
Assuming a bigger encoder would eventually gain generation ability. Scale does not add a missing mechanism. bert-large cannot generate text any more than bert-base can — the limitation is architectural, not a matter of capacity.
Forgetting pad_token_id when generating with GPT-2 and getting confusing warnings. The warnings ("attention mask and pad token id were not set") do not stop generation from running, but they signal a real, silently-handled ambiguity in the input — worth fixing rather than ignoring.
Try it yourself
Run the same AutoConfig check against roberta-base, microsoft/deberta-v3-xsmall, and gpt2-medium. Confirm every BERT-family model in this section reports is_encoder_decoder=False and has no working .generate(), regardless of how the position-handling or attention mechanism inside it differs.
What to learn next
- T5 and the text-to-text idea — building on the
is_encoder_decoder=Truemodel from this lesson. - Repetition and frequency penalties — fixing the repetition loop seen in the GPT-2 output above.
- Decoder-only vs encoder-decoder — a deeper architectural comparison of the two generative designs.
Researcher — Mathematics and papers.
Attention masking as the defining difference
All three architectures share the same core building block — scaled dot-product attention (see Attention's researcher block). What differs is the mask applied to the attention score matrix before softmax:
Encoder self-attention: no mask -- every position attends to every position
Decoder self-attention: causal mask -- position i attends only to positions <= i
Encoder-decoder cross-attention: decoder positions attend to ALL encoder positions,
unmasked, since the full encoder input is already
complete before decoding beginsAn encoder-decoder model like T5 therefore has three distinct attention patterns operating in the same forward pass: unmasked self-attention within the encoder, causally-masked self-attention within the decoder, and unmasked cross-attention from decoder to encoder. This is architecturally more machinery than either a pure encoder or pure decoder needs, which is the direct cost side of an encoder-decoder design's flexibility.
Why encoders cannot generate
Generation requires, at minimum, an autoregressive factorisation of the output distribution:
P(y) = product over t=1..T of P(y_t | y_{<t})y_{<t}is every output token generated before stept.
An encoder trained with a bidirectional objective (masked language modelling — see Masked language modelling) never learns P(y_t | y_{<t}) for arbitrary t; it learns P(x_i | x_{\i}), conditioned on both sides of a gap. These are different training objectives producing models with fundamentally different information available at each position, not only a missing convenience method — this is why AutoModelForMaskedLM.generate() fails structurally in the developer block, rather than being a method the library authors happened not to write yet.
Complexity comparison
| Component | Encoder-only | Decoder-only | Encoder-decoder |
|---|---|---|---|
| Self-attention cost | O(n^2 d), one pass | O(n^2 d) per generated token during naive decoding (amortised via KV cache — see The KV cache) | Encoder: O(n^2 d) once. Decoder: as decoder-only, plus cross-attention |
| Cross-attention cost | none | none | O(n * m * d) per decoder layer, for encoder length n, decoder length m |
| Inference for a fixed-output task (classification) | Single forward pass | Requires generating and parsing an output token, an awkward fit | Requires running the full generation loop even for short structured output |
This table is the formal version of the practical argument made in ModernBERT's researcher block: for tasks with a fixed, already-known output space (classification, extraction), an encoder-only forward pass is strictly cheaper than any generation-based approach, since it needs exactly one forward pass with no autoregressive loop at all.
Key references
- Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762 — the original encoder-decoder transformer, for machine translation.
- Devlin, J. et al. (2019). BERT: Pre-training of Deep Bidirectional Transformers. arXiv:1810.04805
- Radford, A. et al. (2018). Improving Language Understanding by Generative Pre-Training. The original GPT paper, establishing decoder-only pretraining via next-token prediction.
- Raffel, C. et al. (2020). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 — T5.
Current state and open problems
The field's practical centre of gravity has shifted heavily toward decoder-only architectures since roughly 2022 — GPT-family, Llama, and most current frontier LLMs are decoder-only, even for tasks (translation, summarisation) that an encoder-decoder architecture was originally designed for, prompting and instruction-following having proven flexible enough to absorb most of the encoder-decoder use case within a single unified architecture. Encoder-only models have not disappeared, and remain the more compute-efficient choice specifically for classification, retrieval and extraction, as argued in ModernBERT. Whether encoder-decoder models retain a durable niche, or are gradually absorbed entirely into decoder-only designs with sufficiently good prompting, is an open, actively contested question — see Decoder-only vs encoder-decoder for the fuller architectural argument on both sides.
What to learn next
- T5 and the text-to-text idea — the encoder-decoder design examined in full.
- Decoder-only vs encoder-decoder — the deeper architectural trade-off analysis.
- The KV cache — why decoder inference is not actually
O(n^2)per generated token in practice.