AI glossary

Transformer

In one sentence The transformer is the network design behind almost every modern language model — it reads all tokens at once and uses attention to decide which ones matter.

By Updated

A transformer is the neural network design behind nearly every modern language model. It reads all the tokens at once and uses attention to work out which ones matter to which.

Imagine a committee reading a sentence together, where every member is responsible for one word and may consult every other member before deciding what their word means in this context. The person holding "bank" asks around, hears "river" from a colleague, and settles on the water meaning rather than the money one. Everyone does this at the same time, in one round.

That simultaneity is the invention. The models before it, RNNs and LSTMs, read one word at a time and passed a summary forward, which made them slow to train and forgetful across long passages. Transformers process a whole sequence in parallel, which is exactly the shape of work a GPU is fast at — and that is what made training on trillions of tokens possible at all.

The pieces, in the order they run

tokens → embeddings + position information
       → [ attention  →  feed-forward network ]  × N layers
       → probability for every possible next token

Position information has to be added explicitly, because attention on its own has no sense of order — "Ravi called Anil" and "Anil called Ravi" would otherwise look identical. Each layer then alternates between attention (mix information across positions) and a small feed-forward network (process each position on its own), with residual connections and layer normalisation holding the training steady across dozens of layers.

Three variants exist. Encoder-only models like BERT read text for understanding. Decoder-only models, which covers GPT, Llama, Mistral and most of what people mean by LLM today, generate text left to right. Encoder-decoder models such as T5 take one sequence and produce another, which suits translation. The paper that started it is "Attention Is All You Need", Vaswani et al., 2017.

Where to go next

Learn this properly

Full lessons that use this term in context.

  • Deep Learning

    Transformers

    A transformer reads every word at once and lets each word decide which other words matter to it, which is the architecture behind almost every modern AI model.

Back to the glossary