How it works
How ChatGPT works
ChatGPT chops your message into tokens, scores every possible next token, and streams the winners back — shaped by training that taught it which answers people prefer.
- 4 min read
- Updated
On this page 9
- The pipeline at a glance
- Stage 1 — your text becomes tokens
- Stage 2 — tokens become numbers
- Stage 3 — the layers mix meaning together
- Stage 4 — score every possible next token
- Stage 5 — pick one token
- Stage 6 — glue it on and go again
- Where the model's judgement came from
- Which lessons teach each stage
You type a question. Under a second later, an answer starts appearing word by word. This page walks that path end to end: what happens to your text, and why the reply comes back in pieces.
The pipeline at a glance
"Why is the sky blue?"
|
v
+-----------------+ +-----------------+ +---------------------+
| 1. Tokenize | -> | 2. Embed | -> | 3. Transformer |
| chop text | | tokens into | | layers, each |
| into pieces | | numbers | | token looks at |
+-----------------+ +-----------------+ | the others |
+---------------------+
|
v
+-----------------+ +-----------------+ +---------------------+
| 6. Stream back | <- | 5. Sample one | <- | 4. Score every |
| and loop | | token | | possible next |
| until done | | | | token |
+-----------------+ +-----------------+ +---------------------+Stage 1 — your text becomes tokens
The model never sees letters or words. Your message is first cut into tokens — small chunks of text, usually a word or part of one. "unbelievable" might become "un", "bel", "ievable".
Your message is not alone in there. The app quietly adds a system prompt — hidden instructions about tone and behaviour — plus the earlier turns of your chat. All of it must fit inside the context window, the fixed limit on how many tokens the model can read at once. Long chats hit that wall, and the oldest turns fall out. That is why the model can "forget" the start of a long conversation.
Stage 2 — tokens become numbers
Each token is swapped for an embedding: a long list of numbers that acts like coordinates for meaning. Tokens with similar meaning sit near each other. Nobody typed those positions in; the model learned them during training.
Stage 3 — the layers mix meaning together
The numbers pass through a stack of transformer layers, often dozens of them. In every layer, each token looks back at the earlier tokens and pulls in what it needs. That looking-back step is called attention, and it is the heart of the whole machine.
You do this too. When you read "he" in a message, your eyes flick back to find the name. Attention is that flick, done with arithmetic, at every token, in every layer.
Stage 4 — score every possible next token
After the last layer, the model produces one score for every token it knows — its whole vocabulary, tens of thousands of options. A high score means "this token is a likely next piece". At this point the model has not chosen anything. It has only ranked the choices.
Stage 5 — pick one token
A setting called temperature decides how the pick happens. Low temperature takes the top choice almost every time, which reads as careful and repetitive. Higher temperature lets lower-ranked tokens win sometimes, which reads as creative. One token is picked. That is the entire output of one pass: a single token.
Stage 6 — glue it on and go again
The chosen token is appended to the text, and the whole pipeline runs again on the longer sequence. One token per loop, until the model emits a special stop token. The app streams each token to your screen as it lands — which is why the answer appears in pieces rather than all at once.
Where the model's judgement came from
The loop above explains the mechanics, not the manners. Three rounds of training produced those.
First, pretraining: the model reads a huge slice of the internet and learns one skill — predict the next token. This is where its knowledge of language and facts comes from.
Second, instruction tuning: the model is shown many examples of questions paired with good answers, written by people. It learns to answer rather than to ramble on.
Third, RLHF — reinforcement learning from human feedback. People compare pairs of model answers and mark which one they prefer. A second model learns to predict those preferences, then scores the main model's practice answers. The main model is nudged toward answers that score well. This is why ChatGPT is polite, refuses some requests, and writes in that recognisable style.
One honest caveat: nothing in this pipeline checks facts. The model ranks plausible continuations. When plausible and true disagree, you get a confident wrong answer — a hallucination.
Which lessons teach each stage
- Stages 1–2, tokens and embeddings: Tokenization and Embeddings
- Stage 3, attention and layers: Attention and Transformers
- Stages 4–6, the full loop: How LLMs actually work
- The training rounds: What is an LLM? and RLHF