Generative AI

How LLMs actually work

Inside an LLM the text is chopped into pieces, turned into numbers, and mixed together many times so every piece can look at the others.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. The four steps, start to finish
  4. How it works, in one picture
  5. The one rule that makes it work
  6. Why "attention" is the clever bit
  7. Where you have already seen the effects
  8. Honest limits
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

An LLM breaks your text into pieces, turns each piece into numbers, then scores every possible next piece.

The interesting part is the middle step, where the model works out which of your words matter.

The analogy you have already lived

Read a long WhatsApp message from a friend. Halfway down you hit the word "he". Your eyes flick backwards to find the name that "he" refers to.

You did that without deciding to. That backward glance is the most important idea inside an LLM. It has a name: attention. Attention means each word gets to look at the other words. It then takes whatever it needs from them.

Everything else in this lesson is scaffolding around that one move.

The four steps, start to finish

Step one, chopping. Your sentence is cut into tokens. A token is a small chunk of text. It is often a whole word, or a piece of one. "unbelievable" might become "un", "bel", "ievable". This is called tokenization, and it happens before the model sees anything.

Step two, numbers. Each token is swapped for a long list of numbers called an embedding. An embedding is a set of coordinates. It places the token near tokens with similar meaning. "king" lands near "queen". "mango" lands near "banana". Nobody typed those positions in; the model learned them.

Step three, mixing. The lists of numbers pass through a stack of layers — repeated processing steps, often dozens of them. In every layer, each token looks at the tokens before it and pulls in what is useful. After enough layers, the numbers for the last token carry the meaning of the whole sentence.

Step four, scoring. The final numbers are turned into a score for every token the model knows. The highest scoring token is usually chosen, glued on the end, and the whole thing runs again.

How it works, in one picture

"The capital of France is"
        |
        v
 [ chop into tokens ]  ->  "The" "capital" "of" "France" "is"
        |
        v
 [ swap each for numbers ]         (embeddings)
        |
        v
 [ layer 1 ] -> [ layer 2 ] -> ... -> [ layer N ]
        |          each layer lets tokens look at earlier tokens
        v
 [ score every token in the vocabulary ]
        |
        v
   " Paris"  wins  ->  glue it on  ->  run the whole thing again

The one rule that makes it work

A token may look backwards only. Never forwards.

This sounds like a limitation. It is actually the reason training works at all. Because the model never peeks ahead, every position in a training sentence becomes its own practice question. A sentence of a hundred tokens gives a hundred lessons, not one.

This part confuses almost everyone the first time. Read it again slowly. It is worth the second pass.

Why "attention" is the clever bit

Take the sentence: "The cat sat on the mat because it was warm."

What does "it" mean? The mat. Now change one word: "The cat sat on the mat because it was tired." Now "it" means the cat.

Nothing else in the sentence moved. A system matching keywords cannot tell these apart. Attention can. The token "it" looks back at every earlier token. It then decides which one to lean on, for this sentence only.

Where you have already seen the effects

  • Autocomplete that knows you are writing an angry email, not a birthday wish.
  • Translation that gets gendered words right by looking back at the subject.
  • A code assistant that uses the variable name you defined forty lines ago.

Honest limits

  • The model has no scratchpad. All its "thinking" happens in those layers, unless it writes steps out as text.
  • Older tokens can get lost when the conversation grows very long.
  • The chopping step is not neutral. Numbers and non-English scripts are often split badly, which hurts accuracy and raises cost.

Remember this

  • Tokens in, numbers in the middle, scores out. Then repeat.
  • Attention lets each token borrow meaning from earlier tokens.
  • Looking backwards only is a feature, not a restriction.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

That is genuinely all. Attention is small enough to write from scratch, and doing so once will teach you more than reading ten diagrams.

Self-attention in twenty-five lines

We give four tokens hand-made vectors so you can check every number by hand. A real model learns these vectors during training.

attention.py
import numpy as np

tokens = ["the", "cat", "sat", "it"]

# Hand-made vectors so the arithmetic stays checkable.
# "it" shares a dimension with "cat" because, in this sentence, it refers to the cat.
X = np.array([
    [1.0, 0.0, 0.0, 0.0],   # the
    [0.0, 1.0, 0.0, 0.0],   # cat
    [0.0, 0.0, 1.0, 0.0],   # sat
    [0.0, 1.0, 0.0, 1.0],   # it
])

def attend(V):
    d = V.shape[1]
    scores = V @ V.T / np.sqrt(d)          # how well each token matches every other one
    scores = scores + np.triu(np.ones_like(scores), k=1) * -1e9   # no peeking forwards
    exp = np.exp(scores - scores.max(axis=1, keepdims=True))      # stable softmax
    weights = exp / exp.sum(axis=1, keepdims=True)
    return weights, weights @ V

weights, blended = attend(X)

print("        " + "".join(f"{t:>8}" for t in tokens))
for i, t in enumerate(tokens):
    print(f"{t:>8}" + "".join(f"{w:>8.2f}" for w in weights[i]))
print("\nnew vector for 'it':", np.round(blended[3], 2))

Y = X.copy()
Y[3] = [0.0, 0.0, 1.0, 1.0]                # change ONLY the last token
_, blended2 = attend(Y)
print("earlier tokens unchanged:", np.allclose(blended[:3], blended2[:3]))
Output
             the     cat     sat      it
     the    1.00    0.00    0.00    0.00
     cat    0.38    0.62    0.00    0.00
     sat    0.27    0.27    0.45    0.00
      it    0.16    0.26    0.16    0.43

new vector for 'it': [0.16 0.69 0.16 0.43]
earlier tokens unchanged: True

Reading that table

Each row is one token deciding where to look. Each row adds up to one, because the weights are shares of a fixed budget of attention.

Row the is all zeros except itself. It is the first token, so there is nothing behind it to look at.

Row it is the payoff. Look at the second column: it gives 0.26 of its attention to cat, more than it gives to the or sat. That happened purely because their vectors point in a similar direction. Nobody wrote a rule about pronouns.

The last line is the KV cache in miniature. Changing the final token left the first three blended vectors byte-for-byte identical, because those rows have exactly zero weight on position three. That is why a serving system can cache the earlier work and never recompute it as generation proceeds.

Line-by-line, for the parts that trip people up

V @ V.T gives every pairwise match score at once. Row i, column j is how strongly token i matches token j.

/ np.sqrt(d) keeps those scores from growing as the vectors get longer. Without it, large dot products push the softmax into a near-one-hot spike, and gradients vanish. d here is the vector width.

np.triu(..., k=1) * -1e9 writes a huge negative number into everything above the diagonal, which is exactly the "future" positions. After the softmax those become zero. This is the causal mask, meaning the rule that a token may look backwards only.

Subtracting the row maximum before np.exp is standard numerical hygiene. Without it, np.exp of a large score overflows to infinity and the whole row turns into nan.

weights @ V is the actual retrieval. Each token's new vector is a weighted blend of the vectors it chose to look at.

What a real model adds on top

The version above is one attention head with no learned parameters. A production transformer block adds four things:

  1. Learned projections. Instead of using V three times, the model learns three separate matrices producing queries, keys and values. Queries ask, keys advertise, values carry the payload.
  2. Multiple heads. Eight to a hundred copies run in parallel, each free to specialise. One head may track syntax, another may track long-range references.
  3. A feed-forward network. After mixing, each position is processed on its own through a small two-layer network. This is where most parameters live.
  4. Residual connections and normalisation. The block's output is added back to its input, which keeps gradients flowing through very deep stacks.

Stack that block thirty to a hundred times and you have an LLM.

Common mistakes

Softmax over the wrong axis. axis=1 normalises each row, meaning each token's outgoing attention. Using axis=0 normalises columns instead, which quietly produces garbage that still runs without error. Always assert weights.sum(axis=1) is close to one.

Using -np.inf for the mask. It works here, but a row that is entirely masked becomes nan after the softmax. A large finite negative number like -1e9 degrades gracefully. Production code uses the dtype minimum.

Forgetting the mask entirely during training. The model then sees the answer while predicting it, training loss falls beautifully, and the model is worthless at generation time. If your loss looks impossibly good, check the mask first.

Assuming attention weights are explanations. High attention on a token does not prove the model used it to decide. Attention is one of many signals inside a deep stack, and the interpretability literature is clear that reading it as a causal explanation is unsafe.

Try it yourself

Add a fifth token whose vector is [0.0, 0.0, 1.0, 0.0], which is identical to sat. Rerun and look at where it puts its attention. Then set the vector for it to [0.0, 0.0, 1.0, 1.0] and watch its attention move from cat to sat. You have changed what the pronoun refers to by moving numbers.

What to learn next

Researcher — Mathematics and papers.

Scaled dot-product attention

Attention(Q, K, V)  =  softmax( (Q Kᵀ) / √d_k  +  M )  V

Q is the query matrix with shape n × d_k. K is the key matrix, also n × d_k. V is the value matrix with shape n × d_v. n is the sequence length and d_k is the key dimension. M is the causal mask, defined as M_ij = 0 when j ≤ i and −∞ otherwise.

The √d_k divisor is not cosmetic. If the components of Q and K are independent with unit variance, the dot product of two d_k-dimensional vectors has variance d_k. Feeding scores of that magnitude into a softmax saturates it, and the Jacobian of a saturated softmax is close to zero. Dividing by √d_k restores unit variance and keeps gradients alive.

Multi-head attention

head_i    =  Attention( X W_i^Q , X W_i^K , X W_i^V )
MHA(X)    =  Concat( head_1 … head_h ) W^O

X is the input of shape n × d_model. Each projection W_i^Q, W_i^K has shape d_model × d_k and W_i^V has shape d_model × d_v, with d_k = d_v = d_model / h by convention. W^O has shape d_model × d_model. Total parameters stay at 4 · d_model², independent of h.

The motivation is representational, not computational. A single softmax produces one convex combination per position. h heads produce h of them in different learned subspaces, so one head can track subject agreement while another tracks bracket matching.

The rest of the block

x  ←  x + MHA( Norm(x) )          pre-norm residual
x  ←  x + FFN( Norm(x) )

Classic feed-forward network:

FFN(x)  =  W_2 · GELU( W_1 x + b_1 )  +  b_2

W_1 has shape d_ff × d_model with d_ff = 4 · d_model in the original formulation. Modern models mostly use a gated variant:

FFN(x)  =  W_2 · ( Swish(W_1 x) ⊙ W_3 x )

⊙ is elementwise multiplication. Because there are now three matrices, d_ff is usually set near (8/3) · d_model to hold the parameter count roughly constant.

Pre-norm placement — normalising the sublayer input rather than its output — is what made very deep stacks trainable without a learning-rate warmup schedule (Xiong et al., 2020). RMSNorm drops the mean-centring step of LayerNorm and normalises by the root mean square alone, which is cheaper and empirically loses nothing.

Positional information

Self-attention is permutation-equivariant, so position has to be injected. Rotary position embeddings (RoPE, Su et al., 2021) rotate pairs of dimensions in Q and K by an angle proportional to absolute position. The resulting dot product depends only on the relative offset, giving relative-position behaviour at absolute-position cost. Context extension methods such as position interpolation and NTK-aware scaling operate directly on RoPE's frequency base.

Cost

Per layer, for sequence length n and width d:

attention scores and their use   O(n² · d)
Q/K/V/O projections and FFN      O(n · d²)

Attention dominates once n > d. For a model with d = 4096, sequences under about 4k tokens are dominated by the dense projections, not by attention. This is why quadratic-attention panic is often misplaced at short context, and entirely justified at long context.

FlashAttention (Dao et al., 2022) keeps the O(n²) arithmetic but tiles the computation so the n × n score matrix is never materialised in high-bandwidth memory. Memory drops to O(n) and wall-clock time improves substantially, because attention is memory-bandwidth-bound rather than compute-bound on modern accelerators.

KV cache

During autoregressive decoding, keys and values for past positions are reused rather than recomputed. Cache size in bytes:

bytes  =  2 · L · n_kv · d_head · n · b

L is the number of layers, n_kv the number of key/value heads, d_head the per-head width, n the number of cached tokens, and b the bytes per element. The leading 2 covers keys and values.

Worked example with L = 32, n_kv = 8, d_head = 128, n = 8192, b = 2:

2 × 32 × 8 × 128 × 8192 × 2  =  1,073,741,824 bytes  ≈  1 GiB

One gibibyte, per concurrent request. This is why grouped-query attention (Ainslie et al., 2023), which shares key and value heads across several query heads, matters so much for serving economics. It shrinks n_kv directly.

Papers

Open questions

Mechanistic interpretability has identified specific circuits, most famously induction heads that implement a copy-the-continuation-of-a-repeated-prefix rule and correlate with the emergence of in-context learning. What remains unresolved is whether these local explanations compose into an account of a full-scale model, or whether they are a thin readable layer over something largely opaque. Treat any confident claim in either direction with suspicion.

What to learn next