Inside a Transformer Block

Counting a model's parameters by hand

Turn any config file into an exact parameter count with arithmetic you can do on paper, and check it against two real published models.

On this page 6
  1. Why this is a skill worth having
  2. The six numbers
  3. The shape of the count
  4. What the number means in memory
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

You can work out a model's exact size from six numbers in its config file, without downloading anything.

Think about tiling a room. You do not count the tiles one by one. You measure the room, check the size of one tile, and multiply. The answer is exact, and you got it in a minute with a tape measure.

A transformer is far more regular than a room. It is one block repeated, so counting one block and multiplying gives you the whole thing.

Why this is a skill worth having

Four practical reasons, and every one of them comes up in real work.

Will it fit? A model's weights need memory. Knowing the count before you download forty gigabytes saves an afternoon.

Is the name honest? "7B" is marketing rounded to one digit. Two models both called 7B can differ by hundreds of millions.

Where did the size go? Most of it sits in the feedforward layers, or in the vocabulary table. That tells you which knob to turn.

Does my code match the paper? Build the architecture from a description and count it. If the number is off by a few percent, you have a bug. This is one of the few cheap, exact tests available in machine learning.

The six numbers

Open any model's config.json and look for these.

What it is calledWhat it means
vocab_sizehow many different tokens the model knows
hidden_sizethe width carried between blocks
num_hidden_layershow many copies of the block
intermediate_sizethe wide middle of the feedforward layer
num_key_value_headskey and value heads, if fewer than query heads
tie_word_embeddingswhether the input and output tables are shared

That is everything. The rest of the file is tokenizer details and defaults.

The shape of the count

Three groups, added together.

The vocabulary table. One row per token, one column per slot of the model width. If the input and output tables are separate, count it twice.

The blocks. Work out one block, then multiply by the number of blocks. Inside a block, the attention part has four matrices. The feedforward part has two or three, depending on whether the model gates.

The odds and ends. The normalisation layers. Tiny by comparison, and included only because we want an exact answer, not a rough one.

   total  =  vocabulary table (once or twice)
           + one block, multiplied by the number of blocks
           + the normalisation layers

What the number means in memory

Each parameter is stored as a number, and how much room it takes depends on the format.

  • Full precision: four bytes each.
  • Half precision, the normal choice for serving: two bytes each.
  • Eight-bit: one byte each.
  • Four-bit: half a byte each.

A model with eight billion parameters needs about sixteen gigabytes at half precision. At four-bit it needs about four.

Two warnings. Training needs several times more than serving, because the optimiser keeps extra copies. Serving also needs room for the running cache of what has been read. For long inputs that cache can rival the weights themselves.

Remember this

  • Six numbers from a config file give you the exact size.
  • Count the vocabulary table, count one block and multiply, then add the norms.
  • Multiply the count by the bytes per number to get the memory it needs.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Two formulas, checked twice

The right way to trust a formula is to build a model that matches it exactly and compare. Then apply it to published models and compare again.

count_params.py
def gpt2_style(V, d, L, d_ff, max_pos):
    """GPT-2 shape: LayerNorm with bias, biases on every linear, learned positions, tied head."""
    embeddings = V * d + max_pos * d
    attn = (3 * d * d + 3 * d) + (d * d + d)          # packed QKV, then the output projection
    ffn = (d * d_ff + d_ff) + (d_ff * d + d)
    norms = 2 * (2 * d)                                # two LayerNorms, each a gain and a bias
    return embeddings + L * (attn + ffn + norms) + 2 * d   # + the final LayerNorm

def llama_style(V, d, L, d_ff, n_heads, n_kv_heads):
    """Llama shape: RMSNorm, no biases anywhere, grouped-query attention, SwiGLU, untied head."""
    head = d // n_heads
    attn = d * d + 2 * (d * n_kv_heads * head) + d * d
    ffn = 3 * (d * d_ff)                               # gate, up and down
    norms = 2 * d                                      # two RMSNorms, gain only
    return V * d + L * (attn + ffn + norms) + d + V * d

print("check the formulas against models actually built in PyTorch")
import torch.nn as nn, torch, torch.nn.functional as F

class G2Block(nn.Module):
    def __init__(self, d, d_ff):
        super().__init__()
        self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
        self.qkv, self.proj = nn.Linear(d, 3*d), nn.Linear(d, d)
        self.f1, self.f2 = nn.Linear(d, d_ff), nn.Linear(d_ff, d)

class LlBlock(nn.Module):
    def __init__(self, d, d_ff, nh, nkv):
        super().__init__()
        h = d // nh
        self.n1, self.n2 = nn.RMSNorm(d), nn.RMSNorm(d)
        self.q = nn.Linear(d, d, bias=False)
        self.k = nn.Linear(d, nkv*h, bias=False)
        self.v = nn.Linear(d, nkv*h, bias=False)
        self.o = nn.Linear(d, d, bias=False)
        self.gate = nn.Linear(d, d_ff, bias=False)
        self.up   = nn.Linear(d, d_ff, bias=False)
        self.down = nn.Linear(d_ff, d, bias=False)

V, d, L, d_ff, mp = 500, 32, 3, 128, 64
g2 = nn.ModuleList([nn.Embedding(V, d), nn.Embedding(mp, d), nn.LayerNorm(d)] +
                   [G2Block(d, d_ff) for _ in range(L)])
print(f"  GPT-2 shape : formula {gpt2_style(V,d,L,d_ff,mp):>7,}   "
      f"built {sum(p.numel() for p in g2.parameters()):>7,}")

nh, nkv = 4, 2
ll = nn.ModuleList([nn.Embedding(V, d), nn.RMSNorm(d), nn.Linear(d, V, bias=False)] +
                   [LlBlock(d, d_ff, nh, nkv) for _ in range(L)])
print(f"  Llama shape : formula {llama_style(V,d,L,d_ff,nh,nkv):>7,}   "
      f"built {sum(p.numel() for p in ll.parameters()):>7,}")

print("\nnow run the formulas on two real, published models")
gpt2 = gpt2_style(V=50257, d=768, L=12, d_ff=3072, max_pos=1024)
llama = llama_style(V=128256, d=4096, L=32, d_ff=14336, n_heads=32, n_kv_heads=8)
print(f"  GPT-2 small : {gpt2:>15,}   published 124,439,808   "
      f"{'MATCH' if gpt2 == 124_439_808 else 'MISMATCH'}")
print(f"  Llama 3 8B  : {llama:>15,}   published 8,030,261,248 "
      f"{'MATCH' if llama == 8_030_261_248 else 'MISMATCH'}")

print("\nwhere Llama 3 8B's 8 billion actually sit")
d_, L_, dff_, V_ = 4096, 32, 14336, 128256
parts = [("token embedding", V_ * d_),
         ("output head (untied)", V_ * d_),
         ("attention, all layers", L_ * (d_*d_ + 2*(d_*1024) + d_*d_)),
         ("feedforward, all layers", L_ * 3 * d_ * dff_),
         ("norms", L_ * 2 * d_ + d_)]
total = sum(p for _, p in parts)
for name, p in parts:
    print(f"  {name:24s} {p:>15,} {p/total:>7.1%}")
print(f"  {'TOTAL':24s} {total:>15,}")
Output
check the formulas against models actually built in PyTorch
  GPT-2 shape : formula  56,224   built  56,224
  Llama shape : formula  78,304   built  78,304

now run the formulas on two real, published models
  GPT-2 small :     124,439,808   published 124,439,808   MATCH
  Llama 3 8B  :   8,030,261,248   published 8,030,261,248 MATCH

where Llama 3 8B's 8 billion actually sit
  token embedding              525,336,576    6.5%
  output head (untied)         525,336,576    6.5%
  attention, all layers      1,342,177,280   16.7%
  feedforward, all layers    5,637,144,576   70.2%
  norms                            266,240    0.0%
  TOTAL                      8,030,261,248

Both formulas are exact, and that is the point

The small models match to the parameter. 56,224 and 78,304, formula against PyTorch. If a formula is right on a toy model built to the same recipe, it is right at scale.

Both published totals match exactly. 124,439,808 for GPT-2 small. Verify it yourself with sum(p.numel() for p in GPT2LMHeadModel(GPT2Config()).parameters()). 8,030,261,248 for Llama 3 8B. No rounding, no fudge factor.

Seventy percent of Llama 3 8B is feedforward. Not attention. Not embeddings. Three matrices per block, each 4096 by 14336, across 32 blocks. When people talk about where a model's knowledge is stored, this is the part they mean.

The norms are 0.0 percent. 266,240 parameters out of eight billion. Remove them and the model does not train at all. Parameter share is a measure of storage, never of importance.

Where the two shapes differ

GPT-2Llama 3
NormLayerNorm, gain and biasRMSNorm, gain only
Biases on linear layersyesno
Positiona learned table of max_pos rowsrotary, no parameters
Feedforward2 matrices3 matrices, gated
Key/value headssame as query heads8 against 32
Output headtied to the input tableseparate

Six differences, and each one changes the arithmetic. This is why one universal formula does not exist. Read the config rather than reuse a formula from a blog post.

The rough version, for a quick estimate

For a decoder-only model with a four-times feedforward multiplier:

non-embedding parameters  ≈  12 * L * d * d

Four units for attention plus eight for the feedforward layer. For GPT-2 small, 12 times 12 times 768 squared is 84,934,656. The true figure is 85,054,464, so the estimate is within 0.15 percent. Good enough for a back-of-envelope estimate. It is the same 12 L d² that underlies the training-cost rule in FLOPs per token.

From parameters to memory

python
def bytes_needed(n_params, bits=16):
    return n_params * bits / 8

for bits in (32, 16, 8, 4):
    gb = bytes_needed(8_030_261_248, bits) / 1024**3
    print(f"  {bits:>2}-bit: {gb:6.1f} GB of weights")
Output
  32-bit:   29.9 GB of weights
  16-bit:   15.0 GB of weights
   8-bit:    7.5 GB of weights
   4-bit:    3.7 GB of weights

Weights alone. For training, add gradients at the same size as the weights. Adam's state costs roughly twice that again. Plan on four to six times the weight memory. For serving, add the KV cache. It grows with the number of tokens held and with concurrent requests.

Common mistakes

Counting a tied output head twice. sum(p.numel() for p in model.parameters()) already deduplicates shared tensors in PyTorch. A hand-written formula does not. Check tie_word_embeddings.

Forgetting that grouped-query attention shrinks K and V. Llama 3 8B's key and value projections are 4096 by 1024, not 4096 by 4096. Assuming square projections overstates attention parameters by roughly fifty percent.

Assuming the feedforward multiplier is 4. Llama 3 8B uses 3.5. Read intermediate_size rather than computing it.

Confusing parameter count with download size. A repository holds the weights in their stored format plus tokenizer files and metadata. A 4-bit quantized checkpoint of an 8B model is a few gigabytes, not sixteen.

Try it yourself

Take a config file for a model you actually use and run it through the right formula. Then load the model with sum(p.numel() for p in model.parameters()) and compare. If they disagree, the difference tells you exactly which component you got wrong. A mismatch of V * d means the head is tied and you counted it twice. A mismatch of L * 2 * d means you assumed the wrong normalisation.

What to learn next

Researcher — Mathematics and papers.

The general form

Take a decoder-only transformer with vocabulary $V$, width $d$ and depth $L$. Let $d_{\text{ff}}$ be the feedforward width, $h$ the query heads, $g$ the key-value groups and $d_h = d/h$:

$$ N = \underbrace{(2 - \tau) V d}_{\text{embeddings}} + L \underbrace{\left( 2d^2 + 2 g d_h d \right)}{\text{attention}} + L \underbrace{\left( m\, d\, d{\text{ff}} \right)}{\text{feedforward}} + \underbrace{N{\text{norm}} + N_{\text{pos}}}_{\text{small}} $$

where $\tau = 1$ if the output head is tied and $0$ otherwise, and $m \in {2, 3}$ is the number of feedforward matrices. Norms contribute $L \cdot c \cdot d + c' d$ with $c = 2$ for two norms per layer and $c \in {2, 4}$ counting gain and bias. Learned position embeddings add $T_{\max} d$; rotary and ALiBi add nothing.

Substituting $g = h$, $m = 2$, $d_{\text{ff}} = 4d$ gives the familiar $12 L d^2$ for the non-embedding term.

Which count to report

Three counts are in circulation and they differ substantially at small scale:

  • Total. Everything, deduplicating tied tensors. What sum(p.numel() for p in model.parameters()) returns.
  • Non-embedding. Excludes the token embedding table and, in older conventions, position embeddings. This is what Kaplan et al. (2020), arXiv:2001.08361, use for their scaling laws. Embedding parameters do not participate in per-token compute the way weight matrices do.
  • Activated. For mixture-of-experts models, the parameters actually used for a given token. A model with 100B total and 6B activated has the compute cost of a 6B dense model. Its memory footprint is that of a 100B one.

For GPT-2 small: 124,439,808 total, 85,056,000 non-embedding. The 32 percent gap is why the distinction matters below a billion parameters. It stops mattering above about 10B.

Always state which count you mean. Papers that do not are a recurring source of failed reproductions.

Memory, precisely

Inference weights: $N \cdot b/8$ bytes at $b$ bits.

Training with Adam in mixed precision costs 16 bytes per parameter. That is 2 bytes for the bf16 weight and 4 for the fp32 master copy. Then 4 each for the two Adam moments, and 2 for the gradient. Roughly 8 times the inference footprint. Rajbhandari et al. (2020), ZeRO, arXiv:1910.02054, work through this budget and show how to shard it across devices.

KV cache, per token, per sequence:

$$ 2 \cdot L \cdot g \cdot d_h \cdot b/8 \ \text{bytes} $$

Llama 3 8B at bf16: $2 \times 32 \times 8 \times 128 \times 2 = 131{,}072$ bytes per token, or 128 KiB. At 32,768 tokens that is 4 GiB for a single sequence. Comparable to a quarter of the 15 GiB of weights. Without grouped-query attention it would be four times that.

Where the budget goes as models scale

Feedforward share is $m d_{\text{ff}} / (m d_{\text{ff}} + 2d + 2 g d_h)$. For Llama 3 8B that is 70 percent. It stays in the 60 to 75 percent band across almost all dense decoders. Embedding share falls monotonically with scale. It is 31 percent for GPT-2 small and 13 percent for Llama 3 8B.

This is the arithmetic behind two current design trends. Mixture-of-experts targets the 70 percent that is feedforward. Grouped-query attention and KV compression target the cache rather than the weights. The cache is what limits concurrency in serving.

Papers

What to learn next