Counting a model's parameters by hand
Turn any config file into an exact parameter count with arithmetic you can do on paper, and check it against two real published models.
- 13 min read
- 3 reading levels
- Updated
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
You can work out a model's exact size from six numbers in its config file, without downloading anything.
Think about tiling a room. You do not count the tiles one by one. You measure the room, check the size of one tile, and multiply. The answer is exact, and you got it in a minute with a tape measure.
A transformer is far more regular than a room. It is one block repeated, so counting one block and multiplying gives you the whole thing.
Why this is a skill worth having
Four practical reasons, and every one of them comes up in real work.
Will it fit? A model's weights need memory. Knowing the count before you download forty gigabytes saves an afternoon.
Is the name honest? "7B" is marketing rounded to one digit. Two models both called 7B can differ by hundreds of millions.
Where did the size go? Most of it sits in the feedforward layers, or in the vocabulary table. That tells you which knob to turn.
Does my code match the paper? Build the architecture from a description and count it. If the number is off by a few percent, you have a bug. This is one of the few cheap, exact tests available in machine learning.
The six numbers
Open any model's config.json and look for these.
| What it is called | What it means |
|---|---|
vocab_size | how many different tokens the model knows |
hidden_size | the width carried between blocks |
num_hidden_layers | how many copies of the block |
intermediate_size | the wide middle of the feedforward layer |
num_key_value_heads | key and value heads, if fewer than query heads |
tie_word_embeddings | whether the input and output tables are shared |
That is everything. The rest of the file is tokenizer details and defaults.
The shape of the count
Three groups, added together.
The vocabulary table. One row per token, one column per slot of the model width. If the input and output tables are separate, count it twice.
The blocks. Work out one block, then multiply by the number of blocks. Inside a block, the attention part has four matrices. The feedforward part has two or three, depending on whether the model gates.
The odds and ends. The normalisation layers. Tiny by comparison, and included only because we want an exact answer, not a rough one.
total = vocabulary table (once or twice)
+ one block, multiplied by the number of blocks
+ the normalisation layersWhat the number means in memory
Each parameter is stored as a number, and how much room it takes depends on the format.
- Full precision: four bytes each.
- Half precision, the normal choice for serving: two bytes each.
- Eight-bit: one byte each.
- Four-bit: half a byte each.
A model with eight billion parameters needs about sixteen gigabytes at half precision. At four-bit it needs about four.
Two warnings. Training needs several times more than serving, because the optimiser keeps extra copies. Serving also needs room for the running cache of what has been read. For long inputs that cache can rival the weights themselves.
Remember this
- Six numbers from a config file give you the exact size.
- Count the vocabulary table, count one block and multiply, then add the norms.
- Multiply the count by the bytes per number to get the memory it needs.
What to learn next
- FLOPs per token — turning the parameter count into a training bill.
- Quantization in practice — cutting the bytes per parameter.
- Model compression — cutting the parameter count itself.
Developer — Code and libraries.
Setup
pip install torchTwo formulas, checked twice
The right way to trust a formula is to build a model that matches it exactly and compare. Then apply it to published models and compare again.
def gpt2_style(V, d, L, d_ff, max_pos):
"""GPT-2 shape: LayerNorm with bias, biases on every linear, learned positions, tied head."""
embeddings = V * d + max_pos * d
attn = (3 * d * d + 3 * d) + (d * d + d) # packed QKV, then the output projection
ffn = (d * d_ff + d_ff) + (d_ff * d + d)
norms = 2 * (2 * d) # two LayerNorms, each a gain and a bias
return embeddings + L * (attn + ffn + norms) + 2 * d # + the final LayerNorm
def llama_style(V, d, L, d_ff, n_heads, n_kv_heads):
"""Llama shape: RMSNorm, no biases anywhere, grouped-query attention, SwiGLU, untied head."""
head = d // n_heads
attn = d * d + 2 * (d * n_kv_heads * head) + d * d
ffn = 3 * (d * d_ff) # gate, up and down
norms = 2 * d # two RMSNorms, gain only
return V * d + L * (attn + ffn + norms) + d + V * d
print("check the formulas against models actually built in PyTorch")
import torch.nn as nn, torch, torch.nn.functional as F
class G2Block(nn.Module):
def __init__(self, d, d_ff):
super().__init__()
self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
self.qkv, self.proj = nn.Linear(d, 3*d), nn.Linear(d, d)
self.f1, self.f2 = nn.Linear(d, d_ff), nn.Linear(d_ff, d)
class LlBlock(nn.Module):
def __init__(self, d, d_ff, nh, nkv):
super().__init__()
h = d // nh
self.n1, self.n2 = nn.RMSNorm(d), nn.RMSNorm(d)
self.q = nn.Linear(d, d, bias=False)
self.k = nn.Linear(d, nkv*h, bias=False)
self.v = nn.Linear(d, nkv*h, bias=False)
self.o = nn.Linear(d, d, bias=False)
self.gate = nn.Linear(d, d_ff, bias=False)
self.up = nn.Linear(d, d_ff, bias=False)
self.down = nn.Linear(d_ff, d, bias=False)
V, d, L, d_ff, mp = 500, 32, 3, 128, 64
g2 = nn.ModuleList([nn.Embedding(V, d), nn.Embedding(mp, d), nn.LayerNorm(d)] +
[G2Block(d, d_ff) for _ in range(L)])
print(f" GPT-2 shape : formula {gpt2_style(V,d,L,d_ff,mp):>7,} "
f"built {sum(p.numel() for p in g2.parameters()):>7,}")
nh, nkv = 4, 2
ll = nn.ModuleList([nn.Embedding(V, d), nn.RMSNorm(d), nn.Linear(d, V, bias=False)] +
[LlBlock(d, d_ff, nh, nkv) for _ in range(L)])
print(f" Llama shape : formula {llama_style(V,d,L,d_ff,nh,nkv):>7,} "
f"built {sum(p.numel() for p in ll.parameters()):>7,}")
print("\nnow run the formulas on two real, published models")
gpt2 = gpt2_style(V=50257, d=768, L=12, d_ff=3072, max_pos=1024)
llama = llama_style(V=128256, d=4096, L=32, d_ff=14336, n_heads=32, n_kv_heads=8)
print(f" GPT-2 small : {gpt2:>15,} published 124,439,808 "
f"{'MATCH' if gpt2 == 124_439_808 else 'MISMATCH'}")
print(f" Llama 3 8B : {llama:>15,} published 8,030,261,248 "
f"{'MATCH' if llama == 8_030_261_248 else 'MISMATCH'}")
print("\nwhere Llama 3 8B's 8 billion actually sit")
d_, L_, dff_, V_ = 4096, 32, 14336, 128256
parts = [("token embedding", V_ * d_),
("output head (untied)", V_ * d_),
("attention, all layers", L_ * (d_*d_ + 2*(d_*1024) + d_*d_)),
("feedforward, all layers", L_ * 3 * d_ * dff_),
("norms", L_ * 2 * d_ + d_)]
total = sum(p for _, p in parts)
for name, p in parts:
print(f" {name:24s} {p:>15,} {p/total:>7.1%}")
print(f" {'TOTAL':24s} {total:>15,}")check the formulas against models actually built in PyTorch GPT-2 shape : formula 56,224 built 56,224 Llama shape : formula 78,304 built 78,304 now run the formulas on two real, published models GPT-2 small : 124,439,808 published 124,439,808 MATCH Llama 3 8B : 8,030,261,248 published 8,030,261,248 MATCH where Llama 3 8B's 8 billion actually sit token embedding 525,336,576 6.5% output head (untied) 525,336,576 6.5% attention, all layers 1,342,177,280 16.7% feedforward, all layers 5,637,144,576 70.2% norms 266,240 0.0% TOTAL 8,030,261,248
Both formulas are exact, and that is the point
The small models match to the parameter. 56,224 and 78,304, formula against PyTorch. If a formula is right on a toy model built to the same recipe, it is right at scale.
Both published totals match exactly. 124,439,808 for GPT-2 small. Verify it yourself with sum(p.numel() for p in GPT2LMHeadModel(GPT2Config()).parameters()). 8,030,261,248 for Llama 3 8B. No rounding, no fudge factor.
Seventy percent of Llama 3 8B is feedforward. Not attention. Not embeddings. Three matrices per block, each 4096 by 14336, across 32 blocks. When people talk about where a model's knowledge is stored, this is the part they mean.
The norms are 0.0 percent. 266,240 parameters out of eight billion. Remove them and the model does not train at all. Parameter share is a measure of storage, never of importance.
Where the two shapes differ
| GPT-2 | Llama 3 | |
|---|---|---|
| Norm | LayerNorm, gain and bias | RMSNorm, gain only |
| Biases on linear layers | yes | no |
| Position | a learned table of max_pos rows | rotary, no parameters |
| Feedforward | 2 matrices | 3 matrices, gated |
| Key/value heads | same as query heads | 8 against 32 |
| Output head | tied to the input table | separate |
Six differences, and each one changes the arithmetic. This is why one universal formula does not exist. Read the config rather than reuse a formula from a blog post.
The rough version, for a quick estimate
For a decoder-only model with a four-times feedforward multiplier:
non-embedding parameters ≈ 12 * L * d * dFour units for attention plus eight for the feedforward layer. For GPT-2 small, 12 times 12 times 768 squared is 84,934,656. The true figure is 85,054,464, so the estimate is within 0.15 percent. Good enough for a back-of-envelope estimate. It is the same 12 L d² that underlies the training-cost rule in FLOPs per token.
From parameters to memory
def bytes_needed(n_params, bits=16):
return n_params * bits / 8
for bits in (32, 16, 8, 4):
gb = bytes_needed(8_030_261_248, bits) / 1024**3
print(f" {bits:>2}-bit: {gb:6.1f} GB of weights")32-bit: 29.9 GB of weights 16-bit: 15.0 GB of weights 8-bit: 7.5 GB of weights 4-bit: 3.7 GB of weights
Weights alone. For training, add gradients at the same size as the weights. Adam's state costs roughly twice that again. Plan on four to six times the weight memory. For serving, add the KV cache. It grows with the number of tokens held and with concurrent requests.
Common mistakes
Counting a tied output head twice. sum(p.numel() for p in model.parameters()) already deduplicates shared tensors in PyTorch. A hand-written formula does not. Check tie_word_embeddings.
Forgetting that grouped-query attention shrinks K and V. Llama 3 8B's key and value projections are 4096 by 1024, not 4096 by 4096. Assuming square projections overstates attention parameters by roughly fifty percent.
Assuming the feedforward multiplier is 4. Llama 3 8B uses 3.5. Read intermediate_size rather than computing it.
Confusing parameter count with download size. A repository holds the weights in their stored format plus tokenizer files and metadata. A 4-bit quantized checkpoint of an 8B model is a few gigabytes, not sixteen.
Try it yourself
Take a config file for a model you actually use and run it through the right formula. Then load the model with sum(p.numel() for p in model.parameters()) and compare. If they disagree, the difference tells you exactly which component you got wrong. A mismatch of V * d means the head is tied and you counted it twice. A mismatch of L * 2 * d means you assumed the wrong normalisation.
What to learn next
- FLOPs per token — turning the parameter count into a training bill.
- Quantization in practice — cutting the bytes per parameter.
- Model compression — cutting the parameter count itself.
Researcher — Mathematics and papers.
The general form
Take a decoder-only transformer with vocabulary $V$, width $d$ and depth $L$. Let $d_{\text{ff}}$ be the feedforward width, $h$ the query heads, $g$ the key-value groups and $d_h = d/h$:
$$ N = \underbrace{(2 - \tau) V d}_{\text{embeddings}} + L \underbrace{\left( 2d^2 + 2 g d_h d \right)}{\text{attention}} + L \underbrace{\left( m\, d\, d{\text{ff}} \right)}{\text{feedforward}} + \underbrace{N{\text{norm}} + N_{\text{pos}}}_{\text{small}} $$
where $\tau = 1$ if the output head is tied and $0$ otherwise, and $m \in {2, 3}$ is the number of feedforward matrices. Norms contribute $L \cdot c \cdot d + c' d$ with $c = 2$ for two norms per layer and $c \in {2, 4}$ counting gain and bias. Learned position embeddings add $T_{\max} d$; rotary and ALiBi add nothing.
Substituting $g = h$, $m = 2$, $d_{\text{ff}} = 4d$ gives the familiar $12 L d^2$ for the non-embedding term.
Which count to report
Three counts are in circulation and they differ substantially at small scale:
- Total. Everything, deduplicating tied tensors. What
sum(p.numel() for p in model.parameters())returns. - Non-embedding. Excludes the token embedding table and, in older conventions, position embeddings. This is what Kaplan et al. (2020), arXiv:2001.08361, use for their scaling laws. Embedding parameters do not participate in per-token compute the way weight matrices do.
- Activated. For mixture-of-experts models, the parameters actually used for a given token. A model with 100B total and 6B activated has the compute cost of a 6B dense model. Its memory footprint is that of a 100B one.
For GPT-2 small: 124,439,808 total, 85,056,000 non-embedding. The 32 percent gap is why the distinction matters below a billion parameters. It stops mattering above about 10B.
Always state which count you mean. Papers that do not are a recurring source of failed reproductions.
Memory, precisely
Inference weights: $N \cdot b/8$ bytes at $b$ bits.
Training with Adam in mixed precision costs 16 bytes per parameter. That is 2 bytes for the bf16 weight and 4 for the fp32 master copy. Then 4 each for the two Adam moments, and 2 for the gradient. Roughly 8 times the inference footprint. Rajbhandari et al. (2020), ZeRO, arXiv:1910.02054, work through this budget and show how to shard it across devices.
KV cache, per token, per sequence:
$$ 2 \cdot L \cdot g \cdot d_h \cdot b/8 \ \text{bytes} $$
Llama 3 8B at bf16: $2 \times 32 \times 8 \times 128 \times 2 = 131{,}072$ bytes per token, or 128 KiB. At 32,768 tokens that is 4 GiB for a single sequence. Comparable to a quarter of the 15 GiB of weights. Without grouped-query attention it would be four times that.
Where the budget goes as models scale
Feedforward share is $m d_{\text{ff}} / (m d_{\text{ff}} + 2d + 2 g d_h)$. For Llama 3 8B that is 70 percent. It stays in the 60 to 75 percent band across almost all dense decoders. Embedding share falls monotonically with scale. It is 31 percent for GPT-2 small and 13 percent for Llama 3 8B.
This is the arithmetic behind two current design trends. Mixture-of-experts targets the 70 percent that is feedforward. Grouped-query attention and KV compression target the cache rather than the weights. The cache is what limits concurrency in serving.
Papers
- Kaplan et al., Scaling Laws for Neural Language Models, 2020 — arxiv.org/abs/2001.08361
- Hoffmann et al., Training Compute-Optimal Large Language Models, 2022 — arxiv.org/abs/2203.15556
- Rajbhandari et al., ZeRO, 2020 — arxiv.org/abs/1910.02054
- Ainslie et al., GQA, 2023 — arxiv.org/abs/2305.13245
What to learn next
- FLOPs per token — turning the parameter count into a training bill.
- Quantization in practice — cutting the bytes per parameter.
- Model compression — cutting the parameter count itself.