Tokeniser Internals

Adding tokens and resizing embeddings

Adding a token to a tokeniser without resizing the model raises an IndexError. Resizing it badly is worse, because the new rows start nowhere useful.

Read these first

On this page 6
  1. Why anyone adds tokens
  2. What goes wrong if you skip the resize
  3. Where the new rows should start
  4. The bigger warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Adding a word to the tokeniser is half the job. The model also needs a new row in its lookup table, and that row starts out meaningless.

Think of a hotel with 300 rooms and 300 keys. You build five new rooms. Guests now ask for room 301, and there is no key on the board.

Cutting five new keys is the easy part. The hard part is that nobody has ever stayed in those rooms, so the staff have no idea what they are like.

A new token is exactly that. The slot exists once you resize. What it means has to be learned.

Why anyone adds tokens

Control markers. Chat role markers, tool-call delimiters, section boundaries. These need to be single ids so the model can never confuse them with text.

Domain vocabulary. A medical model reads drug names constantly, and a legal model reads citation formats. Giving those single entries saves a lot of tokens.

A new language. The original tokeniser may handle a script byte by byte. Adding pieces for it can cut that language's token count sharply.

What goes wrong if you skip the resize

The model's lookup table has one row per known id. Ask for a row past the end and the program stops with an index error.

That is the good outcome. It is loud and it happens immediately.

The bad outcome is resizing but never training the new rows. Then the token exists, the model reads a meaningless row, and you have manufactured your own glitch token.

Where the new rows should start

A freshly created row is random noise, sitting far from every real word. The model's first reaction to it is unpredictable, and early training steps can be violent.

A better starting point is the average of the existing rows. That places the new token in the middle of the crowd rather than outside the building. Training then moves it somewhere sensible from a reasonable start.

Modern libraries do a version of this for you by default, which is worth knowing before you write your own.

   random start:   far from everything, model reacts unpredictably
   averaged start: in the middle of known words, training nudges it outward

The bigger warning

Adding tokens changes the model. The new rows are untrained, and they will stay untrained unless your finetuning data actually uses them, many times.

Adding fifty domain terms and then finetuning on a thousand examples is not enough. Those rows will barely move.

If you cannot train them properly, you are usually better off leaving them out and paying the extra tokens.

Remember this

  • Adding a token to the tokeniser does nothing to the model until you resize it.
  • New rows start meaningless, so start them near the average of the existing ones.
  • If your training data does not use the new tokens heavily, do not add them.

What to learn next

  • Fine-tuning — the training step that makes new tokens mean something.
  • Glitch tokens — what added tokens become if you skip that step.
  • LoRA — adapting a model when full finetuning is out of reach.

Developer — Code and libraries.

Setup

bash
pip install torch "transformers==5.6.2"

Written against transformers 5.6.2 and PyTorch 2.5.1. The model is built from a config with random weights, so nothing is downloaded and it runs on CPU in seconds.

Every case, in one script

resize.py
import torch
from transformers import LlamaConfig, LlamaForCausalLM
from transformers.utils import logging

logging.set_verbosity_error()          # the resize call is chatty by default

BASE = dict(vocab_size=300, hidden_size=64, intermediate_size=128,
            num_hidden_layers=2, num_attention_heads=4, num_key_value_heads=4,
            max_position_embeddings=64)

def report(m, label):
    inp = m.get_input_embeddings().weight
    out = m.get_output_embeddings().weight
    print(f"{label:<26} input {tuple(inp.shape)}  output {tuple(out.shape)}  "
          f"params {sum(p.numel() for p in m.parameters()):,}")

torch.manual_seed(0)
model = LlamaForCausalLM(LlamaConfig(**BASE, tie_word_embeddings=False))
report(model, "before")

old = model.config.vocab_size
model.resize_token_embeddings(old + 5)
report(model, "after adding 5 tokens")
print("config.vocab_size updated to:", model.config.vocab_size)

print(f"\nid {old} now runs through the model:")
print("  logits shape:", tuple(model(torch.tensor([[1, 2, old]])).logits.shape))

print("\nhow the new rows were initialised (this is the part that bites people):")
for mean_resizing in [True, False]:
    torch.manual_seed(0)
    m = LlamaForCausalLM(LlamaConfig(**BASE, tie_word_embeddings=False))
    ref = m.get_input_embeddings().weight.detach().clone()
    m.resize_token_embeddings(old + 5, mean_resizing=mean_resizing)
    new_rows = m.get_input_embeddings().weight[old:].detach()
    dist = (new_rows[:, None, :] - ref[None, :, :]).norm(dim=-1).min(dim=1).values
    print(f"  mean_resizing={str(mean_resizing):<5} new row length "
          f"{new_rows.norm(dim=1).mean():.4f}, existing rows {ref.norm(dim=1).mean():.4f}, "
          f"distance to nearest existing row {dist.mean():.4f}")

print("\nwhat happens if you skip the resize entirely:")
torch.manual_seed(0)
small = LlamaForCausalLM(LlamaConfig(**BASE, tie_word_embeddings=False))
try:
    small(torch.tensor([[1, 2, 400]]))
except IndexError as e:
    print("  IndexError:", str(e).split("\n")[0])

print("\npadding the table to a multiple of 64, which many GPU kernels prefer:")
model.resize_token_embeddings(model.config.vocab_size, pad_to_multiple_of=64)
report(model, "after padding")
print("  the extra rows are real parameters that no token ever selects")

print("\ntied weights: one table serving both ends")
torch.manual_seed(0)
tied = LlamaForCausalLM(LlamaConfig(**BASE, tie_word_embeddings=True))
same = lambda m: (m.get_input_embeddings().weight.data_ptr()
                  == m.get_output_embeddings().weight.data_ptr())
print("  shared storage before resize:", same(tied))
tied.resize_token_embeddings(305)
print("  shared storage after resize :", same(tied))
Output
before                     input (300, 64)  output (300, 64)  params 120,640
after adding 5 tokens      input (305, 64)  output (305, 64)  params 121,280
config.vocab_size updated to: 305

id 300 now runs through the model:
  logits shape: (1, 3, 305)

how the new rows were initialised (this is the part that bites people):
  mean_resizing=True  new row length 0.0095, existing rows 0.1595, distance to nearest existing row 0.1222
  mean_resizing=False new row length 0.1518, existing rows 0.1595, distance to nearest existing row 0.1760

what happens if you skip the resize entirely:
  IndexError: index out of range in self

padding the table to a multiple of 64, which many GPU kernels prefer:
after padding              input (320, 64)  output (320, 64)  params 123,200
  the extra rows are real parameters that no token ever selects

tied weights: one table serving both ends
  shared storage before resize: True
  shared storage after resize : True

Reading that output

5 tokens cost 640 parameters here, 5 rows times 64 dimensions times two tables. On a real model at width 4096 with untied embeddings, 5 tokens cost 40,960 parameters, and a thousand added tokens cost 8.2 million.

config.vocab_size was updated for you. That matters when saving: a checkpoint whose config disagrees with its tensor shapes fails to load.

The two initialisation rows are the interesting part. With mean_resizing=True, the default, new rows have length 0.0095 against 0.1595 for existing rows, and they sit 0.1222 from the nearest real row. With it off, new rows have length 0.1518 — the same scale as trained rows — but sit 0.1760 away, further from everything.

The default draws new rows from a multivariate normal fitted to the mean and covariance of the existing embeddings. So they land inside the existing cloud rather than at an arbitrary point on a sphere around it. That is why they are shorter and closer.

Doing the averaging by hand is now redundant. Older guidance says to set new rows to the mean of the old ones after resizing. transformers has done a better version of that by default since v4.42. Overwriting with a plain mean afterwards makes every new token identical, which is worse: they cannot be distinguished at initialisation and gradients for all of them start identical.

IndexError: index out of range in self is what nn.Embedding gives you for an out-of-range id. Memorise this message. It almost always means an added token, a wrong tokeniser, or an off-by-one in a data pipeline.

Padding from 305 to 320 added 15 rows no token can ever select. They are still parameters, they still receive optimiser state, and the tradeoff is worth it because many matmul kernels are markedly faster on aligned shapes.

Tied weights survive the resize. data_ptr() matching before and after confirms the two tables are the same storage. If you resize a tied model and then assign to the output table directly, you are also writing to the input table.

Adding the tokens themselves

Resizing is step two. Step one is the tokeniser:

add_tokens.py
from tokenizers import Tokenizer, models, trainers, pre_tokenizers, decoders
from transformers import PreTrainedTokenizerFast

tok = Tokenizer(models.BPE())
tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=False)
tok.decoder = decoders.ByteLevel()
tok.train_from_iterator(["the patient was given paracetamol"] * 20,
                        trainers.BpeTrainer(vocab_size=300, show_progress=False,
                                            initial_alphabet=pre_tokenizers.ByteLevel.alphabet()))
t = PreTrainedTokenizerFast(tokenizer_object=tok)

print("before:", len(t), t.tokenize("give 500mg amoxicillin"))
added = t.add_tokens(["amoxicillin", "500mg"])                 # ordinary text tokens
special = t.add_special_tokens({"additional_special_tokens": ["<|dose|>"]})
print(f"added {added} ordinary and {special} special token(s)")
print("after :", len(t), t.tokenize("give 500mg amoxicillin"))
print("marker stays one piece:", t.tokenize("dose <|dose|> here"))
Output
before: 281 ['gi', 'v', 'e', 'Ġ', '5', '0', '0', 'm', 'g', 'Ġ', 'am', 'o', 'x', 'i', 'c', 'i', 'l', 'l', 'i', 'n']
added 2 ordinary and 1 special token(s)
after : 284 ['gi', 'v', 'e', 'Ġ', '500mg', 'Ġ', 'amoxicillin']
marker stays one piece: ['d', 'o', 's', 'e', 'Ġ', '<|dose|>', 'Ġ', 'he', 'r', 'e']

20 tokens down to 7. Two of the additions collapsed sixteen tokens into two, and <|dose|> stays a single piece surrounded by character-level noise from a tokeniser that never saw the word "dose". That is the gain, and it is only realised if the model learns what those three rows mean.

The full procedure

  1. Add the tokens to the tokeniser and save it.
  2. Call model.resize_token_embeddings(len(tokenizer)). Use len(tokenizer), never a hand-counted number.
  3. Optionally pass pad_to_multiple_of=64 for kernel alignment.
  4. Save the model and the tokeniser together. They are now a matched pair and neither works with an old copy of the other.
  5. Finetune on data that uses the new tokens frequently.
  6. Verify: check that the new rows moved, using the norm and cosine checks from the glitch-tokens lesson.

Step 6 is the one people skip, and skipping it is how added tokens become glitch tokens.

Common mistakes

Resizing to a hand-counted number. len(tokenizer) accounts for added tokens, added special tokens and any gaps. A literal drifts out of date the moment anyone edits the token list.

Saving the model without the tokeniser. The mismatch is silent. Ids mean different things, and the model produces fluent, wrong text.

Adding tokens to a model you cannot finetune. Untrained rows are worse than the token savings are worth. Measure the fertility gain before deciding it is worth the risk.

Assuming add_tokens returns the count you passed. It returns the number actually added. A string already in the vocabulary is skipped, so the return value is what to check.

Overwriting new rows with a single mean vector. Older tutorials advise this. It makes every new token identical at initialisation, which is worse than the library's default. Leave mean_resizing alone unless you have measured a reason.

Shrinking the vocabulary casually. resize_token_embeddings accepts a smaller number and truncates rows. Any id past the new end becomes invalid, and the tokeniser must be trimmed to match.

Try it yourself

Resize a model, then run the glitch-token detector from the previous lesson on the five new ids before any finetuning. They will score as under-trained by a wide margin. Finetune briefly on text that uses them and rerun the detector. Watching the score move is the only real confirmation that added tokens have learned anything.

What to learn next

  • Fine-tuning — the training step that makes new tokens mean something.
  • Glitch tokens — what added tokens become if you skip that step.
  • LoRA — adapting a model when full finetuning is out of reach.

Researcher — Mathematics and papers.

What the operation does

resize_token_embeddings(n) replaces the input embedding $E \in \mathbb{R}^{V \times d}$ with $E' \in \mathbb{R}^{n \times d}$, copying $\min(V, n)$ rows, and does the same for the output projection when embeddings are untied. It updates config.vocab_size and re-ties the two tensors if the config requests tying.

Parameter delta is $(n - V)d$ with tied embeddings and $2(n - V)d$ without. At $d = 4096$ and 1000 added tokens, that is 4.2M or 8.4M parameters.

Initialisation of new rows

The default since transformers v4.42 is mean_resizing=True, which samples new rows from a multivariate normal fitted to the mean and covariance of the existing embedding rows:

$$ e_{\text{new}} \sim \mathcal{N}!\big(\mu_E,\, \Sigma_E\big), \qquad \mu_E = \frac{1}{V}\sum_i E_i, \quad \Sigma_E = \frac{1}{V}\sum_i (E_i - \mu_E)(E_i - \mu_E)^\top $$

The method follows Hewitt (2021), Initializing New Word Embeddings for Pretrained Language Models. The argument is that the trained embedding distribution occupies a specific region of $\mathbb{R}^d$ that a standard normal initialisation does not match, and a mismatch produces large early gradients and can destabilise the first optimisation steps.

mean_resizing=False falls back to the model's ordinary initialiser, typically $\mathcal{N}(0, \sigma^2)$ with $\sigma$ from the config. The developer output shows the practical difference: sampled-from-distribution rows are shorter and closer to existing rows than standard-normal rows.

An older recipe sets every new row to $\mu_E$ exactly. That is degenerate: all new tokens are identical at initialisation, so their gradients are identical for the first step and they separate only through their differing contexts. It is strictly worse than sampling and should be considered superseded.

Vocabulary extension for new languages

The main serious use of this operation. The procedure:

  1. Train a tokeniser on the target-language corpus.
  2. Merge its vocabulary into the base tokeniser, keeping the base ids stable.
  3. Resize the model.
  4. Initialise new rows, either by the multivariate method or from a translation-based mapping of existing rows.
  5. Continue pretraining on target-language text, often with the base model's other parameters frozen for an initial phase.

Cui et al. (2023), Chinese-LLaMA, extend Llama's 32k vocabulary by roughly 20k Chinese tokens and report substantially reduced fertility and improved Chinese performance after continued pretraining. Fujii et al. (2024), Swallow, do the equivalent for Japanese. Both report that the continued-pretraining budget, not the initialisation, dominates the outcome.

Minixhofer et al. (2022), WECHSEL, and follow-up work initialise new embeddings from a cross-lingual word-embedding alignment rather than from a distribution fit, and report faster convergence. That approach requires an aligned static embedding space for the language pair.

Interaction with tied embeddings

With $E$ shared between input and output, resizing produces one tensor used at both ends. Two consequences:

  • Parameter cost halves.
  • The output-side gradient, which pushes every non-target logit down, also acts on the input embedding. The clean separation exploited by glitch-token detection disappears, and diagnostics based on comparing the two tables no longer apply.

Alignment padding

pad_to_multiple_of exists because GEMM kernels are faster on aligned dimensions. A vocabulary of 128256 rather than 128000 or 128001 is the visible artefact of this in published models. Padding rows are unreachable by construction and are therefore glitch tokens by definition, though harmless ones as long as nothing can emit their ids. Constraining generation to exclude them is cheap insurance.

Practical guidance

Adding tokens is justified when:

  • the tokens appear frequently enough in the finetuning corpus to receive substantial gradient, on the order of thousands of occurrences each;
  • the fertility gain is large, which in practice means a new script or a genuinely dense domain vocabulary;
  • you control the tokeniser and the model together and can ship them as a pair.

It is not justified for a handful of terms and a small finetune. The token savings are small, and the untrained rows are a live failure mode.

Papers

What to learn next

  • Fine-tuning — the training step that makes new tokens mean something.
  • Glitch tokens — what added tokens become if you skip that step.
  • LoRA — adapting a model when full finetuning is out of reach.

What to learn next

These follow on from what you just read.

  • How Text Is Generated

    Prefill and decode

    A model reads your entire prompt in one fast pass, then writes the answer one token at a time in a slow loop. Almost every delay in text generation comes from that split.

  • How Text Is Generated

    The KV cache

    Instead of re-reading the whole conversation for every new word, the model keeps a running summary of what each earlier token contributed and appends one row per new token.

  • How Text Is Generated

    How much memory the KV cache eats

    The cache costs a fixed number of bytes per token, and multiplying that by context length and users decides how many people your GPU can serve at once.