Tokeniser Internals

Training a tokeniser on your own text

A tokeniser has five separate stages, and each one changes the result. Here is the whole pipeline built from scratch, saved, and loaded into transformers.

Read these first

On this page 7
  1. Why you might train your own
  2. The five stages
  3. What each stage decides
  4. The rule to remember
  5. Where you have seen this
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Training a tokeniser means showing it your text and letting it work out which chunks are worth having.

Think of a new cook learning your kitchen. On day one they fetch every spice separately. After a month they reach for "the garam masala mix" as one thing, because in your kitchen those spices always go together.

A tokeniser trained on your text does the same. If your documents are full of medical terms, those become single tokens. If they are full of Python, def and self. become single tokens.

A general-purpose tokeniser would break all of those into fragments, and every fragment costs money and context.

Why you might train your own

Three honest reasons, and one reason not to.

Your text is unusual. Legal, medical, chemical or code-heavy text tokenises badly under a general vocabulary. Fewer tokens means cheaper requests and more content per window.

Your language is underserved. Many tokenisers were trained mostly on English. The same sentence in your language may cost three times more tokens.

You are pretraining a model from scratch. Then the tokeniser is yours to choose, and it should match your data.

The reason not to: if you are finetuning an existing model, you must keep its tokeniser. The model's embedding table is indexed by that exact vocabulary. Swapping it makes every learned weight meaningless.

The five stages

A tokeniser is not one thing. It is a pipeline, and each stage changes the answer.

   raw text
      |
      v
  1. NORMALISER      clean-up: lowercase, unicode fixes
      |
      v
  2. PRE-TOKENISER   the boundaries no chunk may cross
      |
      v
  3. MODEL           the actual chunking: BPE, WordPiece, Unigram
      |
      v
  4. POST-PROCESSOR  bolt on the start and end markers
      |
      v
   token ids

Decoding runs the reverse, with a fifth component: the decoder, which knows how to turn Ġthe back into the.

What each stage decides

Normaliser. Should uppercase and lowercase be the same token? Should accents survive? Lowercasing halves your vocabulary pressure and destroys information a name-recognition task needs.

Pre-tokeniser. Where are chunks forbidden to cross? Splitting on whitespace means no token spans two words. Splitting digits individually means numbers behave predictably.

Model. BPE, WordPiece or Unigram, covered in the earlier lessons.

Post-processor. Which special markers get added automatically, and where.

Decoder. Must match the pre-tokeniser. A mismatch produces text with spaces in the wrong places.

The rule to remember

Every one of these must be identical at training time and at serving time. A tokeniser is not a preprocessing convenience. It is part of the model.

Where you have seen this

  • A code assistant that handles your indentation sensibly.
  • A model card listing "32,000 token vocabulary trained on the pretraining corpus".
  • A biomedical model tokenising drug names as single units.
  • A model in your language costing far fewer tokens than a general model does.

Remember this

  • Five stages: normaliser, pre-tokeniser, model, post-processor, decoder.
  • Each stage changes the output, and all five must match between training and serving.
  • Train your own only when pretraining, or when your domain or language is genuinely underserved.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install tokenizers "transformers==5.6.2"

Written against tokenizers 0.22.2 and transformers 5.6.2. Nothing is downloaded; the corpus is inline.

The whole pipeline, assembled stage by stage

train_tokenizer.py
from tokenizers import (Tokenizer, models, normalizers, pre_tokenizers,
                        processors, trainers, decoders)
from transformers import PreTrainedTokenizerFast

CORPUS = [
    "the cat sat on the mat", "the dog sat on the log",
    "a cat and a dog met on the road", "cats and dogs sat and sat",
    "the road was long and the cat was slow", "the mat was red",
    "a red mat and a brown log", "the slow dog met the fast cat",
] * 6

# 1. the model: what a token is, and how text is cut into them
tok = Tokenizer(models.BPE(unk_token="<unk>"))

# 2. the normaliser: text clean-up that happens before anything else
tok.normalizer = normalizers.Sequence([normalizers.NFKC(), normalizers.Lowercase()])

# 3. the pre-tokeniser: the boundaries no merge is ever allowed to cross
tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=True)

# 4. the decoder: how to turn tokens back into text
tok.decoder = decoders.ByteLevel()

trainer = trainers.BpeTrainer(
    vocab_size=320,
    special_tokens=["<pad>", "<unk>", "<s>", "</s>"],   # ids 0,1,2,3 in this order
    initial_alphabet=pre_tokenizers.ByteLevel.alphabet(),
    show_progress=False,
)
tok.train_from_iterator(CORPUS, trainer)

# 5. the post-processor: which special tokens get bolted on automatically
tok.post_processor = processors.TemplateProcessing(
    single="<s> $A </s>",
    pair="<s> $A </s> <s> $B </s>",
    special_tokens=[("<s>", tok.token_to_id("<s>")), ("</s>", tok.token_to_id("</s>"))],
)

print("vocabulary size:", tok.get_vocab_size())
enc = tok.encode("The CAT sat on the MAT")
print("\ninput : 'The CAT sat on the MAT'")
print("tokens:", enc.tokens)
print("ids   :", enc.ids)
print("decoded:", repr(tok.decode(enc.ids, skip_special_tokens=True)))

print("\nthe normaliser lower-cased before anything else, so these are identical:")
print(" ", tok.encode("the cat").ids)
print(" ", tok.encode("THE CAT").ids)

tok.save("my_tokenizer.json")
print("\nsaved to my_tokenizer.json")

fast = PreTrainedTokenizerFast(
    tokenizer_file="my_tokenizer.json",
    unk_token="<unk>", pad_token="<pad>", bos_token="<s>", eos_token="</s>",
)
batch = fast(["the cat sat", "a dog"], padding=True, return_tensors=None)
print("\nwrapped for transformers:")
print("  input_ids     :", batch["input_ids"])
print("  attention_mask:", batch["attention_mask"])
print("  pad token id  :", fast.pad_token_id)
Output
vocabulary size: 305

input : 'The CAT sat on the MAT'
tokens: ['<s>', 'Ġthe', 'Ġcat', 'Ġsat', 'Ġon', 'Ġthe', 'Ġmat', '</s>']
ids   : [2, 263, 270, 277, 281, 263, 283, 3]
decoded: ' the cat sat on the mat'

the normaliser lower-cased before anything else, so these are identical:
  [2, 263, 270, 3]
  [2, 263, 270, 3]

saved to my_tokenizer.json

wrapped for transformers:
  input_ids     : [[2, 263, 270, 277, 3], [2, 264, 276, 3, 0]]
  attention_mask: [[1, 1, 1, 1, 1], [1, 1, 1, 1, 0]]
  pad token id  : 0

Reading that output

305, not 320. The trainer stopped short because the corpus ran out of useful merges. Always print get_vocab_size() after training instead of trusting the argument you passed.

Every word became one token, and every one carries Ġ. With add_prefix_space=True the first word gets a leading space too, which is why Ġthe appears at the start.

'The CAT' and 'the cat' give the same ids. The normaliser lower-cased first, so case is gone before the model ever sees the text. That is a choice: it makes the vocabulary go further and it makes case-sensitive tasks impossible. Delete Lowercase() for a case-preserving tokeniser.

The decoded text has a leading space that the input did not. add_prefix_space=True inserted it, and the decoder faithfully returns it. If exact round-tripping matters, set add_prefix_space=False.

The padded batch shows the pieces fitting together. The short sequence is filled to length 5 with id 0, and its mask ends in 0. Attention will ignore that position.

Choosing each stage

Normaliser. NFKC folds compatibility characters and is nearly always wanted. Lowercase roughly halves the effective vocabulary pressure and destroys case, so keep it for search and drop it for named entities and code. StripAccents is usually a mistake outside specific retrieval settings.

Pre-tokeniser. ByteLevel for general text and code. Add Digits(individual_digits=True) if arithmetic matters. Whitespace when you want a simple, readable tokeniser and coverage is not a concern. Chain them with pre_tokenizers.Sequence([...]), and note the order changes the result.

Model. BPE for generative models, WordPiece for BERT-style encoders, Unigram when you want segmentation probabilities.

Post-processor. TemplateProcessing covers most needs. Get the special_tokens ids right or the template silently inserts the wrong id.

Decoder. Must mirror the pre-tokeniser exactly: ByteLevel with ByteLevel, WordPiece with WordPiece, Metaspace with Metaspace.

Training from files instead of memory

train_from_iterator takes any Python iterable, so a generator over files works without loading everything:

train_from_files.py
from pathlib import Path

def lines(paths):
    for p in paths:
        with open(p, encoding="utf-8") as f:
            for line in f:
                line = line.strip()
                if line:
                    yield line

# tok.train_from_iterator(lines(Path("corpus").glob("*.txt")), trainer)

There is also tok.train([...]) taking file paths directly, which is faster because the Rust side reads them in parallel. Use the iterator form when you need filtering or when your data is not already on disk as plain text.

Common mistakes

Changing a finetuned model's tokeniser. The embedding table is indexed by vocabulary id. A different tokeniser makes every id mean something else, and the model produces confident nonsense. Keep the tokeniser, or add tokens and resize, which has its own lesson.

A decoder that does not match the pre-tokeniser. ByteLevel pre-tokenising with a WordPiece decoder returns text littered with Ġ. It looks like a bug in your model and it is a two-line config error.

Forgetting initial_alphabet with ByteLevel. Without it, the alphabet comes from your corpus alone, and you lose the guarantee that any input is representable.

Special tokens in the wrong order. special_tokens=["<pad>", "<unk>", "<s>", "</s>"] fixes ids 0 to 3. Reorder that list between two training runs and two checkpoints become incompatible in a way that is very hard to debug.

Training on a sample that does not match the model's data. The tokeniser should see a representative slice of the pretraining corpus, including the code, the languages and the formatting the model will meet.

Try it yourself

Train the same corpus twice, once with Lowercase() and once without, and compare get_vocab_size() and the token count for a paragraph containing names. Then add pre_tokenizers.Digits(individual_digits=True) and re-encode "invoice 4429183". Each of those three changes is one line, and each one visibly changes the tokeniser.

What to learn next

Researcher — Mathematics and papers.

The pipeline as a contract

The huggingface/tokenizers architecture formalises a tokeniser as an ordered composition:

$$ \text{ids} = \mathrm{Post} \circ \mathrm{Model} \circ \mathrm{Pre} \circ \mathrm{Norm}\,(\text{text}) $$

with a Decoder providing an approximate inverse of Pre and Model. Serialisation to a single tokenizer.json captures all five components plus the vocabulary and merges, which is what makes a tokeniser reproducible across languages and runtimes.

The decoder is only an approximate inverse in general. Exact round-tripping requires the normaliser to be injective, which Lowercase, StripAccents and NFKC are not. Byte-level pre-tokenisation with an identity normaliser is invertible.

Pre-tokenisation is a hard constraint on the learned vocabulary

Because merges are computed within pre-token boundaries, the pre-tokeniser determines the space of learnable tokens before the algorithm runs. This is the most consequential of the five choices and the least discussed.

Liu et al. (2025), SuperBPE, demonstrate the cost. Training in two phases — merges restricted to within pre-tokens, then unrestricted — yields up to 33 percent fewer tokens at 200k vocabulary and, at matched training compute, +4.0 points average across 30 downstream tasks with +8.2 on MMLU, at 27 percent less inference compute. The whitespace boundary was carrying a real efficiency cost.

Schmidt et al. (2024) study pre-tokenisation choices systematically and report that the split pattern affects downstream quality independently of the merge algorithm.

Training data selection

A tokeniser trained on a distribution different from the pretraining corpus produces high fertility on the corpus that matters. Concretely:

  • Domain. A general-web tokeniser fragments chemical formulae, ICD codes, legal citations and identifiers.
  • Language mixture. Tokeniser training data should reflect the intended serving mixture, not the raw corpus mixture, since the two differ when a corpus is dominated by one language.
  • Code and formatting. Indentation, brackets and camelCase identifiers need representation, or code becomes expensive.

Ali et al. (2024) measure the effect of tokeniser choice on multilingual model training and find fertility differences translate directly into training-compute differences for equivalent content.

Evaluation without training a model

Tokenisers can be compared before any model exists:

  • Fertility: tokens per word on held-out text, per language and per domain. Lower is better, all else equal.
  • Proportion of continued words: the fraction of words split into more than one piece.
  • Compression ratio: bytes per token.
  • Coverage: the fraction of held-out characters representable without fallback.
  • Morphological alignment: agreement between boundaries and gold morpheme boundaries, where an analyser exists.

These correlate with downstream quality but do not determine it. Rust et al. (2021) find that a language-specific tokeniser improves downstream performance over a shared multilingual one even when model size and data are held fixed, and attribute part of the gain to fertility.

Vocabulary size

Covered in its own lesson. The short version: fewer tokens per document trades against a larger embedding and output layer, and the optimum depends on model size, corpus language mixture, and whether the embedding is tied.

Reproducibility notes

  • Special-token order fixes ids. Record it.
  • The tokenizers trainers are deterministic given the same corpus order, the same version and the same arguments. Version bumps have changed tie-breaking in the past; pin the version in your environment file.
  • tokenizer.json is the artefact to version, not the training script. It fully determines behaviour.

Papers

What to learn next

What to learn next

These follow on from what you just read.

  • Tokeniser Internals

    Choosing a vocabulary size

    A bigger vocabulary means shorter sequences and a fatter embedding table. Here is the tradeoff measured on real numbers, including the point where it stops helping.

  • Tokeniser Internals

    Special tokens

    A handful of vocabulary ids stand for no text at all. They mark where a sequence begins, where it ends, and which positions are filler the model must ignore.

  • Tokeniser Internals

    Chat templates

    A model never sees a list of messages. A template flattens the conversation into one string with role markers, and using the wrong template quietly ruins quality.