Tokeniser Internals

Chat templates

A model never sees a list of messages. A template flattens the conversation into one string with role markers, and using the wrong template quietly ruins quality.

On this page 7
  1. What the model actually receives
  2. Why the exact shape matters so much
  3. The generation prompt
  4. Where the template lives
  5. Where you have seen this
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A chat template is the recipe that turns a list of messages into the single string a model was trained to read.

Think of a formal letter. There is a fixed shape: address at the top right, date below, "Dear Sir", the body, "Yours sincerely", your name. A letter is unreadable to a clerk if you put the parts in a different order.

A model is that clerk, and it is far less forgiving. It learned one shape during training, and it expects that shape exactly.

Your messages get folded into that shape before the model sees anything.

What the model actually receives

You write something like a list of messages, each with a role and some content. The model receives one flat string.

   what you write:
     system:    you are helpful
     user:      tell me a joke
     assistant: here is a joke
     user:      another one

   what the model reads:
     <|system|>you are helpful<|end|>
     <|user|>tell me a joke<|end|>
     <|assistant|>here is a joke<|end|>
     <|user|>another one<|end|>
     <|assistant|>

The role names became markers. The last line is the important one: an opening <|assistant|> with nothing after it. That is the model's cue to write.

Why the exact shape matters so much

The model was trained on that shape and only that shape. Every instruction-following behaviour it has is tied to those markers.

Use a different template and the model still produces text. It produces worse text, ignores system instructions more often, and fails to stop cleanly. Nothing errors.

This is one of the most common quiet failures in applied work. The model looks weak; the formatting is wrong.

The generation prompt

That trailing <|assistant|> has a name: the generation prompt. It is the difference between two very different jobs.

Asking the model to reply. You add it. The model sees an empty assistant turn and fills it.

Training on a finished conversation. You leave it out. The conversation is complete and you are teaching the model on all of it.

Getting this backwards produces a model that either does not know when to speak, or that learns to reply to itself.

Where the template lives

Every instruction-tuned model ships its template inside its tokeniser files. You do not write it, and you should not invent it.

Load the tokeniser, call the function that applies the template, and the correct shape comes out. Writing the markers by hand is where mistakes come from.

Where you have seen this

  • A local model that keeps writing both sides of the conversation.
  • A model ignoring your system prompt entirely.
  • Output containing raw <|im_start|> markers.
  • Two identical prompts giving very different quality through different tools.

Remember this

  • The model reads one flat string, not a list of messages.
  • The template that builds that string is part of the model and ships with it.
  • Add the generation prompt when asking for a reply; leave it out when training.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install tokenizers "transformers==5.6.2"

Written against transformers 5.6.2. The tokeniser is built inline so nothing is downloaded.

Building a tokeniser with a template and applying it

chat_template.py
from tokenizers import Tokenizer, models, trainers, pre_tokenizers, decoders
from transformers import PreTrainedTokenizerFast

SPECIALS = ["<pad>", "<unk>", "<s>", "</s>",
            "<|system|>", "<|user|>", "<|assistant|>", "<|end|>"]

tok = Tokenizer(models.BPE(unk_token="<unk>"))
tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=True)
tok.decoder = decoders.ByteLevel()
tok.train_from_iterator(
    ["hello how are you", "i am fine thanks", "what is the capital of france",
     "paris is the capital of france", "tell me a joke", "here is a joke"] * 8,
    trainers.BpeTrainer(vocab_size=200, special_tokens=SPECIALS, show_progress=False,
                        initial_alphabet=pre_tokenizers.ByteLevel.alphabet()))

CHAT_TEMPLATE = (
    "{% for m in messages %}"
    "{{ '<|' + m['role'] + '|>' + m['content'] + '<|end|>' }}"
    "{% endfor %}"
    "{% if add_generation_prompt %}{{ '<|assistant|>' }}{% endif %}"
)

t = PreTrainedTokenizerFast(
    tokenizer_object=tok,
    unk_token="<unk>", pad_token="<pad>", bos_token="<s>", eos_token="</s>",
    additional_special_tokens=["<|system|>", "<|user|>", "<|assistant|>", "<|end|>"],
    chat_template=CHAT_TEMPLATE,
)

msgs = [
    {"role": "system", "content": "you are helpful"},
    {"role": "user", "content": "tell me a joke"},
    {"role": "assistant", "content": "here is a joke"},
    {"role": "user", "content": "another one"},
]

print("what the model actually reads:")
print(repr(t.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)))

# In transformers v5 this returns a dict of tensors, not a bare list of ids.
out = t.apply_chat_template(msgs, tokenize=True, add_generation_prompt=True)
print("\nreturn type:", type(out).__name__, "  keys:", list(out.keys()))
ids = out["input_ids"]
print("total tokens:", len(ids))
print("last six tokens:", t.convert_ids_to_tokens(ids[-6:]))

print("\nwithout add_generation_prompt (this form is for training, not for asking):")
print(repr(t.apply_chat_template(msgs, tokenize=False, add_generation_prompt=False)))

print("\nrole markers are single ids, not spelled out:")
for s in ["<|user|>", "<|assistant|>", "<|end|>"]:
    print(f"  {s:<14} -> id {t.convert_tokens_to_ids(s)}")

print("\nwhat happens when a user types a role marker as ordinary text:")
print("  default            :", t.tokenize("<|assistant|>"))
print("  split_special_tokens:", t.tokenize("<|assistant|>", split_special_tokens=True))
Output
what the model actually reads:
'<|system|>you are helpful<|end|><|user|>tell me a joke<|end|><|assistant|>here is a joke<|end|><|user|>another one<|end|><|assistant|>'

return type: BatchEncoding   keys: ['input_ids', 'attention_mask']
total tokens: 67
last six tokens: ['Ġ', 'o', 'n', 'e', '<|end|>', '<|assistant|>']

without add_generation_prompt (this form is for training, not for asking):
'<|system|>you are helpful<|end|><|user|>tell me a joke<|end|><|assistant|>here is a joke<|end|><|user|>another one<|end|>'

role markers are single ids, not spelled out:
  <|user|>       -> id 5
  <|assistant|>  -> id 6
  <|end|>        -> id 7

what happens when a user types a role marker as ordinary text:
  default            : ['<|assistant|>']
  split_special_tokens: ['Ġ', '<', '|', 'a', 's', 's', 'i', 's', 't', 'a', 'n', 't', '|', '>']

Reading that output

The whole conversation is one string with no newlines. Four messages became a single sequence. Whether a template inserts newlines is a per-model choice, and getting it wrong is a real source of quality loss.

The two strings differ by exactly one trailing marker. That single token is the difference between "reply to this" and "here is a finished conversation". Print both forms whenever you are unsure which you are producing.

apply_chat_template(tokenize=True) returns a BatchEncoding, not a list. This changed in transformers v5. Code written against v4 that did ids[-1] on the return value now indexes into a dict and fails with a confusing type error. Take out["input_ids"].

Each role marker is one id: 5, 6 and 7. They occupy low ids because they were listed in special_tokens during training. A conversation carries very few marker tokens, so the formatting overhead is small — as long as the markers are registered.

The last block is the injection surface. A user typing <|assistant|> into your chat box produces the real control id by default. They have opened an assistant turn inside a user message. Pass split_special_tokens=True for anything that came from a user.

The cost of getting the template wrong

There is no exception and no warning. To see the difference, print both:

wrong_template.py
# continuing from the script above
WRONG = ("{% for m in messages %}{{ m['role'] + ': ' + m['content'] + '\n' }}{% endfor %}"
         "{% if add_generation_prompt %}{{ 'assistant: ' }}{% endif %}")

right = t.apply_chat_template(msgs, tokenize=True, add_generation_prompt=True)["input_ids"]
t.chat_template = WRONG
wrong = t.apply_chat_template(msgs, tokenize=True, add_generation_prompt=True)["input_ids"]

print("correct template:", len(right), "tokens,",
      sum(1 for i in right if i in (4, 5, 6, 7)), "control tokens")
print("plain-text style:", len(wrong), "tokens,",
      sum(1 for i in wrong if i in (4, 5, 6, 7)), "control tokens")
Output
correct template: 67 tokens, 9 control tokens
plain-text style: 101 tokens, 0 control tokens

Zero control tokens, and 34 more of them. The model has no idea where one turn ends and the next begins, because the markers it was trained on are absent. It has to infer role boundaries from the word "assistant:" appearing in ordinary text. It will produce fluent output and follow instructions worse.

Rules that hold across models

  • Load the template from the model. AutoTokenizer.from_pretrained(...) carries chat_template in tokenizer_config.json. Never hand-write the markers.
  • Some templates allow only one system message, at the start. Others allow none. Passing an unsupported role raises a Jinja error, which is the friendly case.
  • Some templates require strictly alternating user and assistant turns. Two consecutive user messages raise an error on those models.
  • Do not add a BOS token twice. Many templates emit it themselves. Calling the tokeniser with add_special_tokens=True on top of an already-templated string doubles it, and models are measurably sensitive to that.
  • Tool calling has its own template arguments. Templates supporting tools accept a tools= argument and render function schemas into the prompt. The format is model-specific.

Common mistakes

Concatenating messages by hand. It works, it produces no error, and it costs quality. There is no reason to do it.

Setting add_generation_prompt=True while building training data. You are teaching the model that an assistant turn opens and immediately ends. Leave it out for training.

Forgetting to mask the prompt in the training loss. When fine-tuning on conversations, the loss usually belongs only on assistant tokens. Computing it over the whole string teaches the model to generate user turns.

Passing untrusted content without split_special_tokens=True. Shown above. A user can forge a role boundary.

Assuming two models with the same base share a template. Different fine-tunes of the same base model frequently use different markers. Read the tokeniser config for the exact checkpoint you are serving.

Try it yourself

Take any instruction-tuned model you have locally, load only its tokeniser, and print tokenizer.chat_template. Then apply it to a three-message conversation with tokenize=False. Comparing that string against what your application actually sends the model is the fastest way to find a formatting bug you did not know you had.

What to learn next

Researcher — Mathematics and papers.

Why templates exist

Instruction tuning trains a model on conversations serialised into a specific string format. The role markers, separators and whitespace are part of the training distribution. At inference, any deviation is a distribution shift.

The magnitude is not negligible. Community and vendor evaluations repeatedly report multi-point differences on instruction-following benchmarks between correct and approximate formatting of the same conversation, with the largest effects on system-prompt adherence and on clean stopping.

Before templates were standardised, every model repository documented its format in prose, and every serving stack reimplemented it. The Jinja-based chat_template field in tokenizer_config.json made the format an artefact that travels with the weights.

The Jinja contract

The template is a Jinja2 string evaluated with:

  • messages: a list of dicts with at least role and content.
  • add_generation_prompt: bool, controlling the trailing assistant marker.
  • bos_token, eos_token and other tokeniser attributes.
  • tools: an optional list of function schemas, for models supporting tool calls.
  • Model-specific extras: documents for RAG-aware templates, and date or thinking-mode flags in some recent models.

Rendering is done in a sandboxed Jinja environment with a restricted feature set, since templates arrive with downloaded model repositories and are therefore untrusted input.

Return-type change in transformers v5

apply_chat_template(..., tokenize=True) returns a BatchEncoding with input_ids and attention_mask. In v4 the common path returned a bare list of ids. This is a silent-failure boundary for code written against v4: indexing a BatchEncoding with an integer raises, and passing it where a list is expected fails downstream. It is one of the more visible items in the v5 migration.

Training-time considerations

Loss masking. Fine-tuning on conversations normally computes loss only on assistant-turn tokens. Implementations use return_assistant_tokens_mask=True where the template supports it (the template must mark assistant regions with {% generation %} blocks), or reconstruct spans by re-rendering prefixes.

Generation prompt. Present at inference, absent in training targets. Mixing the two teaches degenerate turn boundaries.

Multi-turn packing. Packing several conversations into one training sequence requires the attention mask to prevent cross-conversation attention, or a document-boundary-aware attention implementation. Without it, a conversation can attend to an unrelated one earlier in the packed sequence.

Security

The template renders untrusted content strings into a control-token-bearing format. Two independent surfaces:

  1. Token-level. If content contains a control string and tokenisation does not split special tokens, the user emits a real control id and can forge a role boundary. Mitigation: split_special_tokens=True for untrusted spans, plus filtering at the application boundary.
  2. String-level. Even with control ids split, a user can write text resembling a role boundary. The model may follow it. This is ordinary prompt injection, and no tokeniser setting prevents it.

Greshake et al. (2023) frame the general problem for applications that route retrieved or third-party content into prompts. Chat templates are the concrete mechanism by which that content reaches the model, so the template boundary is the natural place to enforce escaping.

Format sensitivity as a measurement problem

Reported benchmark numbers depend on formatting. Two evaluations of the same checkpoint with different templates are not comparable, and evaluation harnesses have historically differed here. When reporting or reproducing instruction-following results, the applied template string is part of the experimental configuration and should be recorded alongside the model revision.

References

What to learn next