Chat templates
A model never sees a list of messages. A template flattens the conversation into one string with role markers, and using the wrong template quietly ruins quality.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A chat template is the recipe that turns a list of messages into the single string a model was trained to read.
Think of a formal letter. There is a fixed shape: address at the top right, date below, "Dear Sir", the body, "Yours sincerely", your name. A letter is unreadable to a clerk if you put the parts in a different order.
A model is that clerk, and it is far less forgiving. It learned one shape during training, and it expects that shape exactly.
Your messages get folded into that shape before the model sees anything.
What the model actually receives
You write something like a list of messages, each with a role and some content. The model receives one flat string.
what you write:
system: you are helpful
user: tell me a joke
assistant: here is a joke
user: another one
what the model reads:
<|system|>you are helpful<|end|>
<|user|>tell me a joke<|end|>
<|assistant|>here is a joke<|end|>
<|user|>another one<|end|>
<|assistant|>The role names became markers. The last line is the important one: an opening <|assistant|> with nothing after it. That is the model's cue to write.
Why the exact shape matters so much
The model was trained on that shape and only that shape. Every instruction-following behaviour it has is tied to those markers.
Use a different template and the model still produces text. It produces worse text, ignores system instructions more often, and fails to stop cleanly. Nothing errors.
This is one of the most common quiet failures in applied work. The model looks weak; the formatting is wrong.
The generation prompt
That trailing <|assistant|> has a name: the generation prompt. It is the difference between two very different jobs.
Asking the model to reply. You add it. The model sees an empty assistant turn and fills it.
Training on a finished conversation. You leave it out. The conversation is complete and you are teaching the model on all of it.
Getting this backwards produces a model that either does not know when to speak, or that learns to reply to itself.
Where the template lives
Every instruction-tuned model ships its template inside its tokeniser files. You do not write it, and you should not invent it.
Load the tokeniser, call the function that applies the template, and the correct shape comes out. Writing the markers by hand is where mistakes come from.
Where you have seen this
- A local model that keeps writing both sides of the conversation.
- A model ignoring your system prompt entirely.
- Output containing raw
<|im_start|>markers. - Two identical prompts giving very different quality through different tools.
Remember this
- The model reads one flat string, not a list of messages.
- The template that builds that string is part of the model and ships with it.
- Add the generation prompt when asking for a reply; leave it out when training.
What to learn next
- Prompt engineering — what goes inside the template.
- Prompt injection — the attack the template boundary has to survive.
- Function calling — the tool-calling half of a chat template.
Developer — Code and libraries.
Setup
pip install tokenizers "transformers==5.6.2"Written against transformers 5.6.2. The tokeniser is built inline so nothing is downloaded.
Building a tokeniser with a template and applying it
from tokenizers import Tokenizer, models, trainers, pre_tokenizers, decoders
from transformers import PreTrainedTokenizerFast
SPECIALS = ["<pad>", "<unk>", "<s>", "</s>",
"<|system|>", "<|user|>", "<|assistant|>", "<|end|>"]
tok = Tokenizer(models.BPE(unk_token="<unk>"))
tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=True)
tok.decoder = decoders.ByteLevel()
tok.train_from_iterator(
["hello how are you", "i am fine thanks", "what is the capital of france",
"paris is the capital of france", "tell me a joke", "here is a joke"] * 8,
trainers.BpeTrainer(vocab_size=200, special_tokens=SPECIALS, show_progress=False,
initial_alphabet=pre_tokenizers.ByteLevel.alphabet()))
CHAT_TEMPLATE = (
"{% for m in messages %}"
"{{ '<|' + m['role'] + '|>' + m['content'] + '<|end|>' }}"
"{% endfor %}"
"{% if add_generation_prompt %}{{ '<|assistant|>' }}{% endif %}"
)
t = PreTrainedTokenizerFast(
tokenizer_object=tok,
unk_token="<unk>", pad_token="<pad>", bos_token="<s>", eos_token="</s>",
additional_special_tokens=["<|system|>", "<|user|>", "<|assistant|>", "<|end|>"],
chat_template=CHAT_TEMPLATE,
)
msgs = [
{"role": "system", "content": "you are helpful"},
{"role": "user", "content": "tell me a joke"},
{"role": "assistant", "content": "here is a joke"},
{"role": "user", "content": "another one"},
]
print("what the model actually reads:")
print(repr(t.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)))
# In transformers v5 this returns a dict of tensors, not a bare list of ids.
out = t.apply_chat_template(msgs, tokenize=True, add_generation_prompt=True)
print("\nreturn type:", type(out).__name__, " keys:", list(out.keys()))
ids = out["input_ids"]
print("total tokens:", len(ids))
print("last six tokens:", t.convert_ids_to_tokens(ids[-6:]))
print("\nwithout add_generation_prompt (this form is for training, not for asking):")
print(repr(t.apply_chat_template(msgs, tokenize=False, add_generation_prompt=False)))
print("\nrole markers are single ids, not spelled out:")
for s in ["<|user|>", "<|assistant|>", "<|end|>"]:
print(f" {s:<14} -> id {t.convert_tokens_to_ids(s)}")
print("\nwhat happens when a user types a role marker as ordinary text:")
print(" default :", t.tokenize("<|assistant|>"))
print(" split_special_tokens:", t.tokenize("<|assistant|>", split_special_tokens=True))what the model actually reads: '<|system|>you are helpful<|end|><|user|>tell me a joke<|end|><|assistant|>here is a joke<|end|><|user|>another one<|end|><|assistant|>' return type: BatchEncoding keys: ['input_ids', 'attention_mask'] total tokens: 67 last six tokens: ['Ġ', 'o', 'n', 'e', '<|end|>', '<|assistant|>'] without add_generation_prompt (this form is for training, not for asking): '<|system|>you are helpful<|end|><|user|>tell me a joke<|end|><|assistant|>here is a joke<|end|><|user|>another one<|end|>' role markers are single ids, not spelled out: <|user|> -> id 5 <|assistant|> -> id 6 <|end|> -> id 7 what happens when a user types a role marker as ordinary text: default : ['<|assistant|>'] split_special_tokens: ['Ġ', '<', '|', 'a', 's', 's', 'i', 's', 't', 'a', 'n', 't', '|', '>']
Reading that output
The whole conversation is one string with no newlines. Four messages became a single sequence. Whether a template inserts newlines is a per-model choice, and getting it wrong is a real source of quality loss.
The two strings differ by exactly one trailing marker. That single token is the difference between "reply to this" and "here is a finished conversation". Print both forms whenever you are unsure which you are producing.
apply_chat_template(tokenize=True) returns a BatchEncoding, not a list. This changed in transformers v5. Code written against v4 that did ids[-1] on the return value now indexes into a dict and fails with a confusing type error. Take out["input_ids"].
Each role marker is one id: 5, 6 and 7. They occupy low ids because they were listed in special_tokens during training. A conversation carries very few marker tokens, so the formatting overhead is small — as long as the markers are registered.
The last block is the injection surface. A user typing <|assistant|> into your chat box produces the real control id by default. They have opened an assistant turn inside a user message. Pass split_special_tokens=True for anything that came from a user.
The cost of getting the template wrong
There is no exception and no warning. To see the difference, print both:
# continuing from the script above
WRONG = ("{% for m in messages %}{{ m['role'] + ': ' + m['content'] + '\n' }}{% endfor %}"
"{% if add_generation_prompt %}{{ 'assistant: ' }}{% endif %}")
right = t.apply_chat_template(msgs, tokenize=True, add_generation_prompt=True)["input_ids"]
t.chat_template = WRONG
wrong = t.apply_chat_template(msgs, tokenize=True, add_generation_prompt=True)["input_ids"]
print("correct template:", len(right), "tokens,",
sum(1 for i in right if i in (4, 5, 6, 7)), "control tokens")
print("plain-text style:", len(wrong), "tokens,",
sum(1 for i in wrong if i in (4, 5, 6, 7)), "control tokens")correct template: 67 tokens, 9 control tokens plain-text style: 101 tokens, 0 control tokens
Zero control tokens, and 34 more of them. The model has no idea where one turn ends and the next begins, because the markers it was trained on are absent. It has to infer role boundaries from the word "assistant:" appearing in ordinary text. It will produce fluent output and follow instructions worse.
Rules that hold across models
- Load the template from the model.
AutoTokenizer.from_pretrained(...)carrieschat_templateintokenizer_config.json. Never hand-write the markers. - Some templates allow only one system message, at the start. Others allow none. Passing an unsupported role raises a Jinja error, which is the friendly case.
- Some templates require strictly alternating user and assistant turns. Two consecutive user messages raise an error on those models.
- Do not add a BOS token twice. Many templates emit it themselves. Calling the tokeniser with
add_special_tokens=Trueon top of an already-templated string doubles it, and models are measurably sensitive to that. - Tool calling has its own template arguments. Templates supporting tools accept a
tools=argument and render function schemas into the prompt. The format is model-specific.
Common mistakes
Concatenating messages by hand. It works, it produces no error, and it costs quality. There is no reason to do it.
Setting add_generation_prompt=True while building training data. You are teaching the model that an assistant turn opens and immediately ends. Leave it out for training.
Forgetting to mask the prompt in the training loss. When fine-tuning on conversations, the loss usually belongs only on assistant tokens. Computing it over the whole string teaches the model to generate user turns.
Passing untrusted content without split_special_tokens=True. Shown above. A user can forge a role boundary.
Assuming two models with the same base share a template. Different fine-tunes of the same base model frequently use different markers. Read the tokeniser config for the exact checkpoint you are serving.
Try it yourself
Take any instruction-tuned model you have locally, load only its tokeniser, and print tokenizer.chat_template. Then apply it to a three-message conversation with tokenize=False. Comparing that string against what your application actually sends the model is the fastest way to find a formatting bug you did not know you had.
What to learn next
- Prompt engineering — what goes inside the template.
- Prompt injection — the attack the template boundary has to survive.
- Function calling — the tool-calling half of a chat template.
Researcher — Mathematics and papers.
Why templates exist
Instruction tuning trains a model on conversations serialised into a specific string format. The role markers, separators and whitespace are part of the training distribution. At inference, any deviation is a distribution shift.
The magnitude is not negligible. Community and vendor evaluations repeatedly report multi-point differences on instruction-following benchmarks between correct and approximate formatting of the same conversation, with the largest effects on system-prompt adherence and on clean stopping.
Before templates were standardised, every model repository documented its format in prose, and every serving stack reimplemented it. The Jinja-based chat_template field in tokenizer_config.json made the format an artefact that travels with the weights.
The Jinja contract
The template is a Jinja2 string evaluated with:
messages: a list of dicts with at leastroleandcontent.add_generation_prompt: bool, controlling the trailing assistant marker.bos_token,eos_tokenand other tokeniser attributes.tools: an optional list of function schemas, for models supporting tool calls.- Model-specific extras:
documentsfor RAG-aware templates, and date or thinking-mode flags in some recent models.
Rendering is done in a sandboxed Jinja environment with a restricted feature set, since templates arrive with downloaded model repositories and are therefore untrusted input.
Return-type change in transformers v5
apply_chat_template(..., tokenize=True) returns a BatchEncoding with input_ids and attention_mask. In v4 the common path returned a bare list of ids. This is a silent-failure boundary for code written against v4: indexing a BatchEncoding with an integer raises, and passing it where a list is expected fails downstream. It is one of the more visible items in the v5 migration.
Training-time considerations
Loss masking. Fine-tuning on conversations normally computes loss only on assistant-turn tokens. Implementations use return_assistant_tokens_mask=True where the template supports it (the template must mark assistant regions with {% generation %} blocks), or reconstruct spans by re-rendering prefixes.
Generation prompt. Present at inference, absent in training targets. Mixing the two teaches degenerate turn boundaries.
Multi-turn packing. Packing several conversations into one training sequence requires the attention mask to prevent cross-conversation attention, or a document-boundary-aware attention implementation. Without it, a conversation can attend to an unrelated one earlier in the packed sequence.
Security
The template renders untrusted content strings into a control-token-bearing format. Two independent surfaces:
- Token-level. If
contentcontains a control string and tokenisation does not split special tokens, the user emits a real control id and can forge a role boundary. Mitigation:split_special_tokens=Truefor untrusted spans, plus filtering at the application boundary. - String-level. Even with control ids split, a user can write text resembling a role boundary. The model may follow it. This is ordinary prompt injection, and no tokeniser setting prevents it.
Greshake et al. (2023) frame the general problem for applications that route retrieved or third-party content into prompts. Chat templates are the concrete mechanism by which that content reaches the model, so the template boundary is the natural place to enforce escaping.
Format sensitivity as a measurement problem
Reported benchmark numbers depend on formatting. Two evaluations of the same checkpoint with different templates are not comparable, and evaluation harnesses have historically differed here. When reporting or reproducing instruction-following results, the applied template string is part of the experimental configuration and should be recorded alongside the model revision.
References
- Hugging Face chat templating documentation — huggingface.co/docs/transformers/chat_templating
- transformers v5 migration guide — github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md
- Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, AISec 2023 — arxiv.org/abs/2302.12173
- Ouyang et al., Training language models to follow instructions with human feedback, NeurIPS 2022 — arxiv.org/abs/2203.02155
What to learn next
- Prompt engineering — what goes inside the template.
- Prompt injection — the attack the template boundary has to survive.
- Function calling — the tool-calling half of a chat template.