Tokeniser Internals

Counting tokens before you send them

Tokens are the unit you are billed in and the unit your context limit is measured in. Estimating from character counts is wrong by up to 90 percent.

Read these first

On this page 6
  1. Why guessing fails
  2. The overheads people forget
  3. Where the counting has to happen
  4. What to do with the number
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Counting tokens means running your text through the exact tokeniser the model uses. Any other method is a guess.

Think of checked luggage. There is a weight limit and a charge above it. You can lift the bag and estimate, or you can put it on the scale.

Estimating works until it does not, and then you pay at the counter. The scale costs nothing and takes a second.

Tokens are the same. Your context limit, your bill and your latency are all measured in tokens, and the scale is one function call.

Why guessing fails

The common rule of thumb is four characters per token. It is roughly right for plain English prose and wrong for nearly everything else.

Code is dense with punctuation, and punctuation tokenises badly. Numbers get chopped unpredictably. Any language not written in the Latin alphabet costs several times more.

The measurements later on this page show the rule of thumb missing by 13 percent on English and 90 percent on Hindi. On the same tokeniser, in the same script.

The overheads people forget

Your message is not the only thing being counted.

The chat wrapper. Role markers, turn separators and the generation prompt all cost tokens on every turn.

The system prompt. Sent with every single request, including the very short ones.

The conversation history. In a multi-turn chat, everything said so far is re-sent each time. The tenth message costs far more than the first.

Tool definitions. Function schemas are rendered into the prompt as text. A set of tools can run to hundreds of tokens per request.

The output. Generated tokens are usually billed at a higher rate than input tokens.

Where the counting has to happen

Every model family has its own tokeniser, and counts do not transfer. The same paragraph gives different totals for different models.

So counting means using the specific tokeniser for the specific model you are calling. Not a similar one, and not an average.

What to do with the number

Budget before you build. Count a typical request, multiply by your expected volume, and you have your monthly bill before writing the feature.

Truncate deliberately. When a document will not fit, cut it at a token boundary you chose, not wherever the API happens to stop.

Watch history growth. A chat that re-sends everything grows quadratically in total cost. Summarising old turns is the standard fix.

Remember this

  • Tokens are the unit of cost, of context and of latency.
  • Character-based estimates are wrong by up to 90 percent depending on the content.
  • Count the chat wrapper, the system prompt, the history and the tool definitions, not only your message.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install tokenizers "transformers==5.6.2"

Written against tokenizers 0.22.2 and transformers 5.6.2. The tokeniser below is trained inline so the script needs no network. For a real model you would load its own tokeniser instead, which is covered further down.

Measuring, rather than estimating

count_tokens.py
import random
from tokenizers import Tokenizer, models, trainers, pre_tokenizers, decoders
from transformers import PreTrainedTokenizerFast

WORDS = ("the a and of to in is was for with on at by from that this it as but "
         "delivery refund warranty invoice payment customer order shipping return "
         "account balance transfer statement").split()
rng = random.Random(0)
CORPUS = [" ".join(rng.choice(WORDS) for _ in range(14)) for _ in range(3000)]

tok = Tokenizer(models.BPE())
tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=False)
tok.decoder = decoders.ByteLevel()
ROLES = ["<pad>", "<|user|>", "<|assistant|>", "<|end|>"]
tok.train_from_iterator(CORPUS, trainers.BpeTrainer(
    vocab_size=600, show_progress=False, special_tokens=ROLES,
    initial_alphabet=pre_tokenizers.ByteLevel.alphabet()))
t = PreTrainedTokenizerFast(tokenizer_object=tok, pad_token="<pad>",
                            additional_special_tokens=ROLES[1:])

SAMPLES = {
    "plain English":      "the customer asked about the refund and the delivery",
    "code":               "for i in range(10): print(i * 2)",
    "a long number":      "invoice 4429183 for 128456 rupees",
    "Hindi (Devanagari)": "ग्राहक ने वापसी के बारे में पूछा",
    "punctuation soup":   "!!! ??? ... ---   <<<>>>",
}

print(f"{'sample':<20} {'chars':>6} {'tokens':>7} {'chars/token':>12}")
for name, s in SAMPLES.items():
    n = len(t(s)["input_ids"])
    print(f"{name:<20} {len(s):>6} {n:>7} {len(s) / n:>12.2f}")

print("\nthe 'four characters per token' rule of thumb, checked:")
for name, s in SAMPLES.items():
    n = len(t(s)["input_ids"])
    print(f"  {name:<20} guess {len(s) // 4:>3}, actual {n:>3}, "
          f"off by {abs(len(s) // 4 - n) / n:>6.1%}")

print("\ncounting a whole batch in one call:")
texts = list(SAMPLES.values())
enc = t(texts, return_length=True)
print("  lengths:", enc["length"], " total:", sum(enc["length"]))

print("\ncutting text to a budget without guessing:")
long_text = " ".join(CORPUS[:3])
cut = t(long_text, truncation=True, max_length=20, return_offsets_mapping=True)
print("  tokens kept:", len(cut["input_ids"]))
end = cut["offset_mapping"][-1][1]
print(f"  that is the first {end} characters:")
print("   ", repr(long_text[:end]))

print("\nthe overhead people forget: chat wrappers cost tokens too")
t.chat_template = ("{% for m in messages %}{{ '<|' + m['role'] + '|>' + m['content'] "
                   "+ '<|end|>' }}{% endfor %}"
                   "{% if add_generation_prompt %}{{ '<|assistant|>' }}{% endif %}")
msg = [{"role": "user", "content": SAMPLES["plain English"]}]
bare = len(t(SAMPLES["plain English"])["input_ids"])
wrapped = len(t.apply_chat_template(msg, tokenize=True,
                                    add_generation_prompt=True)["input_ids"])
print(f"  the message alone : {bare} tokens")
print(f"  inside a template : {wrapped} tokens  (+{wrapped - bare})")
Output
sample                chars  tokens  chars/token
plain English            52      15         3.47
code                     32      25         1.28
a long number            33      24         1.38
Hindi (Devanagari)       32      84         0.38
punctuation soup         24      24         1.00

the 'four characters per token' rule of thumb, checked:
  plain English        guess  13, actual  15, off by  13.3%
  code                 guess   8, actual  25, off by  68.0%
  a long number        guess   8, actual  24, off by  66.7%
  Hindi (Devanagari)   guess   8, actual  84, off by  90.5%
  punctuation soup     guess   6, actual  24, off by  75.0%

counting a whole batch in one call:
  lengths: [15, 25, 24, 84, 24]  total: 172

cutting text to a budget without guessing:
  tokens kept: 20
  that is the first 118 characters:
    'customer shipping and it statement order delivery transfer invoice from for but for is it with delivery is to warranty'

the overhead people forget: chat wrappers cost tokens too
  the message alone : 15 tokens
  inside a template : 18 tokens  (+3)

Reading that output

English is the only row the rule of thumb nearly fits, at 3.47 characters per token against a guess of 4, so 13 percent low. That is the best case.

Hindi costs 84 tokens for 32 characters, 0.38 characters per token. This tokeniser was trained on English only, so Devanagari falls to raw UTF-8 bytes at three bytes per character. A properly multilingual tokeniser does much better than this, and never reaches English parity.

Code and numbers are both around 1.3 characters per token. Punctuation and digits do not merge into long pieces, so both cost roughly one token per character.

The 90.5 percent error is the headline. Budget a Hindi-language product using the four-character rule and you are out by a factor of ten.

return_length=True counts a batch in one pass. Do not loop and call the tokeniser per item; it is slower and the batched form gives you the masks you need anyway.

Truncation with return_offsets_mapping tells you exactly where the cut fell. 20 tokens turned out to be 118 characters of that particular text. That mapping is how you truncate a document and still show the user precisely what was kept.

The chat wrapper added 3 tokens here because the role markers are registered as single ids. Leave them unregistered and the same wrapper costs 23 tokens, since <|user|> spells out character by character. Whether your markers are real vocabulary entries changes your overhead by an order of magnitude.

Counting for a specific model

The script above trains its own tokeniser so it runs offline. In production you count with the tokeniser belonging to the model you are calling:

python
# Hugging Face models. Downloads the tokeniser files on first use.
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("<model-id>")
n = len(tok(text)["input_ids"])

# OpenAI models. pip install tiktoken; downloads the encoding file on first use.
import tiktoken
enc = tiktoken.encoding_for_model("<model-name>")
n = len(enc.encode(text))

No output block for these, deliberately. Both fetch files on first use, counts differ per model and per version, and printing invented numbers here would teach you to expect totals that will not match what you see.

For providers exposing a token-counting endpoint, that endpoint is authoritative and accounts for their exact prompt formatting. Prefer it over any local reimplementation when billing accuracy matters.

A budgeting checklist

Count all of these, not only the user's message:

  • system prompt, sent on every request
  • tool and function schemas rendered into the prompt
  • retrieved context, which is usually the largest item in a RAG system
  • the full conversation history for multi-turn chats
  • chat template markers and the generation prompt
  • expected output length, usually priced higher than input

The item that surprises people is history. Re-sending an $n$-turn conversation on turn $n$ makes total tokens across a session grow with the square of turn count.

Common mistakes

Counting with a different model's tokeniser. Vocabularies differ, so counts differ. A number from one model family is not a valid estimate for another.

Counting the message and not the wrapper. Shown above. For short messages the wrapper can be a significant fraction of the total.

Truncating by characters. Slicing a string mid-token gives the model a partial byte sequence. Truncate with the tokeniser and max_length.

Ignoring output tokens in the budget. Generation is usually the more expensive half, and it is the half you control least.

Assuming a token count is stable across versions. Tokenisers do change between model releases. Re-measure when you upgrade.

Try it yourself

Take a real request from your own application: system prompt, tools, retrieved context, history and the user message. Count each part separately and print the breakdown. Almost everyone is surprised by which part dominates, and that breakdown is the map of where cost reduction is actually available.

What to learn next

Researcher — Mathematics and papers.

What a token count determines

Three separate quantities scale with token count, and they scale differently.

Cost. Provider pricing is per input token and per output token, with output typically priced two to five times higher. Cached or reused prefixes are often discounted, which changes the optimisation: a stable prefix followed by variable content can be substantially cheaper than the reverse.

Context occupancy. The window is measured in tokens, so fertility determines how much content fits. For a document of $W$ words and fertility $f$, occupancy is $fW$.

Latency. Prefill is roughly linear in prompt length for the feed-forward path and quadratic for attention without a KV cache hit. Decode is linear in output length. Prompt tokens and completion tokens therefore have different latency profiles, and a token-count budget that treats them identically will mispredict.

Fertility as the underlying variable

Define fertility $f = \text{tokens}/\text{word}$ on a held-out sample. It varies by:

  • Language. Petrov et al. (2023) and Ahia et al. (2023) measure the spread across languages under commercial tokenisers and find multiples, not percentages. The consequences are pricing, context and latency disparities that fall on speakers of under-served languages.
  • Domain. Code, chemical notation, legal citation and identifiers all tokenise worse than prose under a general vocabulary.
  • Formatting. JSON, markdown tables and indented code carry punctuation-heavy token sequences.

A single average fertility is not a useful operational number. Report it per language and per document class.

Quadratic growth of conversational cost

For a chat re-sending full history, let $t_i$ be the tokens in turn $i$. The prompt for turn $n$ carries $\sum_{i<n} t_i$ plus the system prompt $S$. Total input tokens across a session of $N$ turns is

$$ \sum_{n=1}^{N}\left(S + \sum_{i<n} t_i\right) = NS + \sum_{n=1}^{N}\sum_{i<n} t_i = NS + O(N^2 \bar{t}) $$

for mean turn length $\bar{t}$. Mitigations in decreasing order of usual impact: prompt caching on the stable prefix, summarising or dropping old turns, and retrieving relevant history rather than resending all of it.

Truncation and offset mapping

return_offsets_mapping=True returns, per token, the character span it covers in the original string. This supports:

  • exact truncation reporting, showing the user what was kept;
  • highlighting model-attended spans back onto the source text;
  • aligning token-level labels to character-level annotations, which is the standard requirement for sequence labelling.

Offsets are relative to the string after normalisation. A normaliser that changes length — NFKC on compatibility characters, whitespace collapsing — makes the mapping approximate. For exact alignment, use an identity normaliser.

Counting against a remote provider

Local reimplementation of a provider's prompt assembly is a recurring source of error, because the count depends on the exact serialisation of roles, tool schemas and metadata, and that serialisation is not always documented and does change. Where a provider offers a token-counting endpoint, it is authoritative. Where one does not, treat a local count as a lower bound and reconcile against returned usage metadata.

Papers

What to learn next