Tokeniser Internals

Byte-level BPE

Convert text to raw bytes first, then run BPE. With 256 starting symbols the tokeniser can represent literally anything, and unknown tokens become impossible.

On this page 6
  1. The problem this solves
  2. The trick that makes it work
  3. What it costs
  4. Where you have seen this
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Byte-level BPE turns text into raw bytes before doing anything else, so every possible input is already covered.

Think of a postal system that will carry any parcel, because it stopped caring what is inside. It only handles boxes of a standard size. Whatever you want to send, you pack it into those boxes.

Text has an equivalent standard box: the byte. Every character on Earth, in every script, plus every emoji, is already stored on your computer as a sequence of bytes.

There are exactly 256 possible bytes. Start a tokeniser from those and there is nothing it cannot read.

The problem this solves

Character-based tokenisers have a coverage problem. Train on English text and the tokeniser learns the Latin alphabet. Feed it Tamil, or Arabic, or an emoji, and those characters are not in its list.

The usual result is the unknown token: a single placeholder meaning "something was here and it is gone". Anything replaced by it cannot be recovered.

Unicode has over a hundred thousand assigned characters. You cannot start a tokeniser from all of them. You can start from all 256 bytes.

The trick that makes it work

Bytes are numbers. Most of them are not printable characters, and some of them are spaces, tabs and newlines. Feeding those directly into a text pipeline causes chaos.

GPT-2's solution is a lookup table that gives every byte a visible, unique, non-whitespace character to stand in for it.

   byte 104  ->  'h'      (already printable, unchanged)
   byte  32  ->  'Ġ'      (a space, moved somewhere visible)
   byte  10  ->  'Ċ'      (a newline, likewise)

   "hello world"  ->  "helloĠworld"

That Ġ at the start of a token is not decoration. It is the space before the word. Once you know that, GPT token dumps stop looking strange.

What it costs

Nothing is free. For text written in Latin script, one character is one byte, so nothing changes.

For Devanagari, Tamil, Chinese or emoji, one character is three or four bytes. Unless the tokeniser learned pieces for that script, every character costs three or four tokens.

That is a real bill. The same sentence in Hindi can cost several times more tokens than in English. That means more money per request, and less content fitting in the same window.

Where you have seen this

  • GPT-2, GPT-3, GPT-4 and their tokenisers.
  • Llama 3, which moved to byte-level BPE from its earlier SentencePiece tokeniser.
  • Most models trained with the Hugging Face tokenizers library.
  • The Ġ characters you have seen in a debug print of tokens.

Remember this

  • Convert text to bytes first, then run ordinary BPE over the 256 byte symbols.
  • Unknown tokens become impossible, and every input round-trips exactly.
  • Non-Latin scripts pay for that guarantee in token count.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install tokenizers

Written against tokenizers 0.22.2. The byte mapping needs no libraries.

The byte-to-character table, and BPE on top of it

byte_level.py
from tokenizers import Tokenizer, models, trainers, pre_tokenizers, decoders

def bytes_to_unicode():
    """GPT-2's trick: give all 256 bytes a printable character, none of them a space."""
    keep = (list(range(ord("!"), ord("~") + 1))
            + list(range(ord("¡"), ord("¬") + 1))
            + list(range(ord("®"), ord("ÿ") + 1)))
    chars, n = keep[:], 0
    for b in range(256):
        if b not in keep:
            keep.append(b)
            chars.append(256 + n)      # push the awkward bytes above the printable range
            n += 1
    return dict(zip(keep, (chr(c) for c in chars)))

B2U = bytes_to_unicode()
print("total mapped bytes:", len(B2U), " all distinct:", len(set(B2U.values())) == 256)
print("byte 32 (space)   ->", repr(B2U[32]))
print("byte 10 (newline) ->", repr(B2U[10]))
print("byte 97 ('a')     ->", repr(B2U[97]))

for text in ["hello world", " hello", "café", "नमस्ते"]:
    raw = text.encode("utf-8")
    print(f"\n{text!r}")
    print("  utf-8 bytes :", list(raw)[:12], "..." if len(raw) > 12 else "")
    print("  as visible  :", "".join(B2U[b] for b in raw))

CORPUS = ["the cat sat on the mat", "the dog sat on the log",
          "a cat and a dog met on the road", "cats and dogs sat and sat",
          "the road was long and the cat was slow"] * 4

tok = Tokenizer(models.BPE())
tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=False)
tok.decoder = decoders.ByteLevel()
trainer = trainers.BpeTrainer(vocab_size=300, show_progress=False,
                              initial_alphabet=pre_tokenizers.ByteLevel.alphabet())
tok.train_from_iterator(CORPUS, trainer)
print("\nvocabulary size:", tok.get_vocab_size())

for text in ["the cat sat", " the cat", "café", "🙂", "नमस्ते"]:
    enc = tok.encode(text)
    print(f"\n{text!r}")
    print("  tokens:", enc.tokens)
    print("  ids   :", enc.ids)
    print("  decoded back:", repr(tok.decode(enc.ids)),
          " identical:", tok.decode(enc.ids) == text)

print("\nis there an [UNK] token in this vocabulary?", tok.token_to_id("[UNK]"))
Output
total mapped bytes: 256  all distinct: True
byte 32 (space)   -> 'Ġ'
byte 10 (newline) -> 'Ċ'
byte 97 ('a')     -> 'a'

'hello world'
  utf-8 bytes : [104, 101, 108, 108, 111, 32, 119, 111, 114, 108, 100] 
  as visible  : helloĠworld

' hello'
  utf-8 bytes : [32, 104, 101, 108, 108, 111] 
  as visible  : Ġhello

'café'
  utf-8 bytes : [99, 97, 102, 195, 169] 
  as visible  : café

'नमस्ते'
  utf-8 bytes : [224, 164, 168, 224, 164, 174, 224, 164, 184, 224, 165, 141] ...
  as visible  : नमसà¥įतà¥ĩ

vocabulary size: 292

'the cat sat'
  tokens: ['the', 'Ġcat', 'Ġsat']
  ids   : [258, 269, 267]
  decoded back: 'the cat sat'  identical: True

' the cat'
  tokens: ['Ġthe', 'Ġcat']
  ids   : [265, 269]
  decoded back: ' the cat'  identical: True

'café'
  tokens: ['c', 'a', 'f', 'Ã', '©']
  ids   : [66, 64, 69, 127, 102]
  decoded back: 'café'  identical: True

'🙂'
  tokens: ['ð', 'Ł', 'Ļ', 'Ĥ']
  ids   : [172, 253, 247, 224]
  decoded back: '🙂'  identical: True

'नमस्ते'
  tokens: ['à', '¤', '¨', 'à', '¤', '®', 'à', '¤', '¸', 'à', '¥', 'į', 'à', '¤', '¤', 'à', '¥', 'ĩ']
  ids   : [156, 97, 101, 156, 97, 106, 156, 97, 116, 156, 98, 235, 156, 97, 97, 156, 98, 229]
  decoded back: 'नमस्ते'  identical: True

is there an [UNK] token in this vocabulary? None

Reading that output

The mapping is a bijection over all 256 bytes. That is the whole guarantee. Every byte gets exactly one visible character, and no two bytes share one.

Bytes 33 to 126 map to themselves, so ordinary English text passes through unchanged and stays readable in a debug print. Only the awkward bytes are relocated.

'the' and 'Ġthe' are different tokens with different ids, 258 and 265. That is the space, carried inside the token. It is why prompts ending with a trailing space often behave oddly: the next word arrives without its usual leading space and takes a different id.

café needed five tokens for four characters. The é is two UTF-8 bytes, and neither is a learned piece, so both appear separately. Round trip: exact.

The emoji is four tokens. One emoji, four UTF-8 bytes, four tokens. Round trip: exact.

नमस्ते is six characters and eighteen tokens. That is the token-cost bill, printed. Three bytes per Devanagari character, no learned pieces for the script, so three tokens each. A real multilingual tokeniser learns pieces for common scripts and does much better than this, and the ratio never reaches parity with English.

token_to_id("[UNK]") returns None. There is no unknown token, because there is nothing that could produce one.

What the pre-tokeniser is doing

ByteLevel(add_prefix_space=False) does two jobs. It converts to bytes and applies the mapping. It also splits the text on a regular expression before any merge is allowed, so merges never span certain boundaries.

GPT-2's split pattern keeps contractions together, groups letters, groups digits, and attaches a leading space to a following word. That pattern is why Ġthe exists as a unit rather than the space merging with whatever came before it.

add_prefix_space=True prepends a space to the whole input, so the first word looks like every other word. Which setting a model was trained with matters, and getting it wrong shifts every first-token id.

Common mistakes

Forgetting initial_alphabet. Without initial_alphabet=ByteLevel.alphabet(), the trainer seeds the alphabet only from characters it saw in your corpus. Train on English and you lose the byte guarantee for everything else. This is one line and it is the whole point of the method.

Stripping Ġ before decoding. It is a byte, not formatting. Removing it deletes spaces from the output.

Assuming token count tracks character count. It tracks byte count for anything outside the learned vocabulary. Budget for non-Latin scripts by measuring, never by estimating from character counts.

Mismatching add_prefix_space with the trained model. Different setting, different ids for the first token of every sequence. The model still produces fluent text and quietly performs worse.

Try it yourself

Add a few hundred lines of Hindi or Tamil to CORPUS, raise vocab_size to 800, and re-encode नमस्ते. Watch the token count fall as the trainer learns byte sequences for that script. That experiment is the whole argument for training a tokeniser on the language you actually serve.

What to learn next

Researcher — Mathematics and papers.

Motivation

Radford et al. (2019), GPT-2, state the tradeoff directly. A Unicode-character-level vocabulary would need over 130,000 base symbols before any merges, which is larger than typical model vocabularies of 32k to 50k. A byte-level base needs 256.

Byte-level BPE therefore runs the standard BPE algorithm with $\Sigma = {0, \dots, 255}$ as the base alphabet. Since every string has a UTF-8 encoding, every string is representable, and no unknown token is required.

The bytes-to-unicode bijection

Working directly with byte values inside a text pipeline is awkward: many bytes are control characters, and some are whitespace that would interact with the pre-tokeniser. GPT-2 defines an injective map $\phi: {0,\dots,255} \to \text{Unicode}$ that:

  • fixes the printable ASCII range and two Latin-1 ranges, so English text stays legible in debug output;
  • maps the remaining 68 bytes to codepoints starting at U+0100, none of which are whitespace or control characters.

The map is a bijection onto its image, so decoding is exact. The visible artefacts Ġ (U+0120, byte 32) and Ċ (U+010A, byte 10) come from this.

Pre-tokenisation as a merge constraint

GPT-2 constrains merges with a regular expression, applied before BPE:

's|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S)|\s+

The pattern separates contractions, letter runs, digit runs and symbol runs, and attaches at most one leading space to each. Merges cannot cross these boundaries. Consequences:

  • Tokens like Ġthe exist as units, encoding the word-start distinction without a separate marker.
  • Digit runs are grouped but not split individually, so numbers tokenise irregularly. GPT-4's later pattern and Llama 3's variant change digit handling specifically for this reason.
  • Merges cannot span whitespace, which Liu et al. (2025), SuperBPE, show is a real cost: lifting the restriction in a second training phase yields up to 33 percent fewer tokens at 200k vocabulary and measurable downstream gains.

Fertility across scripts

UTF-8 encodes ASCII in one byte, most Latin-script accented characters and Cyrillic and Greek in two, most Indic and CJK characters in three, and emoji and rarer planes in four.

Absent learned pieces for a script, token count is byte count. Ahia et al. (2023), Do All Languages Cost the Same?, quantify the resulting disparity: for the same content, some languages cost several times more tokens than English under the same tokeniser, with direct consequences for API pricing, effective context length and latency.

Training on a genuinely multilingual corpus reduces the gap substantially, since the trainer learns multi-byte pieces for common scripts. It does not close it. English retains a structural advantage because the base alphabet is ASCII.

Comparison with SentencePiece byte fallback

Both eliminate unknown tokens; the mechanisms differ.

Byte-level BPESentencePiece + byte fallback
Base alphabet256 bytescorpus characters
Merges learned overbytescharacters
Byte piecesthe entire base256 reserved fallback slots
Uncovered scriptbytes, possibly merged into learned piecesbytes, never merged
Whitespace markerĠ (byte 32)▁ (U+2581)

The practical difference is that byte-level BPE can learn merges over the byte sequences of an uncovered script if that script appears in training. SentencePiece byte fallback cannot, because fallback happens after segmentation fails.

Adoption

GPT-2 onward at OpenAI, via tiktoken. RoBERTa, BART and most huggingface/tokenizers defaults. Llama 3 switched from SentencePiece to a 128k byte-level BPE vocabulary, citing improved compression on non-English text and code.

Open direction: removing tokenisation

Byte-level BPE is a compromise: bytes for coverage, merges for efficiency. Several lines of work attempt to remove the merge step entirely.

  • Xue et al. (2022), ByT5, operate directly on bytes with no tokeniser. Robust to noise and to unseen scripts, and considerably slower per unit of text.
  • Pagnoni et al. (2024), Byte Latent Transformer, group bytes into dynamically sized patches using an entropy model, so allocation adapts to content rather than to a fixed merge table. They report matching tokeniser-based models at scale with better inference efficiency.

Neither has displaced byte-level BPE in production, but both are direct attacks on the failures covered later in this section: glitch tokens, character blindness and digit handling.

Papers

  • Radford et al., Language Models are Unsupervised Multitask Learners (GPT-2), 2019
  • Wang et al., Neural Machine Translation with Byte-Level Subwords, AAAI 2020 — arxiv.org/abs/1909.03341
  • Xue et al., ByT5: Towards a token-free future with pre-trained byte-to-byte models, TACL 2022 — arxiv.org/abs/2105.13626
  • Ahia et al., Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models, EMNLP 2023 — arxiv.org/abs/2305.13707
  • Pagnoni et al., Byte Latent Transformer: Patches Scale Better Than Tokens, 2024 — arxiv.org/abs/2412.09871
  • Liu et al., SuperBPE: Space Travel for Language Models, COLM 2025 — arxiv.org/abs/2503.13423

What to learn next