Tokeniser Internals

SentencePiece

SentencePiece treats the space as an ordinary character, so it works on languages that do not use spaces and can rebuild the exact original text.

On this page 7
  1. The visible marker
  2. What this buys you
  3. Byte fallback
  4. The honest catch
  5. Where you have seen this
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

SentencePiece treats the space as a normal character, so it never has to guess where words begin.

Think about writing a shopping list on a strip of paper with no gaps: riceoildalatta. You can still read it, because you know the words. A tokeniser that starts by splitting on spaces has nothing to split on.

Japanese, Chinese and Thai are written this way. So is a lot of real-world text: hashtags, code, URLs, product codes.

SentencePiece skips the splitting step entirely. It reads the raw stream of characters, space included, and learns pieces from that.

The visible marker

Since a space is now part of the text, it needs a printable stand-in. SentencePiece uses a small lower block character, ▁, written before a word.

   "the cat sat"     ->    ▁the   ▁cat   ▁sat

   "▁cat" and "cat" are different pieces:
      ▁cat   is a word starting after a space
      cat    is "cat" inside a longer word

That distinction is why the marker exists. The model gets to know whether a piece began a word or continued one.

What this buys you

It works on any language. No rules about where words end, because it never asks.

It can rebuild your text exactly. Turn ▁ back into a space, join everything up, and you have the original string. Nothing was thrown away.

That second property matters more than it sounds. Think of a tokeniser that splits on spaces and glues them back with one space. It cannot tell you whether your input had one space or three.

Byte fallback

There is one more piece. A tokeniser trained on English has never seen Devanagari or Cyrillic or an emoji.

SentencePiece can be told to keep a spare entry for each of the 256 possible byte values. Anything unrecognised is broken into raw bytes, and every byte has an entry.

The result: nothing is ever unknown, and every input can be reproduced exactly. The cost is that unfamiliar text becomes very long in tokens.

The honest catch

By default SentencePiece cleans your text before tokenising. It collapses runs of spaces, normalises some characters, and adds a space at the start.

That cleaning means the round trip is not exact unless you turn it off. Most people discover this when a code snippet loses its indentation.

The fix is three settings, shown in the code below. Worth knowing before you train on anything where whitespace carries meaning.

Where you have seen this

  • T5, mT5 and Flan-T5.
  • Llama and Llama 2's tokenisers.
  • Gemma and many multilingual models.
  • Whisper, for speech transcription across languages.

Remember this

  • The space is an ordinary character, marked as ▁ at the start of a word.
  • It needs no word-splitting rules, so it works on any script.
  • Byte fallback means nothing is ever unknown.
  • Exact round-tripping requires switching the default text cleaning off.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install sentencepiece

Written against sentencepiece 0.2.1. Everything below trains from an inline string, with no downloads.

Training in memory and inspecting the result

sp_train.py
import io
import sentencepiece as spm

TEXT = """the cat sat on the mat
the dog sat on the log
a cat and a dog met on the road
the mat was red and the log was brown
cats and dogs sat and sat and sat
the road was long and the cat was slow
a slow cat and a long road and a red mat
the dog and the cat and the log and the mat
"""

model = io.BytesIO()
spm.SentencePieceTrainer.train(
    sentence_iterator=iter(TEXT.strip().split("\n")),
    model_writer=model,
    vocab_size=300,
    model_type="unigram",
    character_coverage=1.0,
    byte_fallback=True,        # unknown characters become raw bytes, so nothing is lost
    hard_vocab_limit=False,    # our corpus is tiny, so let it stop short
    minloglevel=2,             # quiet the trainer's own logging
)
sp = spm.SentencePieceProcessor(model_proto=model.getvalue())

print("vocabulary size:", sp.get_piece_size())
print("ids 0-3        :", [sp.id_to_piece(i) for i in range(4)])
print("ids 3-258     : the 256 raw byte pieces, <0x00> to <0xFF>")
print("learned pieces :", [sp.id_to_piece(i) for i in range(259, 271)])

s = "the cat sat on the mat"
print(f"\n{s!r}")
print("  pieces:", sp.encode(s, out_type=str))
print("  ids   :", sp.encode(s))

print("\nthe space is glued to the front of a word as U+2581:")
print("  'the cat'  ->", sp.encode("the cat", out_type=str))
print("  ' the cat' ->", sp.encode(" the cat", out_type=str))

print("\ndecoding:")
for original in ["the cat sat", "  double  spaces  ", "the\tcat"]:
    back = sp.decode(sp.encode(original))
    print(f"  {original!r:>22} -> {back!r}   identical: {original == back}")

print("\nbyte fallback on characters the trainer never saw:")
for word in ["café", "🙂"]:
    pieces = sp.encode(word, out_type=str)
    print(f"  {word!r:>7} -> {pieces}  ({len(pieces)} pieces)")
    print(f"          round trip: {sp.decode(sp.encode(word))!r}")
Output
vocabulary size: 288
ids 0-3        : ['<unk>', '<s>', '</s>', '<0x00>']
ids 3-258     : the 256 raw byte pieces, <0x00> to <0xFF>
learned pieces : ['▁', 'e', '▁a', 't', 'n', 'd', 'h', 's', '▁cat', 'w', '▁sat', 'a']

'the cat sat on the mat'
  pieces: ['▁', 't', 'h', 'e', '▁cat', '▁sat', '▁', 'on', '▁', 't', 'h', 'e', '▁mat']
  ids   : [259, 262, 265, 260, 267, 269, 259, 273, 259, 262, 265, 260, 272]

the space is glued to the front of a word as U+2581:
  'the cat'  -> ['▁', 't', 'h', 'e', '▁cat']
  ' the cat' -> ['▁', 't', 'h', 'e', '▁cat']

decoding:
           'the cat sat' -> 'the cat sat'   identical: True
    '  double  spaces  ' -> 'double spaces'   identical: False
              'the\tcat' -> 'the cat'   identical: False

byte fallback on characters the trainer never saw:
   'café' -> ['▁', 'c', 'a', '<0x66>', '<0xC3>', '<0xA9>']  (6 pieces)
          round trip: 'café'
      '🙂' -> ['▁', '<0xF0>', '<0x9F>', '<0x99>', '<0x82>']  (5 pieces)
          round trip: '🙂'

Reading that output

The first three ids are reserved, then 256 byte pieces occupy ids 3 to 258. byte_fallback=True costs you 256 vocabulary slots before a single piece is learned. On a real 32,000-entry vocabulary that is under one percent. On this toy, it is most of the table.

The learned pieces start at id 259 with a bare ▁. That is the space on its own, used when no ▁word piece exists. Then ▁cat and ▁sat appear as complete word-with-space units.

'the cat' and ' the cat' encode identically. The default normaliser adds a leading space to every input, so a string that already has one gains nothing. Convenient for consistency and a genuine surprise when you need the two to differ.

Two of the three round trips fail. ' double spaces ' comes back as 'double spaces', and the tab becomes a space. The default nmt_nfkc normaliser collapses whitespace runs and strips the ends. SentencePiece is described as lossless, and it is lossless given a normaliser that does nothing. The default normaliser does something.

café needed byte fallback for f, not only for é. The training corpus has no letter f anywhere, so <0x66>, the byte for f, is used. That shows byte fallback triggering on plain ASCII when coverage is thin, not only on exotic scripts.

The emoji took four byte pieces. A single emoji is four UTF-8 bytes. Every one round-trips exactly.

Making it genuinely lossless

Three settings turn the cleaning off:

sp_lossless.py
import io
import sentencepiece as spm

TEXT = """the cat sat on the mat
the dog sat on the log
a cat and a dog met on the road
the mat was red and the log was brown
cats and dogs sat and sat and sat
the road was long and the cat was slow
a slow cat and a long road and a red mat
the dog and the cat and the log and the mat
"""

m = io.BytesIO()
spm.SentencePieceTrainer.train(
    sentence_iterator=iter(TEXT.strip().split("\n")), model_writer=m,
    vocab_size=300, model_type="unigram", byte_fallback=True,
    hard_vocab_limit=False, minloglevel=2,
    normalization_rule_name="identity",   # do not rewrite the text at all
    remove_extra_whitespaces=False,       # keep every space
    add_dummy_prefix=False,               # do not invent a leading space
)
sp = spm.SentencePieceProcessor(model_proto=m.getvalue())

for s in ["the cat sat", "  double  spaces  ", "the\tcat"]:
    back = sp.decode(sp.encode(s))
    print(f"{s!r:>22} -> {back!r}  identical: {s == back}")
Output
         'the cat sat' -> 'the cat sat'  identical: True
  '  double  spaces  ' -> '  double  spaces  '  identical: True
            'the\tcat' -> 'the\tcat'  identical: True

All three now round-trip exactly. Use this configuration for code, for structured text, and for anything where whitespace carries meaning.

Common mistakes

Assuming the defaults are lossless. They are not, as the first output shows. If your model must reproduce indentation, set the three flags above before training, not after.

Setting vocab_size below 256 with byte_fallback=True. The trainer raises Vocabulary size is smaller than required_chars. Byte fallback reserves 256 slots plus the special tokens, so the floor is around 280.

Hitting hard_vocab_limit on a small corpus. Ask for more pieces than the text supports and training fails outright. Set hard_vocab_limit=False to let it stop short, and check get_piece_size() afterwards.

Unicode errors when printing on Windows. The ▁ character is not in the default Windows console codepage. Run with PYTHONIOENCODING=utf-8, or the script raises UnicodeEncodeError on the first print of a piece.

Mixing up the library and the algorithm. SentencePiece implements both unigram and BPE. model_type="bpe" gives you byte pair encoding inside SentencePiece. Llama used exactly that, which is why its tokeniser has ▁ markers despite being a BPE tokeniser.

Try it yourself

Train once with byte_fallback=True and once with byte_fallback=False, then encode "café" with both. Without fallback you get <unk> and the accent is gone forever. Compare the round trip in each case, and you have the argument for byte fallback in two lines of output.

What to learn next

Researcher — Mathematics and papers.

Design

Kudo and Richardson (2018) identify a specific problem: most tokenisers assume pre-tokenised input, usually whitespace-delimited, which makes them language-dependent and makes detokenisation lossy. SentencePiece removes the assumption.

Three decisions follow:

Raw input. The sentence is treated as a Unicode character sequence with no pre-tokenisation. Whitespace is escaped to U+2581 (LOWER ONE EIGHTH BLOCK), so it participates in piece formation like any other character.

Reversibility. Because whitespace is preserved as a character, decode(encode(x)) == x holds by construction, conditional on the normaliser being the identity. The default normaliser is nmt_nfkc, which applies NFKC, collapses whitespace runs, and prepends a dummy space. That default is what breaks the identity, as the developer output demonstrates.

Algorithm-agnostic. SentencePiece implements unigram (Kudo, 2018) and BPE (Sennrich et al., 2016) behind one interface, with the same normalisation, escaping and serialisation.

Character coverage

For languages with large character sets, character_coverage controls what fraction of characters in the training corpus must be representable. The reference recommends 0.9995 for Chinese and Japanese and 1.0 for languages with small alphabets. Characters outside the covered set become <unk>, or raw bytes when byte_fallback is enabled.

Byte fallback

Introduced later than the paper and used by Llama and Gemma. It adds 256 pieces <0x00> to <0xFF>. Any string the vocabulary cannot segment is decomposed into UTF-8 bytes, each mapped to its byte piece.

Consequences:

  • The unknown token becomes unreachable in practice.
  • Any Unicode string round-trips exactly.
  • Fertility on uncovered scripts is high: one Devanagari character is three UTF-8 bytes, so three tokens if no learned piece applies.

This is functionally similar to byte-level BPE but structurally different. Byte-level BPE maps all input to bytes before training, so merges are learned over bytes. SentencePiece with byte fallback learns over characters and drops to bytes only where the vocabulary fails.

Normalisation

SentencePiece ships several rules and accepts a custom TSV mapping:

RuleEffect
nmt_nfkcNFKC plus NMT-specific whitespace handling. Default.
nfkcUnicode NFKC only
nmt_nfkc_cfas nmt_nfkc plus case folding
identityno change

NFKC is aggressive: it folds full-width to half-width, maps ligatures to their components, and normalises many compatibility characters. For most natural-language work that is desirable. For code, for text where visual form matters, and for security-sensitive input, it is not.

Trained pieces are scored

For model_type="unigram", the serialised model stores a log-probability per piece, so the same model file supports Viterbi decoding and sampling. The nbest_size and alpha arguments to encode expose subword regularisation directly.

Adoption

  • T5, mT5, ByT5's non-byte baselines, ALBERT, XLNet: unigram.
  • Llama, Llama 2, Gemma: BPE inside SentencePiece, with byte fallback in the later models.
  • Whisper and several multilingual speech models.

Newer decoder-only models increasingly use huggingface/tokenizers byte-level BPE instead, for speed and for ecosystem alignment rather than for a quality difference. SentencePiece remains the standard where multilingual coverage across many scripts is the primary concern.

Papers

What to learn next