Inside a Transformer Section 078
Tokeniser Internals
Models never see letters. This section shows exactly how your text becomes numbers, and every strange bug that causes.
13 of 13 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- Byte pair encoding, implemented
- WordPiece
- The unigram tokeniser
- SentencePiece
- Byte-level BPE
- Training a tokeniser on your own text
- Choosing a vocabulary size
- Special tokens
- Chat templates
- Glitch tokens
- Why models miscount letters and digits
- Counting tokens before you send them
- Adding tokens and resizing embeddings