BPE (byte pair encoding)
In one sentence BPE builds a tokenizer's vocabulary by repeatedly merging the most frequent pair of adjacent pieces, so common words become single tokens.
Updated
Byte pair encoding, or BPE, learns a tokenizer's vocabulary by starting from single characters and repeatedly gluing together the most frequent adjacent pair.
It works like the shorthand a note-taker invents during a long lecture. Writing "the" hundreds of times, she starts using one symbol for it. Then a symbol for "tion", then for "because". Frequency earns compression: the commonest sequences collapse into single marks, while rare words stay spelled out. BPE runs this process by pure counting.
start: c h a t b o t c h a i
merge 1: "c h" is most frequent → ch a t b o t ch a i
merge 2: "ch a" next → cha t b o t cha i
...continue for 50,000 merges → the vocabularyRun tens of thousands of merges over a huge corpus, and the result is a vocabulary where "the" and "namaste" are single tokens while an unusual surname is three fragments. Every word in every language remains representable — worst case, character by character — which killed the old "unknown word" problem of fixed dictionaries.
Two consequences carry into daily LLM work. The merges mirror the training corpus, so under-represented languages fragment into more tokens and cost more per sentence. And modern GPT-family tokenizers run BPE over raw bytes, so emojis and any script whatsoever stay encodable. Variants like WordPiece (BERT's scheme) and SentencePiece differ in details, not in the core merge-by-frequency idea.
Where to go next
- Full lesson: Tokenization
- Related terms: tokenizer, token, llm, context-window