AI glossary

BPE (byte pair encoding)

In one sentence BPE builds a tokenizer's vocabulary by repeatedly merging the most frequent pair of adjacent pieces, so common words become single tokens.

By Updated

Byte pair encoding, or BPE, learns a tokenizer's vocabulary by starting from single characters and repeatedly gluing together the most frequent adjacent pair.

It works like the shorthand a note-taker invents during a long lecture. Writing "the" hundreds of times, she starts using one symbol for it. Then a symbol for "tion", then for "because". Frequency earns compression: the commonest sequences collapse into single marks, while rare words stay spelled out. BPE runs this process by pure counting.

start:   c h a t   b o t   c h a i
merge 1: "c h" is most frequent  →  ch a t   b o t   ch a i
merge 2: "ch a" next             →  cha t   b o t   cha i
...continue for 50,000 merges → the vocabulary

Run tens of thousands of merges over a huge corpus, and the result is a vocabulary where "the" and "namaste" are single tokens while an unusual surname is three fragments. Every word in every language remains representable — worst case, character by character — which killed the old "unknown word" problem of fixed dictionaries.

Two consequences carry into daily LLM work. The merges mirror the training corpus, so under-represented languages fragment into more tokens and cost more per sentence. And modern GPT-family tokenizers run BPE over raw bytes, so emojis and any script whatsoever stay encodable. Variants like WordPiece (BERT's scheme) and SentencePiece differ in details, not in the core merge-by-frequency idea.

Where to go next