AI glossary

Tokenizer

In one sentence A tokenizer is the component that chops text into the token pieces a language model reads, using a fixed vocabulary learned from data.

By Updated

A tokenizer is the tool that splits raw text into tokens — the units a language model actually processes — and maps each one to a number.

It is the vegetable-chopping station before the cooking. The model never sees your sentence; it sees the chopped pieces, numbered from a fixed menu. "Unbelievable" might become un + believ + able — three pieces, three IDs.

"I love biryani!"  →  ["I", " love", " bir", "iyani", "!"]  →  [40, 3047, 6301, 45001, 0]

The vocabulary — commonly 30,000 to 250,000 pieces — is learned from a text corpus before training, usually with BPE: frequent words earn a single token, rare words get assembled from fragments. Nothing is ever "out of vocabulary", because in the worst case a word decomposes to letters.

Tokenizers explain several LLM oddities that puzzle beginners. Counting letters in a word is hard because the model sees fragments, not letters — the famous "how many r's in strawberry" stumble. Hindi or Tamil text often costs several times more tokens than English, because the vocabulary was trained on mostly-English data — inflating cost and eating context-window. Numbers split unpredictably, which does arithmetic no favours.

One iron rule for developers: a model only understands its own tokenizer. Feeding IDs from a different tokenizer produces confident gibberish, which is why every Hugging Face checkpoint ships its tokenizer alongside the weights.

Where to go next