Tokenizer
In one sentence A tokenizer is the component that chops text into the token pieces a language model reads, using a fixed vocabulary learned from data.
Updated
A tokenizer is the tool that splits raw text into tokens — the units a language model actually processes — and maps each one to a number.
It is the vegetable-chopping station before the cooking. The model never sees your sentence; it sees the chopped pieces, numbered from a fixed menu. "Unbelievable" might become un + believ + able — three pieces, three IDs.
"I love biryani!" → ["I", " love", " bir", "iyani", "!"] → [40, 3047, 6301, 45001, 0]The vocabulary — commonly 30,000 to 250,000 pieces — is learned from a text corpus before training, usually with BPE: frequent words earn a single token, rare words get assembled from fragments. Nothing is ever "out of vocabulary", because in the worst case a word decomposes to letters.
Tokenizers explain several LLM oddities that puzzle beginners. Counting letters in a word is hard because the model sees fragments, not letters — the famous "how many r's in strawberry" stumble. Hindi or Tamil text often costs several times more tokens than English, because the vocabulary was trained on mostly-English data — inflating cost and eating context-window. Numbers split unpredictably, which does arithmetic no favours.
One iron rule for developers: a model only understands its own tokenizer. Feeding IDs from a different tokenizer produces confident gibberish, which is why every Hugging Face checkpoint ships its tokenizer alongside the weights.
Where to go next
- Full lesson: Tokenization
- Related terms: token, bpe, context-window, embedding