AI glossary

Token

In one sentence A token is a chunk of text, often a word or part of a word, and it is the unit a language model actually reads and writes.

By Updated

A token is a chunk of text — usually a whole word, sometimes a piece of one — and it is the unit a language model actually reads and writes.

Before you cook, you chop. Vegetables do not go into the pan whole; they are cut into pieces the pan can handle. A model does the same to your sentence before doing anything with it, and the pieces are tokens. Common words survive whole. Rarer words get cut up.

"Tokenization is helpful for beginners"
   →  ["Token", "ization", " is", " helpful", " for", " beginners"]
       6 tokens, and notice the spaces travel with the words

Why cut at all? A dictionary of every English word would still miss names, typos, code and Hindi. A vocabulary of word pieces covers anything, because an unknown word can always be spelled out of smaller known pieces. That is what algorithms like byte-pair encoding do: they learn which pieces are worth keeping whole from the training text.

The numbers you need

Roughly, for English: 1 token is about 4 characters, and 1,000 tokens is about 750 words. Hindi, Marathi, Tamil and other non-Latin scripts split into far more tokens for the same meaning — sometimes two or three times as many — which makes them more expensive to send and quicker to fill the context window.

Three practical things fall out of this. API pricing is per token, input and output counted separately. Length limits are counted in tokens, not characters, and hitting one gives you a context length exceeded error. And a model cannot see letters inside a token, which is exactly why asking it to count the r's in "strawberry" goes wrong while asking it to write a working function does not.

Each token is converted to an id, then to an embedding, before the model does anything with it. Different models use different tokenizers, so the same sentence has a different token count on different models.

Where to go next

Related terms

Learn this properly

Full lessons that use this term in context.

  • Natural Language Processing

    Tokenization

    Tokenization is cutting text into small pieces called tokens, because a model can only work with a fixed list of known pieces.

Back to the glossary