Why models miscount letters and digits
A model asked how many r's are in "strawberry" has never seen the letters. The word arrives as one token, and arithmetic breaks for the same reason.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A model cannot count the letters in a word because it never receives the letters. The word arrives as one indivisible symbol.
Think of a road sign for a town, written in a script you cannot read. You recognise the shape and you know which town it means. Ask how many curved strokes it contains and you would have to go and look, letter by letter.
You know what the sign means without knowing what it is made of. A model is in that position with common words.
"strawberry" arrives as a single token. The three r's are not in the input.
Why the model cannot look closer
There is no zoom. The token becomes a row of numbers looked up in a table. That row was learned from how the word is used, not from how it is spelled.
Everything the model knows about spelling was learned indirectly. From text that spells things out, from rhymes, from typos, from other languages. That is thin and unreliable evidence.
So a model can be excellent at explaining what strawberries are and wrong about how many r's the word has. Those are different kinds of knowledge and it has plenty of one and little of the other.
Numbers have the same problem, worse
Numbers get chopped in whatever way the tokeniser learned. The pieces do not line up with the columns you would use to add.
what you want to add: what the model receives:
1 2 3 4 "12" "34"
+ 5 6 7 8 + "56" "78"
--------- ---------------
no shared column structure at allChange one digit and the whole split can change. 999 might be one token while 1000 becomes three. The model has to learn arithmetic on a representation that keeps shifting.
The fixes that work
Split every digit. Force the tokeniser to cut numbers into single digits. Now 1234 is four tokens, one per column, and columns line up between two numbers. Several current models do exactly this.
Spell it out. Writing s t r a w b e r r y puts the letters into the input. Then counting is a task the model can actually do.
Use a tool. For real arithmetic, have the model call a calculator. The reliable answer comes from code, not from the token sequence.
What this is not
This is not evidence that the model is stupid, and it is not a reasoning failure. It is a missing-input problem.
Asking a model to count letters is like asking someone to count the pixels in a photo they were shown as a thumbnail. The information was removed before they saw anything.
Where you have seen this
- The famous "how many r's in strawberry" question.
- A model confidently getting a long addition wrong.
- Rhyming and wordplay that nearly works and then does not.
- A model miscounting the words in a sentence.
Remember this
- Common words arrive as one token, so the letters are not in the input.
- Numbers are split by frequency, not by place value, so columns never line up.
- Splitting digits, spelling words out, or calling a tool are the fixes that work.
What to learn next
- Counting tokens before you send them — measuring what tokenisation costs you.
- Function calling — handing arithmetic to code that gets it right.
- Hallucination — the wider class of confident wrong answers.
Developer — Code and libraries.
Setup
pip install tokenizersWritten against tokenizers 0.22.2. The corpus is inline, so nothing is downloaded.
Seeing the missing letters and the misaligned digits
import random
from tokenizers import Tokenizer, models, trainers, pre_tokenizers, decoders
WORDS = ["strawberry", "raspberry", "blueberry", "berry", "straw", "berries",
"banana", "mango", "apple", "the", "a", "and", "is", "was", "fruit"]
rng = random.Random(0)
text = [" ".join(rng.choice(WORDS) for _ in range(10)) for _ in range(600)]
numbers = [" ".join(str(rng.randint(0, 99999)) for _ in range(10)) for _ in range(600)]
CORPUS = text + numbers
def build(split_digits):
pre = [pre_tokenizers.ByteLevel(add_prefix_space=False)]
if split_digits:
pre.insert(0, pre_tokenizers.Digits(individual_digits=True))
tok = Tokenizer(models.BPE())
tok.pre_tokenizer = pre_tokenizers.Sequence(pre)
tok.decoder = decoders.ByteLevel()
tok.train_from_iterator(CORPUS, trainers.BpeTrainer(
vocab_size=600, show_progress=False,
initial_alphabet=pre_tokenizers.ByteLevel.alphabet()))
return tok
plain = build(split_digits=False)
digits = build(split_digits=True)
print("how words look to the model:")
for w in [" strawberry", " raspberry", " berries", " strawberries"]:
print(f" {w!r:>16} -> {plain.encode(w).tokens}")
print("\nthe question 'how many r in strawberry' asks about letters the model never sees:")
t = plain.encode(" strawberry").tokens
print(f" the model receives {len(t)} token(s): {t}")
print(" the letter r appears 3 times in the spelling, and 0 times as a separate token")
print("\nnumbers, without digit splitting:")
for n in ["1234", "12345", "999", "1000", "1001"]:
print(f" {n:>6} -> {plain.encode(' ' + n).tokens}")
print("\nthe same numbers, with a digit pre-tokeniser:")
for n in ["1234", "12345", "999", "1000", "1001"]:
print(f" {n:>6} -> {digits.encode(' ' + n).tokens}")
print("\nwhy that matters for addition. column alignment, without splitting:")
for a, b in [("1234", "5678"), ("999", "1")]:
print(f" {a:>5} -> {plain.encode(' ' + a).tokens}")
print(f" {b:>5} -> {plain.encode(' ' + b).tokens}")
print()
print("with splitting, every column is its own token:")
for a, b in [("1234", "5678")]:
print(f" {a:>5} -> {digits.encode(' ' + a).tokens}")
print(f" {b:>5} -> {digits.encode(' ' + b).tokens}")
print("\nspelling a word out gives the model the letters it was missing:")
spelled = " ".join("strawberry")
print(f" {spelled!r}")
print(f" -> {plain.encode(' ' + spelled).tokens}")how words look to the model:
' strawberry' -> ['Ġstrawberry']
' raspberry' -> ['Ġraspberry']
' berries' -> ['Ġberries']
' strawberries' -> ['Ġstraw', 'berries']
the question 'how many r in strawberry' asks about letters the model never sees:
the model receives 1 token(s): ['Ġstrawberry']
the letter r appears 3 times in the spelling, and 0 times as a separate token
numbers, without digit splitting:
1234 -> ['Ġ12', '34']
12345 -> ['Ġ1', '23', '45']
999 -> ['Ġ999']
1000 -> ['Ġ1', '00', '0']
1001 -> ['Ġ1', '00', '1']
the same numbers, with a digit pre-tokeniser:
1234 -> ['Ġ', '1', '2', '3', '4']
12345 -> ['Ġ', '1', '2', '3', '4', '5']
999 -> ['Ġ', '9', '9', '9']
1000 -> ['Ġ', '1', '0', '0', '0']
1001 -> ['Ġ', '1', '0', '0', '1']
why that matters for addition. column alignment, without splitting:
1234 -> ['Ġ12', '34']
5678 -> ['Ġ56', '78']
999 -> ['Ġ999']
1 -> ['Ġ1']
with splitting, every column is its own token:
1234 -> ['Ġ', '1', '2', '3', '4']
5678 -> ['Ġ', '5', '6', '7', '8']
spelling a word out gives the model the letters it was missing:
's t r a w b e r r y'
-> ['Ġ', 's', 'Ġ', 't', 'Ġ', 'r', 'Ġa', 'Ġ', 'w', 'Ġb', 'Ġ', 'e', 'Ġ', 'r', 'Ġ', 'r', 'Ġ', 'y']Reading that output
One token. That is the whole explanation. The model receives a single id for strawberry. Every question about its spelling is a question about information that is not in the input.
strawberries splits differently from strawberry. Ġstraw plus berries, not Ġstrawberry plus s. Adding one letter changed the entire decomposition. Nothing about the token boundaries tracks morphology reliably.
999 is one token and 1000 is three. Consecutive integers, wildly different representations. A model has to learn that Ġ999 and Ġ1 00 0 are neighbours on the number line, from a representation that gives no hint of it.
1234 and 12345 split differently too, into 12 34 versus 1 23 45. Adding a digit re-aligned every boundary. This is why models are much worse at long arithmetic than short arithmetic: the representation destabilises as numbers grow.
Look at the alignment block. Ġ12 34 above Ġ56 78 happens to line up here, and Ġ999 above Ġ1 does not line up at all. The model has to handle both cases with the same learned circuits.
With digit splitting, every number is one token per column. 1234 and 5678 both give four digit tokens. Now "add the units, carry, add the tens" is expressible over token positions. This single pre-tokeniser change is why several current models handle arithmetic noticeably better than their predecessors.
Spelling the word out costs 18 tokens for 10 letters, because each letter is preceded by a space token. Expensive, and now the letters are genuinely present. That is the trade behind the "spell it out first" prompting trick.
Measuring the gap on your own tokeniser
# continuing from the script above
for w in ["strawberry", "raspberry", "banana"]:
whole = len(plain.encode(" " + w).ids)
spelled = len(plain.encode(" " + " ".join(w)).ids)
print(f"{w:>12}: {whole} token(s) whole, {spelled} spelled out, "
f"{spelled/whole:.1f}x the cost") strawberry: 1 token(s) whole, 18 spelled out, 18.0x the cost
raspberry: 1 token(s) whole, 16 spelled out, 16.0x the cost
banana: 1 token(s) whole, 8 spelled out, 8.0x the costCharacter-level access is available and it is expensive. That is the honest framing: this is a cost decision, not an impossibility.
What to do in practice
For counting characters, do not ask the model. Count in code. text.count("r") is exact, free and never wrong.
For arithmetic, call a tool. Give the model a calculator function and let it use it. Even models with digit splitting are unreliable on long multiplication.
For spelling tasks, spell the input out. Insert spaces or hyphens between characters in the prompt. It costs tokens and it supplies the missing information.
When training your own model, split digits. pre_tokenizers.Digits(individual_digits=True) is one line and it is the highest-value tokeniser change for numeric work.
Common mistakes
Concluding the model cannot reason. It counts perfectly well when the letters are present. Test with a spelled-out input before drawing conclusions about reasoning.
Assuming digit splitting fixes arithmetic. It removes the representation obstacle. Long multiplication still requires many sequential steps and remains error-prone. Tools remain the right answer for anything that must be correct.
Using chain-of-thought as a substitute for the missing input. Reasoning step by step helps, because the model often spells the word out as part of its working. When it does not, the information is still absent and the extra steps do not conjure it.
Assuming every model splits digits the same way. Some split every digit, some group in threes from the right, some do neither. Print the tokenisation before assuming anything about a specific model.
Try it yourself
Encode "1234567890" with both tokenisers and count the tokens. Then encode the same digits in reverse order. With the plain tokeniser the two differ, sometimes substantially; with digit splitting they are identical. That symmetry is exactly what arithmetic needs and what frequency-driven merges destroy.
What to learn next
- Counting tokens before you send them — measuring what tokenisation costs you.
- Function calling — handing arithmetic to code that gets it right.
- Hallucination — the wider class of confident wrong answers.
Researcher — Mathematics and papers.
The information-theoretic framing
For a tokenisation function $T$ mapping strings to id sequences, the model observes $T(s)$, not $s$. Any function of $s$ not determined by $T(s)$ together with the model's learned parameters is unavailable.
Character counting is such a function. The map from token id to its character content exists in the tokeniser, and is not exposed to the model. The model can only recover it from statistical regularities in training text: explicit spellings, acrostics, rhyme, typos, and cross-lingual transliteration. That evidence is sparse and biased.
Kaushal and Mahowald (2022) probe whether token embeddings encode orthographic information and find they do, partially: linear probes recover character identity above chance but far from reliably, and accuracy degrades with word length. The information is present in weak form, not absent, which explains why models sometimes get spelling questions right.
Why numbers are the harder case
Number tokenisation under frequency-driven merges has three properties that obstruct arithmetic.
Non-uniform segmentation. Adjacent integers receive different decompositions. The developer output shows 999 as one token and 1000 as three under the same tokeniser.
Boundaries independent of place value. A merge is learned because a digit pair was frequent in the corpus, not because it corresponds to a positional column. Two operands of an addition therefore need not share a boundary structure.
Instability under length. Adding a digit re-runs the merge sequence and can change every boundary, so a circuit learned for 4-digit numbers does not transfer to 5-digit numbers.
Nogueira et al. (2021) show transformers extrapolate poorly on addition and that representation choice dominates the effect. Zhou et al. (2024), Transformers Can Achieve Length Generalization But Not Robustly, show length generalisation on addition is achievable with the right position and format choices and remains fragile.
Digit tokenisation strategies
| Strategy | Example: 12345 | Used by |
|---|---|---|
| frequency-driven merges | varies, unstable | GPT-2, early BPE tokenisers |
| single digits | 1 / 2 / 3 / 4 / 5 | Llama, PaLM, several current models |
| groups of three, right to left | 12 / 345 | some GPT-4-era tokenisers |
| right-to-left with explicit alignment | task-specific | arithmetic-focused research |
Singh and Strouse (2024), Tokenization counts, compare left-to-right and right-to-left digit grouping and find the direction of grouping measurably changes arithmetic accuracy, with right-to-left aligning better to place value. Zhou et al. (2024) report that index hints and reversed-digit formats substantially improve length generalisation on addition.
Chain of thought as a partial workaround
Chain-of-thought prompting improves arithmetic and spelling tasks partly because the intermediate text often materialises the missing information. A model that writes out s-t-r-a-w-b-e-r-r-y before counting has put the characters into its own context, where subsequent attention can read them.
This is worth stating precisely: the benefit is representational, not only computational. The model is repairing its input, not only thinking longer.
Tokeniser-free approaches
The failure class disappears if the model reads characters or bytes directly.
- Xue et al. (2022), ByT5, operate on raw bytes. They report improved robustness to noise, better performance on tasks sensitive to spelling, and substantially higher compute cost per unit of text.
- Pagnoni et al. (2024), Byte Latent Transformer, allocate compute dynamically over byte patches sized by an entropy model, reporting parity with token-based models at scale with better inference efficiency, and improved robustness on character-level tasks.
Neither has displaced subword tokenisation in production, and the failures described here remain live for essentially every deployed model.
The measurement caution
"Model cannot count letters" is frequently cited as evidence about reasoning capability. It is evidence about input representation. A meaningful evaluation supplies the characters — spelled out, or via a character-level tool — and measures whether the model can then perform the task. When it can, the deficit is located in tokenisation, not in reasoning.
Papers
- Nogueira, Jiang and Lin, Investigating the Limitations of Transformers with Simple Arithmetic Tasks, 2021 — arxiv.org/abs/2102.13019
- Kaushal and Mahowald, What do tokens know about their characters and how do they know it?, NAACL 2022 — arxiv.org/abs/2206.02608
- Xue et al., ByT5, TACL 2022 — arxiv.org/abs/2105.13626
- Singh and Strouse, Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs, 2024 — arxiv.org/abs/2402.14903
- Zhou et al., Transformers Can Achieve Length Generalization But Not Robustly, 2024 — arxiv.org/abs/2402.09371
- Pagnoni et al., Byte Latent Transformer, 2024 — arxiv.org/abs/2412.09871
What to learn next
- Counting tokens before you send them — measuring what tokenisation costs you.
- Function calling — handing arithmetic to code that gets it right.
- Hallucination — the wider class of confident wrong answers.