Sequence Labelling and Structure
Why token accuracy lies for NER
A tagger can get almost every individual tag right and still miss every single entity, because entity-level correctness needs the whole span to match, not only most of its tokens.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Token accuracy counts individual word-tags correct or wrong. But an entity is only useful if its whole span is right — one wrong word can wreck an otherwise perfect answer.
Think about a delivery address. A courier gets the house number, street, area, city and PIN code all correct, except the house number is off by one digit. The parcel does not arrive at 95% of the right place. It arrives at the wrong house, completely. Almost right is not right.
Named entities work the same way. Say a model tags "Reserve Bank of" correctly but misses "India". The entity it extracted is not "Reserve Bank of India" with a small error. It is a different, wrong entity altogether.
Why it exists
Most words in a sentence are not part of any entity, tagged as plain O. A model can get nearly every individual word-tag right by predicting O most of the time, and being roughly correct on entity boundaries. That makes token accuracy look excellent, even when the model is actually bad at the one thing that matters: getting complete entity spans right.
Entity-level evaluation fixes this by scoring whole spans, not individual tags. A predicted entity only counts as correct if its exact boundaries and its exact type both match the true entity. Off-by-one-word does not get partial credit.
How it works
Gold entity: [ Reserve Bank of India ] ORG
Model's guess: [ Reserve Bank of ] ORG
^
missed "India" -- boundary is wrong
Token accuracy: 3 out of 4 entity words tagged correctly = 75% looks fine
Entity-level score: the predicted span does not exactly match the gold span
= counted as a complete miss, both for precision and recallPrecision asks: of the entities the model claimed to find, how many were actually correct? Recall asks: of the entities that actually exist, how many did the model find? F1 combines the two into a single number. It only rewards exact, complete matches.
Where you have already seen it
- Resume parsers. A model that extracts "John Sm" instead of "John Smith" has not extracted 90% of a name. It has extracted a name for the wrong person entirely. Treating that as "close enough" causes real downstream errors.
- Redaction software. Missing the last character of an account number defeats the entire point of redacting it, no matter how many other characters were caught.
- Any leaderboard reporting NER results. When a paper reports "F1 = 91.2", that is almost always entity-level F1, not token accuracy. The two numbers can differ enormously for the same model.
Remember this
- Token accuracy can look excellent while entity-level performance is genuinely poor, because most tokens are the easy, non-entity majority class.
- An entity only counts as correctly extracted if its full span and type match exactly — near misses score zero.
- Precision, recall and F1, computed at the entity level, are the standard way to report NER performance honestly.
What to learn next
- Nested and overlapping entities — a structural case that makes even span matching more complex.
- BERTScore and embedding-based metrics — a softer notion of "close enough" used for other NLP tasks.
- Do your labels even agree? — checking whether your gold data itself is trustworthy before you trust any score against it.
Developer — Code and libraries.
Both metrics are short enough to implement directly, which is the best way to see exactly why they disagree.
Setup
Nothing to install. Standard library only.
Minimal runnable code
tokens = ["Reserve", "Bank", "of", "India", "cut", "rates", "."]
gold = ["B-ORG", "I-ORG","I-ORG","I-ORG", "O", "O", "O"]
# The model found almost the right entity -- it missed only the last word.
pred = ["B-ORG", "I-ORG","I-ORG","O", "O", "O", "O"]
def token_accuracy(gold, pred):
correct = sum(g == p for g, p in zip(gold, pred))
return correct / len(gold)
def bio_to_spans(tags):
spans, start, label = [], None, None
for i, tag in enumerate(tags + ["O"]):
if tag.startswith("B-"):
if start is not None:
spans.append((start, i, label))
start, label = i, tag[2:]
elif tag.startswith("I-") and label == tag[2:]:
continue
else:
if start is not None:
spans.append((start, i, label))
start, label = None, None
return set(spans)
def entity_f1(gold, pred):
gold_spans, pred_spans = bio_to_spans(gold), bio_to_spans(pred)
tp = len(gold_spans & pred_spans)
precision = tp / len(pred_spans) if pred_spans else 0.0
recall = tp / len(gold_spans) if gold_spans else 0.0
f1 = 2 * precision * recall / (precision + recall) if (precision + recall) else 0.0
return precision, recall, f1
acc = token_accuracy(gold, pred)
p, r, f1 = entity_f1(gold, pred)
print(f"Token accuracy: {acc:.0%}")
print(f"Entity precision: {p:.0%}")
print(f"Entity recall: {r:.0%}")
print(f"Entity F1: {f1:.0%}")
print()
print("Gold entity: ", bio_to_spans(gold))
print("Pred entity: ", bio_to_spans(pred))Token accuracy: 86%
Entity precision: 0%
Entity recall: 0%
Entity F1: 0%
Gold entity: {(0, 4, 'ORG')}
Pred entity: {(0, 3, 'ORG')}86% token accuracy and 0% entity F1, on the exact same prediction. The model missed one word out of seven, and by entity-level scoring, it found nothing correctly at all.
Line by line
bio_to_spans walks the tag sequence and groups consecutive B-/I- runs into (start, end, label) spans. Appending a sentinel "O" at the end guarantees any entity still open at the last real token gets closed and added, rather than silently dropped.
entity_f1 compares spans as a set, with == on the full tuple. (0, 4, 'ORG') and (0, 3, 'ORG') are different tuples — different end index — so they do not intersect in gold_spans & pred_spans, and the true positive count is zero.
Precision and recall come out identical here (0%) only because there is exactly one gold entity and one predicted entity. With more entities in play, a model can have high precision and low recall (confidently correct on what it does find, but misses many), or the reverse — the two numbers genuinely measure different failure modes.
Common mistakes
Reporting token accuracy as if it were the model's real-world usefulness. A NER system report should always lead with entity-level F1 (or precision and recall separately), not token accuracy — the gap between the two, as shown above, can be enormous.
Requiring exact type match without checking if that is the right call for your use case. Some applications care only about span boundaries, not the exact label — pulling out "a name" is more important than distinguishing PERSON from ORG. Decide this deliberately, and report both scores if it matters.
Comparing F1 scores across differently-labelled datasets. F1 is only meaningful relative to a specific, fixed gold standard. A 91% F1 on an easy, narrow entity schema is not comparable to an 85% F1 on a harder, broader one.
Using this hand-rolled evaluator on real data without double-checking edge cases. The seqeval library is the standard, well-tested implementation of exactly this logic, including handling for malformed tag sequences the toy code above does not defend against. Use it for anything beyond a learning exercise.
Try it yourself
Change pred so the model gets the entity boundary exactly right but predicts the wrong type — B-LOC I-LOC I-LOC I-LOC instead of B-ORG .... Rerun and check: does entity F1 count this as a match?
What to learn next
- Nested and overlapping entities — how span matching gets harder once entities can overlap.
- Entity linking — the next step once spans are correctly extracted.
- Exact match and F1 for question answering — the same exact-match-versus-partial-credit tension in a different task.
Researcher — Mathematics and papers.
Formal definitions
For a set of gold entity spans G and predicted entity spans P, both represented as (start, end, type) tuples:
precision = |G ∩ P| / |P|
recall = |G ∩ P| / |G|
F1 = 2 * precision * recall / (precision + recall)G ∩ Pis the set of predicted spans that exactly match a gold span in start, end, and type — this is "exact match" scoring, the CoNLL-2003 standard.- Micro-averaged F1 pools true positives, false positives and false negatives across all entity types before computing precision and recall; macro-averaged F1 computes per-type F1 first, then averages, weighting rare types equally with common ones.
Both are widely reported, and they can differ substantially on an imbalanced entity schema — a model that is excellent on the common PERSON type and poor on a rare PRODUCT type looks much better under micro-averaging than macro-averaging.
Partial-credit scoring schemes
Exact-match scoring is strict: a one-token boundary error scores identically to a completely wrong prediction. The MUC-5 and ACE evaluation conventions (used in earlier NER shared tasks before CoNLL standardised on exact match) instead score partial boundary overlap and type mismatch separately, producing metrics with names like "partial match" and "type match" that give credit for a correct span with a wrong type, or a correct type with a slightly wrong boundary. CoNLL-2003's stricter exact-match convention won out in practice, largely because it is unambiguous to compute and report, at the cost of being a harsher, more binary measure of quality than the phenomena it evaluates.
Why token accuracy is a poor proxy for entity F1
For a sequence with n tokens, of which a fraction e are inside some entity span (e is typically small — well under 20% for most news and web text), a trivial all-O classifier already achieves token accuracy of 1 - e, often above 85%, while achieving entity-level recall of exactly zero. This makes token accuracy nearly useless as a standalone metric for any corpus where entities are the minority class, which is the normal case. Token-level F1 restricted to the non-O classes is a partial fix, but still permits the kind of boundary near-miss demonstrated in the developer block to score partial credit that entity-level exact match correctly denies.
Key references
- Tjong Kim Sang, E. & De Meulder, F. (2003). Introduction to the CoNLL-2003 Shared Task. Establishes exact-match, entity-level F1 as the standard NER evaluation convention still used today.
- Chinchor, N. (1992). MUC-4 Evaluation Metrics. Defines the earlier partial-match scoring conventions.
- Nakayama, H.
seqeval: a Python framework for sequence labeling evaluation. The de facto standard library implementation of entity-level precision, recall and F1 for BIO/BILOU-tagged data.
Current state and open problems
Entity-level exact-match F1 remains the dominant metric, and its harshness on near-miss boundaries is a known, accepted trade-off for reproducibility and simplicity rather than an oversight. It becomes a genuinely open problem for nested and discontinuous entities (see Nested and overlapping entities), where "the set of correctly predicted spans" is no longer a straightforward concept, and different papers report incompatible variants of what counts as a match. There is no fully agreed-upon standard for nested-entity evaluation the way CoNLL-2003 exact match is agreed upon for flat NER, which makes cross-paper comparisons in that area considerably less reliable than the field's confident F1 numbers suggest.
What to learn next
- Do your labels even agree? — the ceiling your gold-standard data itself places on any F1 score.
- Exact match and F1 for question answering — the same exact-versus-partial tension applied to answer spans.
- Nested and overlapping entities — where entity-level matching itself becomes contested.