Sequence Labelling and Structure

Why token accuracy lies for NER

A tagger can get almost every individual tag right and still miss every single entity, because entity-level correctness needs the whole span to match, not only most of its tokens.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Token accuracy counts individual word-tags correct or wrong. But an entity is only useful if its whole span is right — one wrong word can wreck an otherwise perfect answer.

Think about a delivery address. A courier gets the house number, street, area, city and PIN code all correct, except the house number is off by one digit. The parcel does not arrive at 95% of the right place. It arrives at the wrong house, completely. Almost right is not right.

Named entities work the same way. Say a model tags "Reserve Bank of" correctly but misses "India". The entity it extracted is not "Reserve Bank of India" with a small error. It is a different, wrong entity altogether.

Why it exists

Most words in a sentence are not part of any entity, tagged as plain O. A model can get nearly every individual word-tag right by predicting O most of the time, and being roughly correct on entity boundaries. That makes token accuracy look excellent, even when the model is actually bad at the one thing that matters: getting complete entity spans right.

Entity-level evaluation fixes this by scoring whole spans, not individual tags. A predicted entity only counts as correct if its exact boundaries and its exact type both match the true entity. Off-by-one-word does not get partial credit.

How it works

  Gold entity:      [ Reserve  Bank  of  India ]  ORG
  Model's guess:     [ Reserve  Bank  of ]        ORG
                                         ^
                                         missed "India" -- boundary is wrong

  Token accuracy:    3 out of 4 entity words tagged correctly  =  75% looks fine
  Entity-level score: the predicted span does not exactly match the gold span
                        =  counted as a complete miss, both for precision and recall

Precision asks: of the entities the model claimed to find, how many were actually correct? Recall asks: of the entities that actually exist, how many did the model find? F1 combines the two into a single number. It only rewards exact, complete matches.

Where you have already seen it

  • Resume parsers. A model that extracts "John Sm" instead of "John Smith" has not extracted 90% of a name. It has extracted a name for the wrong person entirely. Treating that as "close enough" causes real downstream errors.
  • Redaction software. Missing the last character of an account number defeats the entire point of redacting it, no matter how many other characters were caught.
  • Any leaderboard reporting NER results. When a paper reports "F1 = 91.2", that is almost always entity-level F1, not token accuracy. The two numbers can differ enormously for the same model.

Remember this

  • Token accuracy can look excellent while entity-level performance is genuinely poor, because most tokens are the easy, non-entity majority class.
  • An entity only counts as correctly extracted if its full span and type match exactly — near misses score zero.
  • Precision, recall and F1, computed at the entity level, are the standard way to report NER performance honestly.

What to learn next

Developer — Code and libraries.

Both metrics are short enough to implement directly, which is the best way to see exactly why they disagree.

Setup

Nothing to install. Standard library only.

Minimal runnable code

ner_eval.py
tokens = ["Reserve", "Bank", "of", "India", "cut", "rates", "."]
gold =    ["B-ORG",  "I-ORG","I-ORG","I-ORG", "O",   "O",     "O"]
# The model found almost the right entity -- it missed only the last word.
pred =    ["B-ORG",  "I-ORG","I-ORG","O",     "O",   "O",     "O"]

def token_accuracy(gold, pred):
    correct = sum(g == p for g, p in zip(gold, pred))
    return correct / len(gold)

def bio_to_spans(tags):
    spans, start, label = [], None, None
    for i, tag in enumerate(tags + ["O"]):
        if tag.startswith("B-"):
            if start is not None:
                spans.append((start, i, label))
            start, label = i, tag[2:]
        elif tag.startswith("I-") and label == tag[2:]:
            continue
        else:
            if start is not None:
                spans.append((start, i, label))
            start, label = None, None
    return set(spans)

def entity_f1(gold, pred):
    gold_spans, pred_spans = bio_to_spans(gold), bio_to_spans(pred)
    tp = len(gold_spans & pred_spans)
    precision = tp / len(pred_spans) if pred_spans else 0.0
    recall = tp / len(gold_spans) if gold_spans else 0.0
    f1 = 2 * precision * recall / (precision + recall) if (precision + recall) else 0.0
    return precision, recall, f1

acc = token_accuracy(gold, pred)
p, r, f1 = entity_f1(gold, pred)

print(f"Token accuracy:  {acc:.0%}")
print(f"Entity precision: {p:.0%}")
print(f"Entity recall:    {r:.0%}")
print(f"Entity F1:         {f1:.0%}")
print()
print("Gold entity: ", bio_to_spans(gold))
print("Pred entity: ", bio_to_spans(pred))
Output
Token accuracy:  86%
Entity precision: 0%
Entity recall:    0%
Entity F1:         0%

Gold entity:  {(0, 4, 'ORG')}
Pred entity:  {(0, 3, 'ORG')}

86% token accuracy and 0% entity F1, on the exact same prediction. The model missed one word out of seven, and by entity-level scoring, it found nothing correctly at all.

Line by line

bio_to_spans walks the tag sequence and groups consecutive B-/I- runs into (start, end, label) spans. Appending a sentinel "O" at the end guarantees any entity still open at the last real token gets closed and added, rather than silently dropped.

entity_f1 compares spans as a set, with == on the full tuple. (0, 4, 'ORG') and (0, 3, 'ORG') are different tuples — different end index — so they do not intersect in gold_spans & pred_spans, and the true positive count is zero.

Precision and recall come out identical here (0%) only because there is exactly one gold entity and one predicted entity. With more entities in play, a model can have high precision and low recall (confidently correct on what it does find, but misses many), or the reverse — the two numbers genuinely measure different failure modes.

Common mistakes

Reporting token accuracy as if it were the model's real-world usefulness. A NER system report should always lead with entity-level F1 (or precision and recall separately), not token accuracy — the gap between the two, as shown above, can be enormous.

Requiring exact type match without checking if that is the right call for your use case. Some applications care only about span boundaries, not the exact label — pulling out "a name" is more important than distinguishing PERSON from ORG. Decide this deliberately, and report both scores if it matters.

Comparing F1 scores across differently-labelled datasets. F1 is only meaningful relative to a specific, fixed gold standard. A 91% F1 on an easy, narrow entity schema is not comparable to an 85% F1 on a harder, broader one.

Using this hand-rolled evaluator on real data without double-checking edge cases. The seqeval library is the standard, well-tested implementation of exactly this logic, including handling for malformed tag sequences the toy code above does not defend against. Use it for anything beyond a learning exercise.

Try it yourself

Change pred so the model gets the entity boundary exactly right but predicts the wrong type — B-LOC I-LOC I-LOC I-LOC instead of B-ORG .... Rerun and check: does entity F1 count this as a match?

What to learn next

Researcher — Mathematics and papers.

Formal definitions

For a set of gold entity spans G and predicted entity spans P, both represented as (start, end, type) tuples:

text
precision = |G ∩ P| / |P|
recall    = |G ∩ P| / |G|
F1        = 2 * precision * recall / (precision + recall)
  • G ∩ P is the set of predicted spans that exactly match a gold span in start, end, and type — this is "exact match" scoring, the CoNLL-2003 standard.
  • Micro-averaged F1 pools true positives, false positives and false negatives across all entity types before computing precision and recall; macro-averaged F1 computes per-type F1 first, then averages, weighting rare types equally with common ones.

Both are widely reported, and they can differ substantially on an imbalanced entity schema — a model that is excellent on the common PERSON type and poor on a rare PRODUCT type looks much better under micro-averaging than macro-averaging.

Partial-credit scoring schemes

Exact-match scoring is strict: a one-token boundary error scores identically to a completely wrong prediction. The MUC-5 and ACE evaluation conventions (used in earlier NER shared tasks before CoNLL standardised on exact match) instead score partial boundary overlap and type mismatch separately, producing metrics with names like "partial match" and "type match" that give credit for a correct span with a wrong type, or a correct type with a slightly wrong boundary. CoNLL-2003's stricter exact-match convention won out in practice, largely because it is unambiguous to compute and report, at the cost of being a harsher, more binary measure of quality than the phenomena it evaluates.

Why token accuracy is a poor proxy for entity F1

For a sequence with n tokens, of which a fraction e are inside some entity span (e is typically small — well under 20% for most news and web text), a trivial all-O classifier already achieves token accuracy of 1 - e, often above 85%, while achieving entity-level recall of exactly zero. This makes token accuracy nearly useless as a standalone metric for any corpus where entities are the minority class, which is the normal case. Token-level F1 restricted to the non-O classes is a partial fix, but still permits the kind of boundary near-miss demonstrated in the developer block to score partial credit that entity-level exact match correctly denies.

Key references

  • Tjong Kim Sang, E. & De Meulder, F. (2003). Introduction to the CoNLL-2003 Shared Task. Establishes exact-match, entity-level F1 as the standard NER evaluation convention still used today.
  • Chinchor, N. (1992). MUC-4 Evaluation Metrics. Defines the earlier partial-match scoring conventions.
  • Nakayama, H. seqeval: a Python framework for sequence labeling evaluation. The de facto standard library implementation of entity-level precision, recall and F1 for BIO/BILOU-tagged data.

Current state and open problems

Entity-level exact-match F1 remains the dominant metric, and its harshness on near-miss boundaries is a known, accepted trade-off for reproducibility and simplicity rather than an oversight. It becomes a genuinely open problem for nested and discontinuous entities (see Nested and overlapping entities), where "the set of correctly predicted spans" is no longer a straightforward concept, and different papers report incompatible variants of what counts as a match. There is no fully agreed-upon standard for nested-entity evaluation the way CoNLL-2003 exact match is agreed upon for flat NER, which makes cross-paper comparisons in that area considerably less reliable than the field's confident F1 numbers suggest.

What to learn next

What to learn next

These follow on from what you just read.

  • Sequence Labelling and Structure

    Entity linking

    Entity linking connects a name in text to one specific real-world entity in a knowledge base, resolving the fact that many different things share the same name.

  • Sequence Labelling and Structure

    Coreference resolution

    Coreference resolution figures out which earlier name a pronoun like "she" or "it" refers back to, without which a machine reads a whole paragraph as disconnected sentences about strangers.

  • Sequence Labelling and Structure

    Relation extraction

    Relation extraction turns two entities and a sentence into a structured fact, such as (Zoho, acquired, a startup), which is what lets unstructured text feed a database or a knowledge graph.