Sequence Labelling and Structure

Nested and overlapping entities

A single BIO tag sequence can only give each word one label at a time, so an entity hiding inside another entity needs a different representation entirely.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Nested entities are one entity sitting entirely inside another. A single tag-per-word scheme cannot represent that, because it only lets each word carry one label.

Think about Russian nesting dolls. Open the big doll and there is a smaller, complete doll inside it. It is not a piece of the big doll — it is a whole separate doll in its own right. "Bank of Maharashtra" contains "Maharashtra" the same way. One full entity, an organisation, has another full entity, a place name, nested completely inside it.

A tag sequence gives each word exactly one label. It has no way to say "this word is part of an organisation AND part of a location at the same time." Something has to give.

Why it exists

BIO tagging, covered earlier in this section, gives each word exactly one tag. That works well when entities never overlap. It breaks the moment they do.

"Bank of Maharashtra" is an organisation. "Maharashtra" alone is also a real, separately useful entity — a state name, tagged LOCATION on its own in plenty of other sentences. Both readings are correct. Both are useful for different downstream tasks: one for identifying companies, one for identifying places mentioned in text.

Plain BIO tagging forces a choice. Whichever entity gets labelled second overwrites the tags of the one labelled first, and the other one silently disappears.

How it works

  Tokens:        The   Bank   of    Maharashtra   raised   rates   .
  Entity 1:            [------ ORG -----------]
  Entity 2:                          [ GPE ]

  A single BIO layer can only hold one tag per word.
  Whichever entity gets applied last wins -- the other is lost:

  Single layer:  O     B-ORG  I-ORG  B-GPE        O        O       O
                                     ^
                                     The ORG tag that used to be here is gone.

The fix used in practice keeps more than one layer of tags — one full BIO sequence per entity type. Overlapping entities never have to fight over the same slot.

Where you have already seen it

  • Legal documents. "The Reserve Bank of India Act, 1934" is a law's name. It also contains "Reserve Bank of India", an organisation, and "1934", a date — both nested inside one span.
  • Biomedical text. A drug name is often nested inside a longer phrase describing a drug interaction or dosage, a well-known hard case for NER research.
  • Company filings, where a subsidiary's name is nested inside its parent company's full legal name.

Remember this

  • A nested entity sits entirely inside another entity's span, and both are genuinely correct at the same time.
  • A single BIO tag sequence can only hold one label per word, so it cannot represent this on its own.
  • The standard fix is either multiple independent tag layers, or a model that scores spans directly instead of tagging tokens.

What to learn next

Developer — Code and libraries.

The failure mode, and the standard fix, are both plain Python — no model needed to see the structural problem.

Setup

Nothing to install. Standard library only.

Minimal runnable code

nested_entities.py
tokens = ["The", "Bank", "of", "Maharashtra", "raised", "rates", "."]

# Two real, overlapping entities. "Bank of Maharashtra" is an organisation.
# "Maharashtra" alone is also a location, nested entirely inside the first one.
entities = [
    (1, 4, "ORG"),   # "Bank of Maharashtra"
    (3, 4, "GPE"),   # "Maharashtra" -- overlaps with the ORG above
]

def to_bio(tokens, entities):
    tags = ["O"] * len(tokens)
    for start, end, label in entities:
        tags[start] = f"B-{label}"
        for i in range(start + 1, end):
            tags[i] = f"I-{label}"
    return tags

# A single BIO layer can only hold one label per token. Building it entity by
# entity, the second entity silently overwrites the first one's tags.
tags = to_bio(tokens, entities)
print("Single BIO layer (the ORG tag for 'Maharashtra' has been lost):")
for tok, tag in zip(tokens, tags):
    print(f"  {tok:12} {tag}")

# The standard fix: one independent BIO layer per entity TYPE, not per entity.
# Nothing gets overwritten because ORG and GPE never share a layer.
def to_layered_bio(tokens, entities):
    labels = sorted({label for _, _, label in entities})
    layers = {label: ["O"] * len(tokens) for label in labels}
    for start, end, label in entities:
        layers[label][start] = f"B-{label}"
        for i in range(start + 1, end):
            layers[label][i] = f"I-{label}"
    return layers

layers = to_layered_bio(tokens, entities)
print("\nLayered BIO (one row per entity type, nothing lost):")
for label, row in layers.items():
    print(f"  {label:4}", row)
Output
Single BIO layer (the ORG tag for 'Maharashtra' has been lost):
  The          O
  Bank         B-ORG
  of           I-ORG
  Maharashtra  B-GPE
  raised       O
  rates        O
  .            O

Layered BIO (one row per entity type, nothing lost):
  GPE  ['O', 'O', 'O', 'B-GPE', 'O', 'O', 'O']
  ORG  ['O', 'B-ORG', 'I-ORG', 'I-ORG', 'O', 'O', 'O']

Line by line

to_bio applies entities in list order, and the second one silently wins. Notice "Maharashtra" ends up tagged only B-GPE — the fact that it was also part of an ORG span is gone completely, with no error or warning. This is exactly why nested entities need to be caught during data design, not discovered by a training run behaving strangely.

to_layered_bio groups tags by entity type into separate lists. Because ORG and GPE each get their own independent BIO sequence, both to_layered_bio(...)["ORG"] and to_layered_bio(...)["GPE"] can mark position 3 ("Maharashtra") as part of their respective entity, with no conflict.

This is a data representation fix, not a modelling fix. Training a real tagger on layered BIO means training one tagger per layer — either as separate models, or as separate output heads sharing one encoder — and combining their predictions afterward.

Common mistakes

Not noticing the overwrite at all. Because to_bio runs without error or warning, a nested-entity dataset converted this way looks fine until someone checks the output carefully. Always spot-check converted training data for entities you know should be there.

Assuming a bigger model architecture fixes a data representation problem. Feeding overwritten, single-layer BIO tags into even the best transformer will not recover the lost entity — the information was destroyed before training started. Fix the representation, not the model.

Choosing layered BIO without checking your annotation tool supports it. Many popular annotation tools export single-layer spans by default. Confirm your export format actually preserves overlapping annotations before building a pipeline around them.

Try it yourself

Add a third, non-overlapping entity — (4, 5, "ACTION") for "raised" — to the entities list and rerun both functions. Confirm it lands cleanly in its own new layer, alongside ORG and GPE.

What to learn next

Researcher — Mathematics and papers.

Why token-tagging cannot represent nested structure

A per-token tagging formulation assigns each position i a single label y_i from a fixed tag set. Nested and overlapping entities require assigning position i to more than one span simultaneously, which is structurally outside what a function i -> y_i can express, regardless of tag set size or scheme (BIO, BILOU, or otherwise). This is not a training limitation — it is a representational one, fixed by the choice of output space before any learning happens.

Reported nesting rates make this a real, not theoretical, problem: Finkel & Manning (2009) report roughly 17% of entities in the GENIA biomedical corpus are nested inside another entity; legal and clinical corpora report comparable or higher rates.

Representations that support nesting

Layered tagging (the developer-block approach), formalised for learning as a multi-task or multi-head problem: one tagger (or output head) per entity type, sharing an encoder. Cost scales linearly in the number of types, and the layers are trained largely independently, which misses interactions between overlapping entity types.

Span-based classification. Enumerate candidate spans (i, j) up to a maximum length, represent each with pooled or boundary-token features, and classify each span independently as an entity type or "not an entity" (Sohrab & Miwa, 2018; Eberts & Ulges, 2019, SpERT). Complexity is O(n * L) candidate spans for max span length L, each independently classified — nesting is free, since spans are scored independently rather than tagged sequentially.

Sequence-to-sequence / set-generation formulations. Frame NER as generating a set or sequence of (span, type) pairs directly, e.g. with a pointer network or an autoregressive decoder (Yan et al., 2021, unified generative NER). This naturally handles nesting and even discontinuous entities, at the cost of a more complex, typically slower decoding procedure than tagging.

Biaffine / parsing-inspired scoring. Yu, Bohnet & Poesio (2020) score every (start, end) span pair with a biaffine classifier borrowed from dependency parsing, achieving strong results on nested benchmarks (ACE 2004, ACE 2005, GENIA) while keeping inference a single forward pass rather than autoregressive generation.

Complexity comparison

ApproachCandidate spaceHandles nestingTypical inference cost
Single-layer BIOO(n) tagsNoO(n), Viterbi if CRF
Layered BIOO(n) tags per layerYes, across layersO(n * types)
Span classificationO(n * L) spansYesO(n * L)
Biaffine span scoringO(n^2) span pairsYesO(n^2)
Generative / set predictionUnboundedYes, plus discontinuousSequential decoding

Key references

  • Finkel, J. & Manning, C. (2009). Nested Named Entity Recognition. EMNLP.
  • Sohrab, M. & Miwa, M. (2018). Deep Exhaustive Model for Nested Named Entity Recognition. EMNLP.
  • Eberts, M. & Ulges, A. (2019). Span-based Joint Entity and Relation Extraction with Transformer Pre-training. arXiv:1909.07755 — SpERT.
  • Yu, J., Bohnet, B. & Poesio, M. (2020). Named Entity Recognition as Dependency Parsing. arXiv:2005.07150
  • Yan, H. et al. (2021). A Unified Generative Framework for Various NER Subtasks. arXiv:2106.01223

Current state and open problems

Span-based and biaffine methods are now the standard approach for nested NER benchmarks, having displaced layered tagging in research settings, largely because they scale better to many overlapping types and unify flat and nested NER under one formulation. In production, layered tagging remains common regardless, because it reuses existing token-classification infrastructure, and many real applications only need two or three entity types with modest, predictable overlap — where the added complexity of span or biaffine scoring is not worth the engineering cost. The open research question is less about representation now and more about data: nested-entity annotation is markedly harder and slower for human annotators than flat annotation, and labelled nested-NER corpora remain small relative to flat-NER benchmarks like CoNLL-2003.

What to learn next