Sequence Labelling and Structure

BIO and BILOU tagging schemes

BIO and BILOU are ways of writing down where an entity starts, continues and ends using one tag per token, which is what turns entity spans into something a per-token classifier can learn.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

BIO tagging marks each word as the Beginning of an entity, Inside one, or Outside any entity. A multi-word name can then be labelled with one tag per word.

Think about parcel labels on a conveyor belt at a courier hub. A single parcel might actually be three boxes taped together. The label on the first box says "start of shipment". The labels on the rest say "part of the same shipment". That way the sorting machine knows where one shipment ends and the next begins.

Named entities have the same problem. "New York City" is one entity made of three separate word-tokens. Something has to mark which word starts it. Something has to mark which words continue it, and which words are not part of it at all.

Why it exists

A model that labels text word by word needs one tag per word. That is how per-token classifiers work. But an entity like "Reserve Bank of India" is four words acting as a single unit.

Say you tag each word only with its entity type. "Reserve" gets LOCATION. "of" gets LOCATION. "India" gets LOCATION. Now you cannot tell where one entity ends and the next one starts, when two entities sit right next to each other.

BIO tagging solves this with a small trick. It splits each type into two tags. One tag means "this word starts a new entity". The other means "this word continues the entity before it".

How it works

Every word gets a tag built from two parts: a position marker (B, I, or O) and, when relevant, the entity type.

  Words:   Priya   Sharma   flew   to   New     York    City    yesterday
  BIO:     B-PER   I-PER    O      O    B-LOC   I-LOC   I-LOC   O

  B = Beginning of an entity
  I = Inside (continuing) an entity
  O = Outside any entity

B-PER means "a person's name starts here". I-PER means "still the same person's name". The moment the tag drops to O, the entity is over.

BILOU adds two more letters for extra precision. L marks the Last word of a multi-word entity. U marks a Unit — a complete entity that is only one word long.

  Words:    Priya    Sharma   flew   to   New     York    City    yesterday
  BILOU:    B-PER    L-PER    O      O    B-LOC   I-LOC   L-LOC   O

Where you have already seen it

  • Every named entity recognition tool. spaCy, Hugging Face pipelines and commercial entity extractors represent training data this way behind the scenes. You never see the tags on screen.
  • Resume parsers. Extracting a candidate's name, company and location from free-text resumes relies on exactly this per-word labelling under the hood.
  • Redaction tools. Automatically blacking out names and account numbers in documents starts by tagging where each sensitive span begins and ends.

Remember this

  • BIO gives every word one tag, built from a position marker and an entity type.
  • B starts an entity, I continues one, O means "not part of any entity".
  • BILOU adds L (last word) and U (single-word entity) for extra precision, at the cost of a bigger tag set to learn.

What to learn next

Developer — Code and libraries.

Converting entity spans into BIO or BILOU tags is plain Python — no library, no download.

Setup

Nothing to install. Standard library only.

Minimal runnable code

bio_tags.py
tokens = ["Priya", "Sharma", "flew", "to", "New", "York", "City", "yesterday"]
# Entity spans in (start_index, end_index_exclusive, label) form -- the format
# most annotation tools export.
entities = [(0, 2, "PERSON"), (4, 7, "LOCATION")]

def to_bio(tokens, entities):
    tags = ["O"] * len(tokens)
    for start, end, label in entities:
        tags[start] = f"B-{label}"
        for i in range(start + 1, end):
            tags[i] = f"I-{label}"
    return tags

def to_bilou(tokens, entities):
    tags = ["O"] * len(tokens)
    for start, end, label in entities:
        if end - start == 1:
            tags[start] = f"U-{label}"
        else:
            tags[start] = f"B-{label}"
            for i in range(start + 1, end - 1):
                tags[i] = f"I-{label}"
            tags[end - 1] = f"L-{label}"
    return tags

bio = to_bio(tokens, entities)
bilou = to_bilou(tokens, entities)

print(f"{'token':10} {'BIO':12} {'BILOU':12}")
for tok, b, u in zip(tokens, bio, bilou):
    print(f"{tok:10} {b:12} {u:12}")
Output
token      BIO          BILOU       
Priya      B-PERSON     B-PERSON    
Sharma     I-PERSON     L-PERSON    
flew       O            O           
to         O            O           
New        B-LOCATION   B-LOCATION  
York       I-LOCATION   I-LOCATION  
City       I-LOCATION   L-LOCATION  
yesterday  O            O           

Line by line

entities is a list of (start, end, label) triples, using Python's normal half-open range convention — end is exclusive. This is the format spacy.training.Example and most annotation exports use, so it is worth getting comfortable with it early.

to_bio only needs to know where an entity starts. Everything from start + 1 to end gets the same I- tag, regardless of how long the entity is.

to_bilou special-cases single-word entities, tagging them U- instead of B-...L-. A one-word entity has no "inside" or genuine "last" — it is both the first and only word, so BILOU gives it its own dedicated tag rather than forcing a B- tag that is never followed by an I- or L-.

Does BIO actually break on adjacent entities?

A common claim is that BIO cannot tell two back-to-back entities apart. Test it directly.

adjacent.py
# Two back-to-back ORG entities with no gap between them.
tokens = ["SBI", "HDFC", "Bank", "both", "raised", "rates"]
entities = [(0, 1, "ORG"), (1, 3, "ORG")]   # "SBI" and "HDFC Bank", adjacent

def to_bio(tokens, entities):
    tags = ["O"] * len(tokens)
    for start, end, label in entities:
        tags[start] = f"B-{label}"
        for i in range(start + 1, end):
            tags[i] = f"I-{label}"
    return tags

for tok, tag in zip(tokens, to_bio(tokens, entities)):
    print(f"{tok:8} {tag}")
Output
SBI      B-ORG
HDFC     B-ORG
Bank     I-ORG
both     O
raised   O
rates    O

BIO handles this correctly: a second B-ORG immediately after the first one unambiguously starts a new entity. The real reason BILOU exists is different, and it matters more in practice: Ratinov & Roth (2009) found that models trained to predict BILOU tags make noticeably fewer boundary mistakes than models trained on BIO, because a dedicated "last word" tag gives the model a direct, learnable signal for where an entity ends, instead of leaving "end of entity" to be inferred indirectly from the next tag not being I-.

Common mistakes

Writing I- tags without ever writing the matching B-. A tag sequence like O I-PER I-PER is invalid BIO — it claims to continue an entity that never started. This happens constantly with a model's raw predictions, not with gold data, and needs a repair step before spans can be extracted.

Forgetting entity types must match across a B-/I- run. B-PERSON I-LOCATION is nonsensical and should never appear in training data. If your conversion code allows it, check your span boundaries.

Mixing tagging schemes between training and evaluation. A model trained on BILOU and evaluated with a BIO-only decoder will look far worse than it actually is, because the decoder does not know what to do with L- and U- tags.

Try it yourself

Take three entities that touch each other with zero gap — [(0, 1, "ORG"), (1, 2, "ORG"), (2, 4, "PERSON")] over four tokens — and run both to_bio and to_bilou on them. Check that every entity boundary is still recoverable from each output.

What to learn next

Researcher — Mathematics and papers.

Formal definition

Given tokens x = (x_1, ..., x_n) and a set of non-overlapping entity spans {(s_k, e_k, l_k)}, the BIO encoding assigns:

text
y_i = B-l_k   if i = s_k
y_i = I-l_k   if s_k < i < e_k
y_i = O       otherwise
  • s_k, e_k are the start and end token indices of entity k (e_k exclusive).
  • l_k is the entity type of span k.

This makes the tag set size 2L + 1 for L entity types (B- and I- per type, plus O). BILOU/BIOES uses 4L + 1 (B-, I-, L-/E-, U-/S- per type, plus O), trading a larger output space for a more explicit boundary signal.

Why the scheme choice affects learned models, not only notation

For a model that scores each tag independently or with a linear-chain transition matrix (an HMM or a CRF — see Conditional random fields), the decision boundary the model has to learn differs by scheme. Under BIO, "is this the end of the entity" is not a tag at all — it must be inferred from the next token's tag being O or B-, which requires the model to correctly predict a token it has not conditioned on the label of. Under BILOU, "end of entity" is the tag itself, observable at the current position.

Ratinov & Roth (2009) report consistent F1 gains from switching a linear-chain CRF NER system from BIO to BILOU, holding features and data fixed, attributing the gain to exactly this: cleaner boundary supervision, not additional information. The effect shrinks, though it does not fully disappear, for neural sequence taggers with a wide effective context window, since a BiLSTM or transformer encoder can already condition each position's representation on tokens arbitrarily far to the right.

Tag consistency constraints

A valid BIO sequence obeys: an I-l tag may only follow B-l or I-l for the same type l; it can never follow O or a tag of a different type. Enforcing this is exactly the transition-scoring role a CRF layer plays on top of a neural encoder — see the next lesson for the mechanism.

Raw per-token softmax predictions from a plain classifier (no CRF layer) can and do violate this constraint. A standard repair heuristic converts any illegal I-l immediately following O or a mismatched type into B-l, salvaging a span rather than discarding the prediction outright.

Key references

  • Ramshaw, L. & Marcus, M. (1995). Text Chunking Using Transformation-Based Learning. Introduces the IOB (equivalent to BIO) scheme for chunking.
  • Ratinov, L. & Roth, D. (2009). Design Challenges and Misconceptions in Named Entity Recognition. CoNLL. The paper establishing BILOU's practical edge over BIO for feature-based sequence models.
  • Tjong Kim Sang, E. & De Meulder, F. (2003). Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition. Standardised BIO-tagged benchmark data still used to evaluate NER systems today.

Current state and open problems

BIO and BILOU remain the default output representation even for transformer-based NER systems, largely because the CoNLL-2003 and OntoNotes benchmarks that the field measures itself against were released in this format, and switching representations mid-field would break decades of comparability.

Neither scheme represents nested or discontinuous entities, which is a real limitation for biomedical and legal text where spans genuinely overlap. Span-based and set-prediction formulations (see Nested and overlapping entities) sidestep the token-tagging representation entirely rather than extending it, which is itself evidence that BIO/BILOU are treated as a practical convenience rather than a principled representation of entity structure.

What to learn next