Sequence Labelling and Structure

Training NER on your own entities

Off-the-shelf named entity recognisers only know the entity types they were trained on, so extracting your own categories means labelling examples and training a small model from scratch.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Training a custom NER model teaches a tagger to recognise the entity types you care about. Off-the-shelf models only know a handful of generic categories.

Think about training a new security guard at a housing society gate. A generic guard knows to stop unfamiliar people. But they do not yet know which residents live where, or which vehicles belong to the society, or which vendors are expected on which day. You have to show them examples until the pattern sinks in.

A generic named entity recognizer knows generic categories: person, organisation, location, date. It does not know that "RRN" is a transaction reference number. It does not know that "SKU-4471" is a product code. To recognise entity types specific to your work, you have to show it labelled examples until it learns the pattern.

Why it exists

Pretrained NER models are trained on general news and web text. They tag broad categories like PERSON, ORG, GPE (a country, city or state). This works well for a general-purpose news reader. It does nothing for a hospital that needs to pull out drug names and dosages. It does nothing for a fintech company that needs UPI transaction IDs either.

The entity types a business actually needs are usually specific to that business. No generic model was ever pretrained for "SKU code" or "policy number". Those categories only exist inside one company's documents.

Training your own model means labelling a set of examples yourself. It teaches the model your categories directly, instead of hoping a general model happens to notice a pattern it was never shown.

How it works

You show the model examples of text with the correct spans marked, and it gradually adjusts itself to reproduce those spans on new, unseen text.

  Training example:
    "Priya works at Zoho in Chennai."
     ^^^^^          ^^^^     ^^^^^^^
     PERSON         ORG      GPE

  The model sees many such examples and learns patterns:
  a capitalised word after "works at" often marks an organisation,
  a capitalised word after "in" often marks a place.

  New sentence:            "Kavya works at Wipro in Mumbai."
  Model's prediction:        PERSON        ORG     GPE
                              (never seen "Kavya" or "Wipro" during training)

The model is not memorising names. It is learning the surrounding patterns that tend to signal each entity type, so it can label names it has never seen before.

Where you have already seen it

  • Resume screening tools that pull out skills, degrees and company names specific to a hiring pipeline — categories no general-purpose model ships with by default.
  • Legal document review software that extracts clause types and party names unique to contract law.
  • Customer support ticket routing. A model tags product names and error codes specific to one company's product line.

Remember this

  • Generic NER models only know generic categories — person, organisation, location, date.
  • Business-specific categories need labelled examples and a model trained on them.
  • A well-trained model generalises the pattern around an entity, not the exact names it saw during training.

What to learn next

Developer — Code and libraries.

spaCy's blank pipeline can train a small NER model from a handful of examples on CPU in seconds. A real project needs hundreds of examples per entity type — this is a demonstration of the mechanics, not production-scale training.

Setup

bash
pip install spacy

No pretrained model download needed — spacy.blank("en") starts from nothing but tokenization rules.

Minimal runnable code

train_ner.py
import random
import spacy
from spacy.training import Example

def make_example(text, spans):
    """spans: list of (substring, label). Finds character offsets automatically,
    which avoids the single most common bug in hand-built NER data: miscounted
    character positions."""
    entities = []
    for substring, label in spans:
        start = text.index(substring)
        entities.append((start, start + len(substring), label))
    return (text, {"entities": entities})

RAW_DATA = [
    ("Priya works at Zoho in Chennai.", [("Priya", "PERSON"), ("Zoho", "ORG"), ("Chennai", "GPE")]),
    ("Ravi joined Freshworks last year.", [("Ravi", "PERSON"), ("Freshworks", "ORG")]),
    ("Meera moved to Bengaluru for a job at Infosys.", [("Meera", "PERSON"), ("Bengaluru", "GPE"), ("Infosys", "ORG")]),
    ("Arjun founded a startup called Kutumb.", [("Arjun", "PERSON"), ("Kutumb", "ORG")]),
]
TRAIN_DATA = [make_example(text, spans) for text, spans in RAW_DATA]

nlp = spacy.blank("en")
ner = nlp.add_pipe("ner")
for _, ann in TRAIN_DATA:
    for start, end, label in ann["entities"]:
        ner.add_label(label)

optimizer = nlp.initialize()
random.seed(0)
for epoch in range(30):
    random.shuffle(TRAIN_DATA)
    losses = {}
    for text, ann in TRAIN_DATA:
        example = Example.from_dict(nlp.make_doc(text), ann)
        nlp.update([example], sgd=optimizer, losses=losses)
    if epoch % 10 == 0 or epoch == 29:
        print(f"epoch {epoch:2d}  loss {losses['ner']:.2f}")

doc = nlp("Kavya works at Wipro in Mumbai.")
for ent in doc.ents:
    print(ent.text, "->", ent.label_)
Output
epoch  0  loss 27.17
epoch 10  loss 0.00
epoch 20  loss 0.00
epoch 29  loss 0.00
Kavya -> PERSON
Wipro -> ORG
Mumbai -> GPE

The starting loss at epoch 0 depends on spaCy's random weight initialisation, so your exact number will differ from 27.17 — expect it somewhere in a similar range. The model correctly tags all three entities in a sentence it never saw during training.

Line by line

make_example finds character offsets with .index() instead of counting by hand. Counting characters manually is the single most common source of bugs in hand-built NER training data — see Common mistakes below for what happens when it goes wrong.

nlp.initialize() sets up random weights and returns an optimizer. This has to run after every label has been registered with ner.add_label, or the output layer will be the wrong size for labels added later.

The loss drops to exactly 0.00 by epoch 10. With only four training sentences, the model can essentially memorise them — this is overfitting, and it is expected and harmless for a demonstration this small. On a real dataset with hundreds of varied examples, loss dropping to exactly zero this fast would be a warning sign, not a good result.

doc.ents returns the entity spans the trained pipeline found, each with .text and .label_.

Common mistakes

Miscounting character offsets by hand. This is common enough to demonstrate directly:

misaligned.py
import spacy
from spacy.training import offsets_to_biluo_tags

nlp = spacy.blank("en")
text = "Priya works at Zoho in Chennai."
# Manually counted offsets -- and miscounted by one character (should be 15, 19).
entities = [(16, 20, "ORG")]
tags = offsets_to_biluo_tags(nlp.make_doc(text), entities)
print(tags)
Output
UserWarning: [W030] Some entities could not be aligned in the text "Priya works at
Zoho in Chennai." with entities "[(16, 20, 'ORG')]". Use
`spacy.training.offsets_to_biluo_tags(nlp.make_doc(text), entities)` to check the
alignment. Misaligned entities ('-') will be ignored during training.
['O', 'O', 'O', '-', 'O', 'O', 'O']

The '-' tag marks the position spaCy could not align — this same silent substitution is what happens inside real training on misaligned data.

spaCy does not crash — it silently drops the misaligned entity and trains on whatever is left. A model trained this way will look like it is learning something, while quietly missing entities in its training data. Building offsets with .index(), as the main example does, avoids this entirely.

Forgetting to call ner.add_label for every label before nlp.initialize(). The output layer's size is fixed at initialization time. A label added afterwards will not be predicted correctly.

Judging a model by four training examples. This lesson's model "works" because the test sentence closely mirrors the training pattern. A production model needs hundreds of examples per entity type, covering varied sentence structures, or it will fail the moment real text stops resembling the training set.

Not holding out any evaluation data. Every example here was used for training. A real project always keeps a separate labelled set the model never trains on, purely to measure how well it generalises — see the next lesson on evaluation for why token-level accuracy alone can be misleading.

Try it yourself

Add a fifth training sentence with a new entity type, such as PRODUCT, and register it with ner.add_label("PRODUCT") before training. Check whether the model picks up the new label with this little data — and notice how much less confident (and how much more epochs) it needs compared to the well-represented types.

What to learn next

Researcher — Mathematics and papers.

Training objective

spaCy's transition-based NER parser (Lample et al., 2016 architecture lineage) frames entity recognition as a sequence of actions over a stack-and-buffer state, trained to imitate a gold action sequence derived from the labelled spans. The per-example loss backpropagated through nlp.update is a structured prediction loss over these actions, not a simple per-token cross-entropy — which is why loss values are not directly comparable across different pipeline architectures.

For a pure token-classification formulation instead (common with transformer-based NER), the objective is standard cross-entropy over BIO/BILOU tags per position:

text
L = -(1/n) * sum over i=1..n of log P(y_i | x)
  • n is the number of tokens.
  • P(y_i | x) is the model's predicted probability of the correct tag at position i, given the whole input x.

Sample efficiency and data requirements

Learning curves for NER (Yadav & Bethard, 2018 survey; and empirically across most reported CoNLL-style fine-tuning work) show F1 rising steeply from roughly 50 to several hundred labelled sentences per entity type, then flattening — additional data beyond a few thousand examples yields diminishing returns for a single, narrowly-scoped entity schema. This has a direct practical consequence: for a business-specific schema with 3 to 5 entity types, a few hundred carefully labelled sentences, augmented with active learning to prioritise uncertain examples, typically outperforms a much larger but unfocused labelling effort.

Active learning — selecting which unlabelled examples to hand-label next based on model uncertainty, rather than labelling at random — consistently reduces the number of examples needed to reach a target F1 by a factor of 2 to 4 in reported NER studies (Shen et al., 2004; Settles, 2009 survey), and is the standard approach in production annotation pipelines with a limited labelling budget.

Fine-tuning a pretrained encoder versus training from scratch

Training NER from a blank pipeline, as in the developer block, starts every weight at random. Fine-tuning a pretrained encoder such as BERT for the same task — replacing only the final classification layer, or adding one on top of frozen or lightly-tuned encoder weights — consistently needs far less labelled data to reach a comparable F1, because the encoder already carries general language structure (syntax, common entity shapes) learned from unlabelled pretraining. This is the same transfer-learning argument developed at length in Fine-tuning BERT for classification, applied here to token-level rather than sequence-level output.

Key references

  • Lample, G. et al. (2016). Neural Architectures for Named Entity Recognition. arXiv:1603.01360
  • Settles, B. (2009). Active Learning Literature Survey. University of Wisconsin-Madison Computer Sciences Technical Report 1648.
  • Shen, D. et al. (2004). Multi-Criteria-based Active Learning for Named Entity Recognition. ACL.
  • Yadav, V. & Bethard, S. (2018). A Survey on Recent Advances in Named Entity Recognition from Deep Learning Models. COLING.

Current state and open problems

The practical bottleneck in most custom NER projects is not model architecture — pretrained-encoder fine-tuning is a solved recipe — but annotation quality and consistency. Inter-annotator agreement on span boundaries and entity type assignment is often the ceiling on achievable F1, well before model capacity becomes the limiting factor; see Do your labels even agree? for how that ceiling is measured. Few-shot and LLM-prompted extraction (asking a large language model to extract entities directly, with no fine-tuning) are increasingly used to bootstrap an initial labelled set faster than manual annotation alone, though they still require human review before the result is trustworthy enough to train on.

What to learn next