Named entity recognition
Named entity recognition finds the names inside a sentence — people, places, companies, dates — and the honest way to score it counts whole entities, not individual words.
- 19 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Named entity recognition is finding the names inside a piece of text and saying what kind of name each one is.
The analogy you have already lived
A wedding invitation arrives on WhatsApp. You read it once and then do three things.
You copy the couple's names into a message to your cousin. You copy the venue into Maps. You copy the date into your calendar.
You did not read that invitation word by word with equal care. You hunted for three kinds of thing — a person, a place, a date — and skipped everything between them. That hunt is exactly what a named entity recogniser does, and the short name for it is NER.
Why it exists, and how it differs from what came before
Text classification puts a label on a whole message. This ticket is about billing. That review is positive.
NER goes inside the message. It labels parts of it.
text classification: "Ravi Sharma works at Infosys in Pune." -> employment news
NER: Ravi Sharma works at Infosys in Pune
^^^^^^^^^^^ ^^^^^^^ ^^^^
person company cityThat difference is the whole lesson. One gives you a folder. The other gives you fields you can put into a database.
The problem that makes it awkward
Names are often more than one word. "Ravi Sharma" is one person, written as two words.
So the model cannot look at each word alone and say person or not-person. It has to say something stronger: this word starts a name, or this word continues the name before it.
The standard way to write that down is called BIO tagging — three letters that stand for Beginning, Inside and Outside.
Ravi B-PER <- B means this word Begins a person's name
Sharma I-PER <- I means this word is Inside the same name
works O <- O means Outside, not part of any name
at O
Infosys B-ORG <- a new name begins, this one an organisation
in O
Pune B-LOC <- a location
. ORead that table twice. Once you can produce those tags, you can rebuild the names from them. That is the whole task.
Why the same word can be two different things
This is where NER stops being easy.
"I bought a phone from Apple." -> Apple is a company
"I bought an apple from the market." -> apple is fruit, not a name at all
"Washington signed the treaty." -> a person
"Washington gets cold in winter." -> a placeNothing about the word decides it. The surrounding words decide it. A system built from a list of known names cannot handle this, and lists were how people did it for years.
This is why models that read the whole sentence — the kind described in BERT — took over this task so completely.
Where you have already seen it
- Your phone reading an OTP. The message arrives and the code is offered as a tap-to-fill. Something found the number inside the text.
- Gmail putting a flight in your calendar. It pulled the airline, the date and the airport out of a confirmation email.
- Bank SMS becoming a spending chart. The amount, the merchant and the date were extracted from a line of text.
- Maps offering directions when you paste an address into a chat.
- Resume screening, pulling out names, companies, colleges and skills.
What is honestly hard
Indian names break many systems. Most widely used models were trained on English news, where names look like "John Smith". A name like "Venkata Subrahmanyam Iyer" or a place like "Thiruvananthapuram" is unfamiliar, often split into odd pieces, and frequently missed. This is a real gap, not a small one, and it is why India-specific datasets exist.
Capital letters do most of the work, and that is fragile. Models lean heavily on the capital letter at the start of a name. Feed them a message typed entirely in lowercase, as people type on phones, and performance falls apart. Devanagari, Tamil and most Indian scripts have no capital letters at all, so this crutch is not even available.
Scoring it is easy to get wrong. A model can tag most words correctly and still extract zero complete names. Getting "Ravi" right and "Sharma" wrong does not give you half a person. The developer section shows exactly this happening.
Remember this
- NER labels parts of a sentence, not the whole sentence.
- BIO tags solve the multi-word problem: B begins a name, I continues it, O is everything else.
- The surrounding words decide what a name means, so context-reading models win here.
What to learn next
- BERT — the context-reading encoder most entity models are built on.
- Text classification — labelling whole documents, for contrast.
- Structured output — getting clean fields out of a language model.
Developer — Code and libraries.
We will build a working entity tagger with scikit-learn and eight training sentences. It runs on CPU in under a second and it makes every important mistake, on purpose, where you can see it.
Setup
pip install scikit-learnNo downloads, no GPU, no framework.
A tagger you can read end to end
from sklearn.feature_extraction import DictVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
# Each sentence is a list of (word, tag). B- starts an entity, I- continues one,
# O means "not part of any entity".
TRAIN = [
[("Ravi","B-PER"),("Sharma","I-PER"),("works","O"),("at","O"),("Infosys","B-ORG"),("in","O"),("Pune","B-LOC"),(".","O")],
[("Priya","B-PER"),("Nair","I-PER"),("joined","O"),("Wipro","B-ORG"),("in","O"),("Bengaluru","B-LOC"),(".","O")],
[("Arjun","B-PER"),("Mehta","I-PER"),("left","O"),("Zomato","B-ORG"),("for","O"),("Delhi","B-LOC"),(".","O")],
[("Anita","B-PER"),("Rao","I-PER"),("moved","O"),("to","O"),("Chennai","B-LOC"),("last","O"),("year","O"),(".","O")],
[("She","O"),("met","O"),("Kiran","B-PER"),("Desai","I-PER"),("at","O"),("Flipkart","B-ORG"),(".","O")],
[("The","O"),("team","O"),("at","O"),("Paytm","B-ORG"),("is","O"),("in","O"),("Noida","B-LOC"),(".","O")],
[("Rohit","B-PER"),("Verma","I-PER"),("visited","O"),("Hyderabad","B-LOC"),("on","O"),("Monday","O"),(".","O")],
[("Meera","B-PER"),("Iyer","I-PER"),("reports","O"),("to","O"),("Swiggy","B-ORG"),(".","O")],
]
def features(words, i):
w = words[i]
return {
"word.lower": w.lower(),
"is.title": w[:1].isupper(),
"is.first": i == 0,
"suffix3": w[-3:].lower(),
"prev.lower": words[i - 1].lower() if i else "<START>",
"prev.title": words[i - 1][:1].isupper() if i else False,
"next.lower": words[i + 1].lower() if i < len(words) - 1 else "<END>",
}
X, y = [], []
for sent in TRAIN:
words = [w for w, _ in sent]
for i, (_, tag) in enumerate(sent):
X.append(features(words, i))
y.append(tag)
tagger = make_pipeline(DictVectorizer(), LogisticRegression(max_iter=2000)).fit(X, y)
def tag(words):
return list(tagger.predict([features(words, i) for i in range(len(words))]))
def spans(words, tags):
"""Turn a BIO tag sequence back into (text, type) entities."""
out, cur, typ = [], [], None
for w, t in zip(words, tags):
if t.startswith("B-"):
if cur: out.append((" ".join(cur), typ))
cur, typ = [w], t[2:]
elif t.startswith("I-") and cur and t[2:] == typ:
cur.append(w)
else:
if cur: out.append((" ".join(cur), typ))
cur, typ = [], None
if cur: out.append((" ".join(cur), typ))
return out
for sentence in ["Sunil Kapoor works at Infosys in Chennai .",
"Rahul Bose joined Ola in Kolkata .",
"apple sales grew in march ."]:
words = sentence.split()
tags = tag(words)
print(sentence)
print(" tags :", " ".join(tags))
print(" found:", spans(words, tags) or "nothing")
print()Sunil Kapoor works at Infosys in Chennai .
tags : B-PER I-PER O O B-ORG O B-LOC O
found: [('Sunil Kapoor', 'PER'), ('Infosys', 'ORG'), ('Chennai', 'LOC')]
Rahul Bose joined Ola in Kolkata .
tags : B-PER I-PER O B-ORG O B-LOC O
found: [('Rahul Bose', 'PER'), ('Ola', 'ORG'), ('Kolkata', 'LOC')]
apple sales grew in march .
tags : O O O O O O
found: nothingWhat actually happened there
Look at sentence two. Ola and Kolkata and Rahul Bose appear nowhere in the training data. Every one of them was tagged correctly.
The model did not memorise a list of names. It learned a shape: a capitalised word after "joined" is an organisation, a capitalised word after "in" is a location, two capitalised words at the start of a sentence are a person.
Sentence three is the control. Lowercase "apple" and "march" are left alone, which is right — they are not being used as names here.
Line by line, for the parts that carry the weight
prev.lower and next.lower. These two features are doing most of the work. Without them, the model sees an isolated capitalised word and has no way to guess whether it is a person, a company or a city. The neighbours are the evidence.
is.first. The first word of a sentence is capitalised for grammatical reasons, not because it is a name. Without this feature, "The team at Paytm" gets "The" tagged as a person. Try deleting it.
suffix3. Character endings generalise where whole words cannot. Indian surnames ending in kar, wal, jee, nan share endings across many unseen names, and the suffix feature transfers to names the model never saw.
spans() is defensive on purpose. This model predicts each word independently, so it can produce impossible sequences — I-PER with no B-PER before it, or B-PER followed by I-LOC. Real taggers add a layer that forbids these (a CRF, covered in the researcher section). Without one, your decoder must handle nonsense gracefully instead of crashing.
Now the mistake almost everyone makes
Here is why "accuracy" is close to useless for this task.
Append this to the end of ner.py and run it again. The three sentences above print first; this is what follows them.
gold = [("Sunil","B-PER"),("Kapoor","I-PER"),("works","O"),("at","O"),
("Infosys","B-ORG"),("in","O"),("Chennai","B-LOC"),(".","O")]
words = [w for w, _ in gold]
truth = [t for _, t in gold]
for label, ws in [("normal case ", words), ("all lowercase", [w.lower() for w in words])]:
tags = tag(ws)
token_acc = sum(a == b for a, b in zip(tags, truth)) / len(truth)
print(f"{label}: token accuracy {token_acc:.2f} entities found: {spans(ws, tags)}")normal case : token accuracy 1.00 entities found: [('Sunil Kapoor', 'PER'), ('Infosys', 'ORG'), ('Chennai', 'LOC')]
all lowercase: token accuracy 0.50 entities found: []That gap is the reason NER is scored at span level: an entity counts only when both its boundaries and its type are exactly right. Half a name is worth nothing to whatever database you were filling.
It also demonstrates how much weight capital letters carry. Real user text — WhatsApp messages, search queries, form fields — is often typed entirely in lowercase. If that is your input, train on lowercase text too.
Scoring it properly
pip install seqevalfrom seqeval.metrics import classification_report
y_true = [["B-PER","I-PER","O","O","B-ORG","O","B-LOC","O"]]
y_pred = [["B-PER","O", "O","O","B-ORG","O","B-LOC","O"]]
print(classification_report(y_true, y_pred))seqeval compares whole spans, not tokens. In this example the PER entity is scored as entirely wrong, because Sunil Kapoor and Sunil are different spans. No output block here — seqeval's report formatting has changed between releases, and a printed version would go stale.
Use seqeval, or the original CoNLL evaluation script, and never report token accuracy for an entity task.
Using a real model instead
For anything beyond learning, use a pretrained tagger. Two mainstream options:
pip install spacy && python -m spacy download en_core_web_sm # about 12 MBimport spacy
nlp = spacy.load("en_core_web_sm")
for ent in nlp("Sunil Kapoor works at Infosys in Chennai.").ents:
print(ent.text, ent.label_)Or a fine-tuned transformer through Hugging Face, using pipeline("ner", aggregation_strategy="simple"). The aggregation_strategy argument is what merges subword pieces back into whole entities, and forgetting it is the most common complaint about that API.
No output blocks for either. Both depend on the exact model version you download, and printed results here would drift out of date within months.
Common mistakes
Reporting token accuracy. Covered above. It is inflated by the O class and it hides total failure.
Mishandling subwords. A transformer tokenizer splits "Bengaluru" into "bengal" and "##uru". You must decide which piece carries the label — the convention is the first piece — and merge them back afterwards. Skipping this silently halves your entity count.
Training on cased text and serving lowercase text. The single most common production failure in NER. Match your training distribution to your input.
Ignoring invalid tag sequences. Independent per-token prediction produces O followed by I-LOC. Either add a CRF layer, or write a decoder that repairs these, and log how often it fires. A high repair rate means your model is confused, not that your decoder is clever.
Assuming entity types are obvious. Is "Infosys Pune office" one ORG, or an ORG plus a LOC? Is a college an ORG? Your annotators need a written rule for every such case, decided before labelling starts, or your labels will disagree and cap your score.
Try it yourself
Delete "is.first" from features() and re-run. Watch "The team at Paytm" start tagging "The" as a person, because a capitalised first word now looks exactly like a name.
Then add two training sentences containing a date — [("She","O"),("joined","O"),("on","O"),("12","B-DATE"),("March","I-DATE"),("2024","I-DATE"),(".","O")] — and test on a date the model has never seen. Predict first whether it will work. Digits are a strong shape signal, but your feature set has no digit feature yet. Add one, and measure the difference.
What to learn next
- BERT — the encoder that replaced hand-built features for this task.
- Structured output — extracting fields with a language model instead.
- Text classification — evaluation discipline that carries over here.
Researcher — Mathematics and papers.
The task, formally
NER is sequence labelling. Given a token sequence x = (x_1, ..., x_n), predict a tag sequence y = (y_1, ..., y_n) from a tag set derived from an entity-type inventory T and a chunking scheme.
The chunking scheme matters more than its obscurity suggests.
| Scheme | Tags per type | Notes |
|---|---|---|
| IOB1 | I-, B- | B- used only to split two adjacent entities of the same type. The original CoNLL format. |
| IOB2 / BIO | B-, I- | Every entity starts with B-. The modern default. |
| BIOES / BILOU | B-, I-, E-, S-, O | Explicit end and single-token tags. |
Ratinov & Roth (2009) reported that BILOU consistently outperforms BIO, by roughly 1 F1 on CoNLL-2003 with the same model. The explicit end marker gives the model a direct signal for boundary closure. This finding is old, reproducible and still routinely ignored.
Note also that the CoNLL-2003 files are IOB1. Loading them as IOB2 without conversion silently corrupts the boundaries of adjacent same-type entities. It is a recurring source of unreproducible published numbers.
Why a CRF layer
Independent per-token classification has no mechanism to forbid O -> I-LOC. A linear-chain conditional random field (Lafferty, McCallum & Pereira, 2001) models the whole sequence:
p(y | x) = (1 / Z(x)) * exp( sum over t = 1..n of [ s(x, t, y_t) + A[y_{t-1}, y_t] ] )s(x, t, y_t)is the emission score for tagy_tat positiont, produced by the underlying network.Ais a learned|Y| x |Y|transition matrix over the tag setY.Z(x)is the partition function, summing the exponentiated score over all|Y|^ntag sequences.
Z(x) is computed in O(n |Y|^2) by the forward algorithm, and decoding uses Viterbi at the same cost. Setting A[i, j] = -inf for invalid transitions makes illegal sequences impossible rather than only unlikely.
The empirical gain from a CRF layer is largest when the encoder is weak. On top of a strong pretrained encoder the gain shrinks to a fraction of a point, and many current systems drop it. Whether to include one is an engineering trade-off between a small quality gain and a decoding step that is harder to batch.
Architectures, in order of appearance
- Feature-engineered CRF (Finkel et al., 2005 — Stanford NER). Gazetteers, orthographic features, word shape, context windows. Competitive for years, and still fast enough for high-throughput pipelines.
- BiLSTM-CRF (Huang, Xu & Yu, 2015; Lample et al., 2016). Character-level BiLSTM for morphology plus word-level BiLSTM plus CRF. Lample et al. reached 90.94 F1 on CoNLL-2003 English without gazetteers, which was the result that ended feature engineering for this task.
- Contextual embeddings. ELMo (Peters et al., 2018) then BERT (Devlin et al., 2018) pushed CoNLL-2003 past 92 F1 with a linear head over token representations. See BERT.
- Span-based and biaffine. Yu, Bohnet & Poesio (2020) reframe NER as dependency parsing, scoring every candidate
(start, end)span with a biaffine classifier. This handles nested entities, which BIO tagging structurally cannot represent — "Bank of India" contains aLOCinside anORG, and a flat tag sequence has to choose one. - MRC framing. Li et al. (2020) cast each entity type as a question ("find all locations") answered by span extraction. Strong in low-resource settings because the type description carries information.
Evaluation
The CoNLL convention is exact-match span-level micro-F1. A predicted entity counts as correct only when its start index, end index and type all match a gold entity.
Consequences worth internalising:
- Token-level accuracy is not comparable to reported F1 and should never be reported for this task. In a typical corpus, 80 to 90 percent of tokens are
O, so a model predictingOeverywhere already exceeds 0.8 token accuracy. - Partial matches score zero. A one-token boundary error costs one false positive and one false negative.
- Micro-averaging over entities means frequent types dominate. Report per-type F1 as well;
MISCand rare types are usually far worse than the headline suggests.
MUC-5 style scoring, which gives partial credit for boundary-correct type-wrong matches, exists and is occasionally more informative for diagnosis. It is not comparable to CoNLL numbers.
Use seqeval (Nakayama, 2018) or the original conlleval script. Hand-rolled span scorers get IOB1 and adjacent-entity edge cases wrong with high probability.
Datasets, and their limits
- CoNLL-2003 — Reuters news, English and German, four types. Saturated, small, and heavily contaminated in LLM pretraining corpora. Any zero-shot result on it should be treated with suspicion.
- OntoNotes 5.0 — 18 types across newswire, broadcast, web and telephone conversation. Multi-genre, which makes it far more informative about domain robustness.
- WikiANN / PAN-X (Pan et al., 2017) — silver-standard NER for 282 languages from Wikipedia links. Broad coverage, noisy labels.
- Naamapadam (Mhaske et al., 2022, arXiv:2212.10168) — the largest publicly available Indic NER corpus, covering 11 Indian languages, built by projecting English labels across a parallel corpus. This is the right starting point for Indian-language entity work, with the caveat that projected labels carry alignment errors.
Domain shift between these datasets is severe. A model at 92 F1 on CoNLL news commonly lands in the 60s on user-generated text, and the drop comes disproportionately from casing, tokenisation and unfamiliar name morphology.
LLMs and zero-shot NER
Large decoder models underperform fine-tuned encoders on NER by a wide margin. Wang et al. (2023), GPT-NER (arXiv:2304.10428), document the gap and the reasons: LLMs are trained to generate, not to align outputs to exact input offsets, and they hallucinate boundaries that do not exist in the source.
The interesting counterexample is GLiNER (Zaratiana et al., 2023, arXiv:2311.08526). It uses a bidirectional encoder to match arbitrary entity-type descriptions against candidate spans, giving genuine zero-shot NER at encoder cost and encoder latency. It outperforms prompted LLMs on zero-shot benchmarks while being orders of magnitude cheaper.
The practical decision rule: if the entity types are fixed and you can label two thousand sentences, fine-tune an encoder. If the types change per request, use GLiNER or a similar span-matching model before reaching for a generative one.
Key references
- Lafferty, J., McCallum, A. & Pereira, F. (2001). Conditional Random Fields. ICML.
- Finkel, J., Grenager, T. & Manning, C. (2005). Incorporating Non-local Information into Information Extraction Systems by Gibbs Sampling. ACL.
- Ratinov, L. & Roth, D. (2009). Design Challenges and Misconceptions in Named Entity Recognition. CoNLL.
- Huang, Z., Xu, W. & Yu, K. (2015). Bidirectional LSTM-CRF Models for Sequence Tagging. arXiv:1508.01991
- Lample, G. et al. (2016). Neural Architectures for Named Entity Recognition. arXiv:1603.01360
- Pan, X. et al. (2017). Cross-lingual Name Tagging and Linking for 282 Languages. ACL.
- Li, X. et al. (2020). A Unified MRC Framework for Named Entity Recognition. arXiv:1910.11476
- Yu, J., Bohnet, B. & Poesio, M. (2020). Named Entity Recognition as Dependency Parsing. arXiv:2005.07150
- Mhaske, A. et al. (2022). Naamapadam: A Large-Scale Named Entity Annotated Data for Indic Languages. arXiv:2212.10168
- Zaratiana, U. et al. (2023). GLiNER: Generalist Model for NER using Bidirectional Transformer. arXiv:2311.08526
Open problems
Nested and discontinuous entities. BIO cannot represent either. Span-based models handle nesting; discontinuous entities — "left and right ventricles" denoting two entities sharing a head — remain awkward and are common in clinical text.
Recognition is not linking. Tagging "Apple" as an ORG does not tell you which organisation. Entity linking against a knowledge base is a separate, harder problem, and most downstream applications actually need it.
Low-resource and code-mixed text. Performance on Indian languages trails English substantially, and the cause is layered: tokenizer fertility, projected rather than human labels, and the absence of case as a signal. Progress here is limited by annotation budgets rather than by architecture, which is an unglamorous but accurate diagnosis.
What to learn next
- BERT — the encoder backbone and its token-classification head.
- Text classification — label quality and evaluation discipline.
- Structured output — constrained generation as an alternative extraction route.