Sequence Labelling and Structure
Entity linking
Entity linking connects a name in text to one specific real-world entity in a knowledge base, resolving the fact that many different things share the same name.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Entity linking takes a name found in text and points it at one specific real-world thing, choosing between all the different things that share that name.
Think about calling out "Kumar!" in a crowded classroom. Several students might share that surname. The teacher does more than notice that a name was said. They use context: who was addressed a moment earlier, who is sitting where. That is how they work out exactly which Kumar is meant.
Entity linking does the same job for text. Recognising that "Chennai" is a place name is one step. Working out that this mention of "Chennai" means the city, not the cricket team named after it, is a separate, harder step. That is entity linking.
Why it exists
Named entity recognition, covered earlier in this section, only tells you that a span of text is a LOCATION or an ORG. It does not tell you which location, or which organisation. "Chennai" could mean the city, the cricket franchise "Chennai Super Kings", or the railway station "Chennai Egmore". Three real entities share part of the name.
Two systems can both correctly tag "Chennai" as a LOCATION, and still disagree completely about what it refers to. Some tasks need to look something up — pulling facts, merging duplicate mentions, building a knowledge graph. For those, recognising the entity type is not enough. You need to know exactly which entity it is.
How it works
Entity linking has two stages: find every plausible candidate, then pick the one that actually fits the context.
Mention: "Chennai" in "Chennai won the IPL final after a tense chase."
Stage 1 -- candidate generation:
Chennai (the city)
Chennai Super Kings (the cricket team)
Chennai Egmore (the railway station)
Stage 2 -- disambiguation, using the surrounding words:
"won", "IPL", "final", "chase" all strongly suggest cricket
-> Chennai Super Kings [chosen]The mention text alone is ambiguous. The words around it are what break the tie.
Where you have already seen it
- Wikipedia's blue links. Clicking a linked name inside an article takes you to the specific page for that entity. Wikipedia's own internal linking system solves entity linking at massive scale.
- Search engines' knowledge panels. Typing "Chennai" and seeing a panel about the city, not the cricket team, means the search engine already resolved which entity you probably meant.
- News aggregators. Every article mentioning "the RBI" gets grouped together, even across different phrasings, by linking each mention to one canonical organisation.
Remember this
- Entity linking connects a text mention to one specific entity in a knowledge base, not only to a category like
LOCATION. - The mention text alone is often ambiguous — several different real entities can share a name.
- Disambiguation uses the surrounding context to pick the right one.
What to learn next
- Coreference resolution — connecting later pronouns and references back to an entity once it is linked.
- Named entity recognition — the step that finds the mention entity linking then resolves.
- Semantic search — a related use of context-based matching, applied to whole queries instead of single mentions.
Developer — Code and libraries.
A minimal entity linker is two functions: rank candidates by name similarity, then rerank by context overlap. No model download required to see the core idea.
Setup
Nothing to install. Standard library only — difflib ships with Python.
Minimal runnable code
from difflib import SequenceMatcher
# A tiny knowledge base: entity id -> (canonical name, short description)
kb = {
"Q1156": ("Chennai", "capital city of Tamil Nadu state India"),
"Q1361": ("Chennai Super Kings", "cricket team IPL franchise trophy"),
"Q15180": ("Chennai Egmore", "railway station train platform"),
}
def name_score(mention, name):
return SequenceMatcher(None, mention.lower(), name.lower()).ratio()
def link(mention, context):
# Step 1: candidate generation. Rank every KB entry by how close its
# canonical name is to the mention string alone, with no context yet.
candidates = sorted(kb.items(), key=lambda kv: name_score(mention, kv[1][0]), reverse=True)[:3]
# Step 2: disambiguation. Rerank those candidates by words the surrounding
# sentence shares with each candidate's description. This is a toy stand-in
# for what a real linker's context encoder does with dense vectors.
context_words = set(context.lower().split()) - {mention.lower()}
def context_overlap(entry):
_, (name, desc) = entry
return len(context_words & set(desc.lower().split()))
return max(candidates, key=context_overlap)
mentions = [
("Chennai", "The train departs from Chennai railway station at 6 AM."),
("Chennai", "Chennai lifted the IPL trophy after a tense final."),
]
for mention, context in mentions:
qid, (name, desc) = link(mention, context)
print(f"{mention!r:10} in {context[:40]!r}...")
print(f" -> {qid} {name} ({desc})")'Chennai' in 'The train departs from Chennai railway s'... -> Q15180 Chennai Egmore (railway station train platform) 'Chennai' in 'Chennai lifted the IPL trophy after a te'... -> Q1361 Chennai Super Kings (cricket team IPL franchise trophy)
The exact same mention string, "Chennai", resolves to two different knowledge base entries, purely because the surrounding words changed.
Line by line
name_score uses difflib.SequenceMatcher, a general-purpose string similarity tool, not anything specific to entity linking. It measures how similar two strings look, character by character, which is enough for candidate generation when mentions are close variants of the canonical name.
Candidate generation happens before disambiguation, and deliberately ignores context. This mirrors real systems: generate a short list cheaply and broadly first, then spend the more expensive context-matching step only on that short list, not the entire knowledge base.
context_words - {mention.lower()} removes the mention word itself from the comparison. Without this, every candidate description that happens to repeat the word "Chennai" gets spurious credit for containing the mention alone, which drowns out the actual disambiguating words like "railway" or "trophy".
A real linker replaces both string-matching steps with learned embeddings: a dense vector for the mention in context, compared against dense vectors for each candidate entity, ranked by similarity rather than word overlap. The two-stage shape — generate candidates cheaply, then rerank with more context — stays the same.
Common mistakes
Forgetting to exclude the mention word from the context-overlap check, as shown above — every candidate description containing the entity's own name gets an unfair boost, and disambiguation collapses toward "whichever description happens to repeat the mention word most."
Using name similarity alone, with no disambiguation step. "Chennai" and "Chennai Super Kings" are close string matches to each other by construction — string similarity alone cannot tell them apart without context, since the ambiguity is inherent to the names themselves.
Assuming every mention has a correct entity in the knowledge base. Real text mentions people, places and organisations that a fixed knowledge base has never heard of. A production linker needs a deliberate "no match" option (often called NIL prediction) rather than being forced to pick the closest available candidate every time.
Try it yourself
Add a fourth knowledge base entry for a different "Chennai"-named entity — perhaps a restaurant chain — with its own description, and write a third context sentence that should resolve to it. Check whether the two-stage approach still picks correctly as the candidate pool grows.
What to learn next
- Coreference resolution — resolving pronouns back to an already-linked entity.
- Multilingual sentence embeddings — the kind of learned representation a real linker uses in place of string matching.
- Bi-encoders vs cross-encoders — the two-stage retrieve-then-rerank pattern this lesson's code mirrors.
Researcher — Mathematics and papers.
The task, formally
Given a mention m with surrounding context c, and a knowledge base KB = {e_1, ..., e_K} where each e_k has a canonical name and a textual or structured description, entity linking predicts:
e* = argmax over e in C(m) of score(m, c, e)C(m) subset of KBis the candidate set for mentionm, produced by candidate generation — restricting the search space is necessary sinceKcan be in the millions (Wikidata has over 100 million entities).score(m, c, e)is a compatibility score between the mention-in-context and a candidate entity, learned or heuristic.
Many formulations add an explicit NIL option, e* = NIL when no candidate scores above a threshold, representing "this mention refers to an entity not present in the knowledge base."
Candidate generation methods
Alias tables, built from redirect links, disambiguation pages and anchor text statistics on Wikipedia, mapping surface strings to a ranked list of entities by prior mention frequency (Hoffart et al., 2011). Fast, high-recall, but blind to context.
Dense retrieval, encoding the mention-in-context and every KB entity into the same embedding space and retrieving nearest neighbours (Wu et al., 2020, BLINK). Scales to large knowledge bases via approximate nearest-neighbour search — see FAISS — and captures semantic rather than purely lexical similarity, at higher compute cost than an alias table lookup.
Disambiguation methods
Local models score each mention against each candidate independently, typically as a cross-encoder over [mention + context, candidate description] pairs, giving strong per-mention accuracy at the cost of being unaware of other entities mentioned in the same document.
Global models additionally exploit coherence between entities in the same document — if "Chennai" and "IPL final" both appear, and "IPL final" is confidently linked to the cricket tournament, that raises the probability "Chennai" refers to the cricket team, not the city. Formulated as a joint optimisation over all mentions in a document simultaneously (Hoffart et al., 2011; Ganea & Hofmann, 2017, using a graph neural network over an entity-mention graph), global models consistently outperform local ones on standard benchmarks (AIDA-CoNLL) by several F1 points, at higher inference cost since mentions can no longer be resolved independently.
End-to-end and generative approaches
GENRE (De Cao et al., 2021) reframes entity linking as sequence-to-sequence generation: the model directly generates the entity's unique name as text, using constrained decoding restricted to valid entity titles via a prefix trie. This sidesteps candidate generation and dense retrieval entirely, replacing a retrieve-then-rank pipeline with a single autoregressive model, at the cost of a much larger decoding search space than a fixed candidate list.
Key references
- Hoffart, J. et al. (2011). Robust Disambiguation of Named Entities in Text. EMNLP. Introduces the AIDA-CoNLL benchmark and a global coherence-based linker.
- Ganea, O. & Hofmann, T. (2017). Deep Joint Entity Disambiguation with Local Neural Attention. arXiv:1704.04920
- Wu, L. et al. (2020). Scalable Zero-shot Entity Linking with Dense Entity Retrieval. arXiv:1911.03814 — BLINK.
- De Cao, N. et al. (2021). Autoregressive Entity Retrieval. arXiv:2010.00904 — GENRE.
Current state and open problems
Dense-retrieval linkers (BLINK and its successors) are the current standard for Wikipedia-scale linking, largely displacing alias-table-only approaches by handling mentions with no exact string overlap to the canonical name. Zero-shot entity linking — resolving mentions to entities never seen during training, common in specialised domains like biomedicine or enterprise knowledge bases — remains meaningfully harder than the well-studied Wikipedia setting, since it removes the large mention-frequency priors that alias tables and much of the training data depend on. Emerging mentions — real-world entities that postdate a model's training data or knowledge base snapshot — are a related, unresolved gap: a linker cannot resolve a mention to an entity that does not yet exist in its knowledge base, regardless of how good its disambiguation is.
What to learn next
- Bi-encoders vs cross-encoders — the architectural split behind candidate generation versus reranking.
- Coreference resolution — resolving mentions of an already-linked entity across a document.
- Relation extraction — building structured facts once entities are correctly linked.