Sequence Labelling and Structure

Entity linking

Entity linking connects a name in text to one specific real-world entity in a knowledge base, resolving the fact that many different things share the same name.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Entity linking takes a name found in text and points it at one specific real-world thing, choosing between all the different things that share that name.

Think about calling out "Kumar!" in a crowded classroom. Several students might share that surname. The teacher does more than notice that a name was said. They use context: who was addressed a moment earlier, who is sitting where. That is how they work out exactly which Kumar is meant.

Entity linking does the same job for text. Recognising that "Chennai" is a place name is one step. Working out that this mention of "Chennai" means the city, not the cricket team named after it, is a separate, harder step. That is entity linking.

Why it exists

Named entity recognition, covered earlier in this section, only tells you that a span of text is a LOCATION or an ORG. It does not tell you which location, or which organisation. "Chennai" could mean the city, the cricket franchise "Chennai Super Kings", or the railway station "Chennai Egmore". Three real entities share part of the name.

Two systems can both correctly tag "Chennai" as a LOCATION, and still disagree completely about what it refers to. Some tasks need to look something up — pulling facts, merging duplicate mentions, building a knowledge graph. For those, recognising the entity type is not enough. You need to know exactly which entity it is.

How it works

Entity linking has two stages: find every plausible candidate, then pick the one that actually fits the context.

  Mention:  "Chennai"  in  "Chennai won the IPL final after a tense chase."

  Stage 1 -- candidate generation:
    Chennai (the city)
    Chennai Super Kings (the cricket team)
    Chennai Egmore (the railway station)

  Stage 2 -- disambiguation, using the surrounding words:
    "won", "IPL", "final", "chase" all strongly suggest cricket

    -> Chennai Super Kings   [chosen]

The mention text alone is ambiguous. The words around it are what break the tie.

Where you have already seen it

  • Wikipedia's blue links. Clicking a linked name inside an article takes you to the specific page for that entity. Wikipedia's own internal linking system solves entity linking at massive scale.
  • Search engines' knowledge panels. Typing "Chennai" and seeing a panel about the city, not the cricket team, means the search engine already resolved which entity you probably meant.
  • News aggregators. Every article mentioning "the RBI" gets grouped together, even across different phrasings, by linking each mention to one canonical organisation.

Remember this

  • Entity linking connects a text mention to one specific entity in a knowledge base, not only to a category like LOCATION.
  • The mention text alone is often ambiguous — several different real entities can share a name.
  • Disambiguation uses the surrounding context to pick the right one.

What to learn next

  • Coreference resolution — connecting later pronouns and references back to an entity once it is linked.
  • Named entity recognition — the step that finds the mention entity linking then resolves.
  • Semantic search — a related use of context-based matching, applied to whole queries instead of single mentions.

Developer — Code and libraries.

A minimal entity linker is two functions: rank candidates by name similarity, then rerank by context overlap. No model download required to see the core idea.

Setup

Nothing to install. Standard library only — difflib ships with Python.

Minimal runnable code

entity_link.py
from difflib import SequenceMatcher

# A tiny knowledge base: entity id -> (canonical name, short description)
kb = {
    "Q1156":  ("Chennai",             "capital city of Tamil Nadu state India"),
    "Q1361":  ("Chennai Super Kings", "cricket team IPL franchise trophy"),
    "Q15180": ("Chennai Egmore",      "railway station train platform"),
}

def name_score(mention, name):
    return SequenceMatcher(None, mention.lower(), name.lower()).ratio()

def link(mention, context):
    # Step 1: candidate generation. Rank every KB entry by how close its
    # canonical name is to the mention string alone, with no context yet.
    candidates = sorted(kb.items(), key=lambda kv: name_score(mention, kv[1][0]), reverse=True)[:3]

    # Step 2: disambiguation. Rerank those candidates by words the surrounding
    # sentence shares with each candidate's description. This is a toy stand-in
    # for what a real linker's context encoder does with dense vectors.
    context_words = set(context.lower().split()) - {mention.lower()}
    def context_overlap(entry):
        _, (name, desc) = entry
        return len(context_words & set(desc.lower().split()))
    return max(candidates, key=context_overlap)

mentions = [
    ("Chennai", "The train departs from Chennai railway station at 6 AM."),
    ("Chennai", "Chennai lifted the IPL trophy after a tense final."),
]
for mention, context in mentions:
    qid, (name, desc) = link(mention, context)
    print(f"{mention!r:10} in {context[:40]!r}...")
    print(f"  -> {qid}  {name}  ({desc})")
Output
'Chennai'  in 'The train departs from Chennai railway s'...
  -> Q15180  Chennai Egmore  (railway station train platform)
'Chennai'  in 'Chennai lifted the IPL trophy after a te'...
  -> Q1361  Chennai Super Kings  (cricket team IPL franchise trophy)

The exact same mention string, "Chennai", resolves to two different knowledge base entries, purely because the surrounding words changed.

Line by line

name_score uses difflib.SequenceMatcher, a general-purpose string similarity tool, not anything specific to entity linking. It measures how similar two strings look, character by character, which is enough for candidate generation when mentions are close variants of the canonical name.

Candidate generation happens before disambiguation, and deliberately ignores context. This mirrors real systems: generate a short list cheaply and broadly first, then spend the more expensive context-matching step only on that short list, not the entire knowledge base.

context_words - {mention.lower()} removes the mention word itself from the comparison. Without this, every candidate description that happens to repeat the word "Chennai" gets spurious credit for containing the mention alone, which drowns out the actual disambiguating words like "railway" or "trophy".

A real linker replaces both string-matching steps with learned embeddings: a dense vector for the mention in context, compared against dense vectors for each candidate entity, ranked by similarity rather than word overlap. The two-stage shape — generate candidates cheaply, then rerank with more context — stays the same.

Common mistakes

Forgetting to exclude the mention word from the context-overlap check, as shown above — every candidate description containing the entity's own name gets an unfair boost, and disambiguation collapses toward "whichever description happens to repeat the mention word most."

Using name similarity alone, with no disambiguation step. "Chennai" and "Chennai Super Kings" are close string matches to each other by construction — string similarity alone cannot tell them apart without context, since the ambiguity is inherent to the names themselves.

Assuming every mention has a correct entity in the knowledge base. Real text mentions people, places and organisations that a fixed knowledge base has never heard of. A production linker needs a deliberate "no match" option (often called NIL prediction) rather than being forced to pick the closest available candidate every time.

Try it yourself

Add a fourth knowledge base entry for a different "Chennai"-named entity — perhaps a restaurant chain — with its own description, and write a third context sentence that should resolve to it. Check whether the two-stage approach still picks correctly as the candidate pool grows.

What to learn next

Researcher — Mathematics and papers.

The task, formally

Given a mention m with surrounding context c, and a knowledge base KB = {e_1, ..., e_K} where each e_k has a canonical name and a textual or structured description, entity linking predicts:

text
e* = argmax over e in C(m) of score(m, c, e)
  • C(m) subset of KB is the candidate set for mention m, produced by candidate generation — restricting the search space is necessary since K can be in the millions (Wikidata has over 100 million entities).
  • score(m, c, e) is a compatibility score between the mention-in-context and a candidate entity, learned or heuristic.

Many formulations add an explicit NIL option, e* = NIL when no candidate scores above a threshold, representing "this mention refers to an entity not present in the knowledge base."

Candidate generation methods

Alias tables, built from redirect links, disambiguation pages and anchor text statistics on Wikipedia, mapping surface strings to a ranked list of entities by prior mention frequency (Hoffart et al., 2011). Fast, high-recall, but blind to context.

Dense retrieval, encoding the mention-in-context and every KB entity into the same embedding space and retrieving nearest neighbours (Wu et al., 2020, BLINK). Scales to large knowledge bases via approximate nearest-neighbour search — see FAISS — and captures semantic rather than purely lexical similarity, at higher compute cost than an alias table lookup.

Disambiguation methods

Local models score each mention against each candidate independently, typically as a cross-encoder over [mention + context, candidate description] pairs, giving strong per-mention accuracy at the cost of being unaware of other entities mentioned in the same document.

Global models additionally exploit coherence between entities in the same document — if "Chennai" and "IPL final" both appear, and "IPL final" is confidently linked to the cricket tournament, that raises the probability "Chennai" refers to the cricket team, not the city. Formulated as a joint optimisation over all mentions in a document simultaneously (Hoffart et al., 2011; Ganea & Hofmann, 2017, using a graph neural network over an entity-mention graph), global models consistently outperform local ones on standard benchmarks (AIDA-CoNLL) by several F1 points, at higher inference cost since mentions can no longer be resolved independently.

End-to-end and generative approaches

GENRE (De Cao et al., 2021) reframes entity linking as sequence-to-sequence generation: the model directly generates the entity's unique name as text, using constrained decoding restricted to valid entity titles via a prefix trie. This sidesteps candidate generation and dense retrieval entirely, replacing a retrieve-then-rank pipeline with a single autoregressive model, at the cost of a much larger decoding search space than a fixed candidate list.

Key references

  • Hoffart, J. et al. (2011). Robust Disambiguation of Named Entities in Text. EMNLP. Introduces the AIDA-CoNLL benchmark and a global coherence-based linker.
  • Ganea, O. & Hofmann, T. (2017). Deep Joint Entity Disambiguation with Local Neural Attention. arXiv:1704.04920
  • Wu, L. et al. (2020). Scalable Zero-shot Entity Linking with Dense Entity Retrieval. arXiv:1911.03814 — BLINK.
  • De Cao, N. et al. (2021). Autoregressive Entity Retrieval. arXiv:2010.00904 — GENRE.

Current state and open problems

Dense-retrieval linkers (BLINK and its successors) are the current standard for Wikipedia-scale linking, largely displacing alias-table-only approaches by handling mentions with no exact string overlap to the canonical name. Zero-shot entity linking — resolving mentions to entities never seen during training, common in specialised domains like biomedicine or enterprise knowledge bases — remains meaningfully harder than the well-studied Wikipedia setting, since it removes the large mention-frequency priors that alias tables and much of the training data depend on. Emerging mentions — real-world entities that postdate a model's training data or knowledge base snapshot — are a related, unresolved gap: a linker cannot resolve a mention to an entity that does not yet exist in its knowledge base, regardless of how good its disambiguation is.

What to learn next

What to learn next

These follow on from what you just read.

  • Sequence Labelling and Structure

    Coreference resolution

    Coreference resolution figures out which earlier name a pronoun like "she" or "it" refers back to, without which a machine reads a whole paragraph as disconnected sentences about strangers.

  • Sequence Labelling and Structure

    Relation extraction

    Relation extraction turns two entities and a sentence into a structured fact, such as (Zoho, acquired, a startup), which is what lets unstructured text feed a database or a knowledge graph.

  • Sequence Labelling and Structure

    Dependency parsing

    Dependency parsing draws a tree connecting every word to the word it grammatically depends on, and it is the structure relation extraction, coreference resolution and grammar checkers are all quietly built on.