Sequence Labelling and Structure
Coreference resolution
Coreference resolution figures out which earlier name a pronoun like "she" or "it" refers back to, without which a machine reads a whole paragraph as disconnected sentences about strangers.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Coreference resolution works out which earlier name a word like "she", "he" or "it" is standing in for.
Think about listening to someone tell a story about their day. "I met Priya at the station. She was late, so we grabbed chai." You instantly know "she" means Priya. Nobody spells it out for you. You connect the pronoun to the name automatically, using memory of who was mentioned a moment earlier.
A machine reading text word by word does not have that memory built in. Without coreference resolution, "she" is only a pronoun floating with no owner. "Priya" and "she" look like two unconnected facts about two different, unrelated people.
Why it exists
Real writing is full of pronouns. Repeating a full name in every sentence sounds unnatural. "Priya went to the station. Priya waited for the train. Priya was cold." reads like a robot wrote it. "Priya went to the station. She waited for the train. She was cold." reads normally. Every "she" quietly depends on the reader remembering who was mentioned first.
Any system that needs to actually understand a document has to resolve these connections. This means more than reading individual sentences. Otherwise a summariser might report "an unnamed woman waited for a train". It would completely miss that it is the same Priya mentioned in the sentence before.
How it works
A coreference resolver scans a document. It finds every mention that could refer to an entity — names, pronouns, descriptive phrases. Mentions of the same real entity get grouped into one cluster.
"Priya met Arjun at the station. She told him the train was late."
Mentions found: Priya, Arjun, She, him, the train
Clustering by who refers to whom:
Cluster 1: { Priya, She }
Cluster 2: { Arjun, him }
Cluster 3: { the train } <- only mentioned once, its own clusterDeciding "She" belongs with "Priya" and not "Arjun" uses several clues at once. Grammatical gender rules out Arjun. Recency matters too, along with which role each name played in the sentence before.
Where you have already seen it
- Voice assistants handling follow-up questions. "What's the weather in Chennai?" then "Is it going to rain there tomorrow?" Resolving "it" and "there" back to Chennai is coreference resolution, running behind the scenes.
- Document summarisers. They need to know every sentence mentioning "she" or "the company" is talking about the same entity. Otherwise the summary fragments one story into several unrelated-looking facts.
- Legal and medical document review. Correctly tracking which pronoun refers to which party, or which patient, across a long document has real consequences if it goes wrong.
Remember this
- Coreference resolution groups every mention of the same real-world entity — names, pronouns, descriptions — into one cluster.
- It is what lets a machine understand a pronoun refers to something mentioned earlier, instead of treating it as an unconnected word.
- Real resolvers combine several clues at once: grammatical agreement, recency, and sentence role.
What to learn next
- Relation extraction — building facts that connect entities, which coreference resolution feeds cleaner input into.
- Entity linking — resolving what a name refers to, a related but different problem from what a pronoun refers to.
- Multi-hop questions — a task that regularly needs coreference resolved correctly across sentences.
Developer — Code and libraries.
A simple rule-based resolver — matching pronouns to the nearest name with agreeing gender — needs nothing more than spaCy's part-of-speech tagging. It is a genuine, working baseline, not a toy with no relation to real systems: this was close to the state of the art before neural coreference models arrived.
Setup
pip install spacy
python -m spacy download en_core_web_smMinimal runnable code
import spacy
nlp = spacy.load("en_core_web_sm")
MALE_PRONOUNS = {"he", "him", "his"}
FEMALE_PRONOUNS = {"she", "her", "hers"}
MALE_NAMES = {"Arjun", "Ravi"}
FEMALE_NAMES = {"Priya", "Meera"}
def resolve_pronouns(text):
doc = nlp(text)
candidates = [t for t in doc if t.text in MALE_NAMES or t.text in FEMALE_NAMES]
for token in doc:
low = token.text.lower()
if low not in MALE_PRONOUNS and low not in FEMALE_PRONOUNS:
continue
wanted_gender = "male" if low in MALE_PRONOUNS else "female"
# Walk through every earlier name and keep the closest one whose
# gender agrees -- "closest" is what makes this a real, if simple,
# resolution strategy rather than a random guess.
best = None
for cand in candidates:
if cand.i >= token.i:
continue
cand_gender = "male" if cand.text in MALE_NAMES else "female"
if cand_gender == wanted_gender:
best = cand
if best is not None:
print(f"{token.text!r:6} at position {token.i:2} -> refers to {best.text!r} at position {best.i}")
resolve_pronouns("Priya met Arjun at the station. She told him the train was late.")'She' at position 7 -> refers to 'Priya' at position 0 'him' at position 9 -> refers to 'Arjun' at position 2
Both pronouns resolve correctly, using nothing more than "closest name with matching gender".
Line by line
candidates is collected once, then filtered per pronoun. Only names appearing before the pronoun (cand.i < token.i) are eligible — a pronoun cannot refer to something not yet mentioned in this simple forward-reading model, which matches how coreference almost always works in practice.
Gender agreement is the whole disambiguation mechanism here. "She" can only match a name in FEMALE_NAMES, which is precisely why "him" correctly skips Priya and lands on Arjun even though Priya is mentioned more recently in absolute position.
The loop keeps overwriting best with each later matching candidate, so it naturally lands on the closest eligible name by the time the loop finishes — a cheap way to implement a recency preference without explicit distance math.
Where this breaks
import spacy
nlp = spacy.load("en_core_web_sm")
FEMALE_PRONOUNS = {"she", "her", "hers"}
FEMALE_NAMES = {"Priya", "Meera"}
def resolve_pronouns(text):
doc = nlp(text)
candidates = [t for t in doc if t.text in FEMALE_NAMES]
for token in doc:
if token.text.lower() not in FEMALE_PRONOUNS:
continue
best = None
for cand in candidates:
if cand.i < token.i:
best = cand
if best is not None:
print(f"{token.text!r:6} at position {token.i:2} -> refers to {best.text!r} at position {best.i}")
resolve_pronouns("Priya met Meera at the station. She told her the train was late.")'She' at position 7 -> refers to 'Meera' at position 2 'her' at position 9 -> refers to 'Meera' at position 2
Both pronouns resolve to "Meera" — the closest matching name — even though a human reader cannot actually tell from this sentence alone who told whom. This is a genuine, unresolved ambiguity in the sentence itself, not a bug in the code. This exact failure mode — two entities sharing the same gender, so agreement alone cannot disambiguate — is precisely why real coreference systems need more than agreement rules, and why even state-of-the-art neural resolvers still get sentences like this wrong sometimes.
Common mistakes
Assuming gender agreement is enough for real text. As shown above, it collapses the moment two candidates share a gender. Real resolvers add syntactic clues (subject position is preferred over object position for the next pronoun, following Hobbs' 1978 naive algorithm) and, in neural systems, learned patterns from data.
Ignoring "it". Production coreference has to handle "it" referring to organisations, objects, and entire previous clauses ("The meeting ran long, which annoyed everyone" — "which" refers to the whole preceding clause). This toy example only handles gendered pronouns to keep the code short.
Treating coreference as solved once entities are correctly recognised. Named entity recognition finds "Priya" and "Arjun" as separate mentions. It says nothing about which later pronouns belong to which — that is a genuinely separate task layered on top.
Try it yourself
Add "it" and "its" as neutral pronouns that can match any name in a new ORG_NAMES set, and test on "Zoho released an update. It fixed several bugs." Notice how much weaker the disambiguation signal is without gender agreement to lean on.
What to learn next
- Relation extraction — using resolved entities to build structured facts.
- Entity linking — the related task of resolving what a name refers to, rather than what a pronoun refers to.
- Handling 'what about last year?' — coreference resolution applied to multi-turn conversation instead of a single document.
Researcher — Mathematics and papers.
The task, formally
Given a document with a set of mentions M = {m_1, ..., m_n} (spans that could refer to an entity — names, pronouns, definite noun phrases), coreference resolution partitions M into clusters {C_1, ..., C_k} such that all mentions within a cluster refer to the same real-world entity.
This is typically decomposed into mention detection (finding M) and antecedent linking (for each mention, either linking it to the most recent coreferent mention, or starting a new cluster). Clusters are recovered by following the resulting chain of links.
From rules to end-to-end neural models
Hobbs' naive algorithm (1978) performs a syntactic tree search, preferring subject position over object position, and earlier sentences over later ones — the underlying intuition the developer-block heuristic borrows in simplified form. Surprisingly strong as a baseline for decades, given its complete lack of learned parameters.
Mention-pair and mention-ranking models (Soon et al., 2001; Denis & Baldridge, 2008) frame antecedent selection as binary classification or ranking over candidate (mention, antecedent) pairs, using hand-built features similar in spirit to the CRF features in Conditional random fields.
End-to-end neural coreference (Lee et al., 2017) removes the separate mention-detection step entirely. It scores every span up to a maximum width as a potential mention, and every (mention, antecedent) pair jointly, trained end-to-end to maximise the marginal likelihood of the correct clustering — since a mention can have multiple valid antecedents within a cluster and any one of them counts as correct.
score(i, j) = s_m(i) + s_m(j) + s_a(i, j)s_m(i)is a learned score for spanibeing a valid mention at all.s_a(i, j)is a learned pairwise antecedent score between mentioniand candidate antecedentj.- Spans below a mention-score threshold are pruned before the
O(n^2)pairwise scoring step, keeping the method tractable for document-length input.
Transformer-based systems (Joshi et al., 2019, SpanBERT; Wu et al., 2020, CorefQA reframing coreference as span-based question answering) replace the encoder with a pretrained transformer, improving mention representations substantially, particularly for longer, harder antecedents.
Complexity
Naive mention-pair scoring over all spans is O(n^4) for a document of n tokens (span boundaries squared, then pairs of spans squared again). Lee et al.'s pruning — keeping only the top-scoring O(n) spans by mention score before pairwise scoring — reduces this to a practically tractable O(n^2) in the number of retained mentions, which is what made end-to-end neural coreference feasible on document-length text.
Key references
- Hobbs, J. (1978). Resolving Pronoun References. Lingua 44(4).
- Soon, W., Ng, H. & Lim, D. (2001). A Machine Learning Approach to Coreference Resolution of Noun Phrases. Computational Linguistics 27(4).
- Lee, K., He, L., Lewis, M. & Zettlemoyer, L. (2017). End-to-end Neural Coreference Resolution. arXiv:1707.07045
- Joshi, M. et al. (2019). SpanBERT: Improving Pre-training by Representing and Predicting Spans. arXiv:1907.10529
- Wu, W. et al. (2020). CorefQA: Coreference Resolution as Query-based Span Prediction. ACL.
Current state and open problems
Neural coreference resolvers now substantially outperform rule-based systems on standard benchmarks (OntoNotes), but a meaningful accuracy gap remains on genuinely hard cases: winograd-schema-style sentences where world knowledge, not syntax or gender agreement, decides the correct antecedent ("The trophy doesn't fit in the suitcase because it is too big" — resolving "it" requires knowing trophies and suitcases have different typical sizes, not any grammatical clue). Levesque, Davis & Morgenstern (2012) proposed the Winograd Schema Challenge specifically because such sentences resist the statistical shortcuts — gender, recency, syntactic position — that both rule-based and neural coreference systems otherwise lean on heavily. Cross-document coreference (linking mentions of the same entity across separate documents, rather than within one) and long-document coreference (tracking entities across text far longer than a transformer's context window comfortably handles) remain active, less-solved extensions of the single-document task.
What to learn next
- Multi-hop questions — a task where unresolved coreference directly causes wrong answers.
- Relation extraction — structured facts that depend on entities being correctly resolved first.
- Attention — the mechanism transformer-based coreference models use to connect a pronoun to distant context.