Sequence Labelling and Structure

Relation extraction

Relation extraction turns two entities and a sentence into a structured fact, such as (Zoho, acquired, a startup), which is what lets unstructured text feed a database or a knowledge graph.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Relation extraction turns a sentence into a structured fact: who did what to whom.

Think about a court stenographer. Their job is condensing a witness's rambling testimony into a clean statement for the record: "the accused handed the package to the driver." All the extra words get dropped. What remains is exactly who did what to whom.

Relation extraction does the same thing to any sentence. "Zoho acquired a small startup in 2023" becomes the structured fact (Zoho, acquired, a startup). A database or spreadsheet can actually store and search that clean triple. A computer can only treat a paragraph as an opaque blob of characters.

Why it exists

Named entity recognition finds "Zoho" and "a startup" as entities. It does not say how they relate to each other. Two sentences — "Zoho acquired a startup" and "A startup acquired Zoho" — mention exactly the same two entities. They mean the opposite thing.

Relation extraction fills that gap. It looks at the entities a sentence mentions. It works out the connection between them. The result is structured enough to store in a database, or search over. It can feed into a knowledge graph, a network of entities connected by exactly these kinds of facts.

How it works

A relation extractor is given a sentence with two identified entities. It decides what connects them, usually by looking at the words and grammatical structure between and around them.

  Sentence:  "The RBI raised interest rates on Wednesday."

  Entities found:   RBI (ORG)         interest rates (CONCEPT)

  Relation extractor looks at the connecting word "raised" and the
  grammatical structure around it:

  Result:    ( RBI ,  raised ,  interest rates )
               subject   relation      object

The connecting verb is often the strongest clue, but not the only one. Word order, prepositions, and the grammatical role each entity plays — subject versus object — all matter. A wrong role reverses the meaning of the fact.

Where you have already seen it

  • Google's Knowledge Graph panels, built partly by extracting structured facts like "(company, founded by, person)" from web text at enormous scale.
  • Financial news monitoring. Systems automatically extract "(company, acquired, company)" or "(company, raised, amount)" facts, tracking deal activity without a human reading every article.
  • Drug interaction databases, built partly by extracting "(drug A, interacts with, drug B)" facts from medical literature. It is faster than any team of humans could read.

Remember this

  • Relation extraction produces a structured fact — typically a (subject, relation, object) triple — from a sentence and two entities within it.
  • Word order and grammatical role matter: swapping subject and object usually reverses the meaning.
  • The output is meant to be stored and queried, not read as prose — that is what makes it useful for databases and knowledge graphs.

What to learn next

  • Dependency parsing — the grammatical structure relation extraction leans on most directly.
  • Coreference resolution — resolving pronouns first so relations are not lost across sentences.
  • Entity linking — identifying exactly which real-world entities a relation connects.

Developer — Code and libraries.

A dependency parse already encodes subject-verb-object structure. Pulling relations out of it is direct, and needs nothing beyond spaCy's default pipeline.

Setup

bash
pip install spacy
python -m spacy download en_core_web_sm

Minimal runnable code

relation_extract.py
import spacy

nlp = spacy.load("en_core_web_sm")

def extract_svo_triples(text):
    doc = nlp(text)
    triples = []
    for token in doc:
        if token.pos_ == "VERB":
            subjects = [c for c in token.children if c.dep_ in ("nsubj", "nsubjpass")]
            objects = [c for c in token.children if c.dep_ in ("dobj", "attr", "pobj")]
            for c in token.children:
                if c.dep_ == "prep":
                    objects += [g for g in c.children if g.dep_ == "pobj"]
            for s in subjects:
                for o in objects:
                    triples.append((s.text, token.lemma_, o.text))
    return triples

sentences = [
    "Zoho acquired a small startup in 2023.",
    "Infosys hired thousands of engineers last year.",
    "The RBI raised interest rates on Wednesday.",
]
for sent in sentences:
    for subj, rel, obj in extract_svo_triples(sent):
        print(f"({subj}, {rel}, {obj})")
Output
(Zoho, acquire, startup)
(Zoho, acquire, 2023)
(Infosys, hire, thousands)
(RBI, raise, rates)
(RBI, raise, Wednesday)

Line by line

token.pos_ == "VERB" finds every action word, then looks at its direct grammatical children — see Dependency parsing for what token.children and dep_ mean in detail.

nsubj / nsubjpass find the subject, including passive-voice subjects ("was acquired by" style sentences), while dobj finds a direct object and pobj catches objects hiding behind a preposition, such as "of engineers" attaching to "hired".

token.lemma_ normalises the verb — "acquired" becomes "acquire", "raised" becomes "raise" — so the same relation extracted from different tenses lands on the same relation name, which matters if you are storing these triples for later lookup.

Reading the noise honestly

Look closely at the output: (Zoho, acquire, 2023) and (RBI, raise, Wednesday) are wrong. They come from time expressions — "in 2023", "on Wednesday" — that share the same pobj dependency label as a genuine object, because the parser cannot always tell a true object apart from an adjunct time phrase using dependency labels alone. This is not a bug in the code above. It is a real, well-known limitation of dependency-based relation extraction: the grammar of "acquired a startup" and "acquired ... in 2023" looks similar enough at the dependency-label level that simple pattern rules cannot always separate them.

Production systems handle this by filtering candidate objects with a named-entity or type check — a date-shaped span is very unlikely to be the true object of "acquire" — or by training a proper classifier on the relation itself, rather than deriving it purely from dependency labels. See the researcher block for how that is usually done.

Common mistakes

Trusting every triple a dependency-based extractor produces. As shown above, prepositional phrases attached to time or manner, not a true object, regularly leak into the output. Always sanity-check or filter by entity type before storing extracted relations.

Ignoring passive voice. "A startup was acquired by Zoho" reverses subject and object compared to "Zoho acquired a startup", but both describe the same real-world fact. Code that only checks nsubj and misses nsubjpass will get the direction backwards on passive sentences.

Assuming one verb means one relation type. "Raised" means something completely different in "the RBI raised interest rates" versus "the company raised $2 million". Mapping verb lemmas directly to fixed relation types without checking the object's entity type conflates unrelated facts.

Try it yourself

Add the sentence "A small startup was acquired by Zoho in 2023." and check whether extract_svo_triples correctly finds (startup, acquire, Zoho) given the passive-voice nsubjpass handling already in the code, or whether the direction comes out backwards.

What to learn next

Researcher — Mathematics and papers.

The task, formally

Given a sentence x and two identified entity spans (e_1, e_2) within it, relation extraction predicts a relation type r from a fixed schema R, or NO_RELATION if the entities are unrelated in this sentence:

text
r* = argmax over r in R ∪ {NO_RELATION} of P(r | x, e_1, e_2)

This is a strict simplification of the developer block's open, pattern-based extraction — a closed-relation classifier assumes a fixed, predefined relation schema, rather than discovering arbitrary relation phrases like "acquired" or "raised" directly from text.

Closed versus open relation extraction

Closed (schema-based) relation extraction trains a classifier over (sentence, entity pair) inputs against a fixed relation inventory (per:employee_of, org:founded_by, in the TACRED schema). Standard architecture: encode the sentence with entity-position markers inserted around each entity span, feed a pretrained transformer, classify the pooled or entity-token representations (Zhang et al., 2017; Baldini Soares et al., 2019).

Open information extraction (Open IE) makes no assumption about a fixed relation set, extracting (subject phrase, relation phrase, object phrase) triples directly from surface syntax — the developer block's dependency-pattern approach is a simplified instance of this family. Systems like ReVerb (Fader et al., 2011) and OpenIE 5 add syntactic and lexical constraints specifically to filter out the adjunct-versus-object confusion demonstrated in the developer block.

Distant supervision

Hand-labelling sentences with relation types is expensive. Distant supervision (Mintz et al., 2009) sidesteps this: given a knowledge base fact (e_1, r, e_2) already known to be true, every sentence mentioning both e_1 and e_2 is heuristically labelled with relation r, generating large amounts of noisy training data automatically. The core assumption — that any co-occurring mention expresses the known relation — is frequently false (two entities can co-occur in a sentence for unrelated reasons), so distant supervision trades label quality for label quantity, and most later work (Riedel et al., 2010; multi-instance learning formulations) exists specifically to model and reduce that noise.

Neural relation extraction

Zhang et al. (2017) show that combining a dependency-tree-pruned representation — keeping only the shortest dependency path between the two entities, discarding unrelated branches of the parse — with a graph convolutional network measurably outperforms sequence-only models, providing direct empirical support for using dependency structure (as the developer block does by hand) rather than raw word sequence alone. Baldini Soares et al. (2019), "Matching the Blanks," show a BERT-based relation classifier trained with an entity-blank pretraining objective, requiring very little task-specific labelled data.

Document-level and cross-sentence relation extraction (Yao et al., 2019, DocRED) extends the task beyond single sentences, since many real facts — particularly ones requiring coreference resolution across sentences — cannot be recovered from any single sentence in isolation.

Key references

  • Mintz, M. et al. (2009). Distant Supervision for Relation Extraction Without Labeled Data. ACL.
  • Fader, A., Soderland, S. & Etzioni, O. (2011). Identifying Relations for Open Information Extraction. EMNLP — ReVerb.
  • Riedel, S., Yao, L. & McCallum, A. (2010). Modeling Relations and Their Mentions without Labeled Text. ECML PKDD.
  • Zhang, Y., Qi, P. & Manning, C. (2017). Graph Convolution over Pruned Dependency Trees Improves Relation Extraction. arXiv:1809.10185
  • Baldini Soares, L. et al. (2019). Matching the Blanks: Distributional Similarity for Relation Learning. arXiv:1906.03158
  • Yao, Y. et al. (2019). DocRED: A Large-Scale Document-Level Relation Extraction Dataset. arXiv:1906.06127

Current state and open problems

Sentence-level relation extraction on well-studied schemas (TACRED, SemEval-2010 Task 8) is a largely mature task for transformer-based classifiers, with distant supervision and its noise-reduction techniques the dominant way to obtain training data at scale. Document-level relation extraction, requiring reasoning across multiple sentences and often depending on correct coreference resolution first, remains substantially harder, with DocRED-style benchmarks still far from human performance. Large language models used zero-shot or few-shot for open-domain relation extraction are an active and unsettled area: they can extract plausible-looking triples for arbitrary, previously unseen relation types without task-specific training, but systematic evaluation of their precision at scale, compared to purpose-built extractors, is still an open empirical question rather than a settled one.

What to learn next