Legal and Compliance AI

Extracting clauses from contracts

Clause extraction sorts contract clauses into categories like termination or payment, and a simple keyword approach reliably misses clauses that say the same thing in different words.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Clause extraction sorts each part of a contract into a category, like "termination" or "payment". A reader can then find what matters without reading everything.

A librarian sorts a returned pile of books by genre, glancing at each cover. Most books sort themselves instantly. A book with an unusual, misleading cover can end up on the wrong shelf.

Clause extraction is that sorting job, done on a contract. Given a clause's text, decide what category it belongs to: termination, confidentiality, payment, liability, and more. Simple approaches get most clauses right, and confidently misplace the rest.

Why it exists

A 40-page contract might have one clause a lawyer actually needs right now, say, the termination terms. Reading all 40 pages to find it wastes time that adds up fast across hundreds of contracts.

Clause extraction lets software jump straight to the relevant category. Instead of "read the whole contract", the task becomes "show me every termination clause across these 200 supplier agreements".

That single capability is the foundation nearly every other legal AI tool in this section builds on top of.

How it works

  Clause: "Either party may terminate this Agreement
           with thirty (30) days written notice."

  Check for category keywords:
    "terminat..."  -> found!

  Category: TERMINATION
  Clause: "The parties may bring this partnership
           to a close by mutual agreement."

  Check for category keywords:
    "terminat..."  -> not found
    (no obvious keyword matches any category)

  Category: UNKNOWN     <- this is actually a termination clause too

The second example is the real lesson here. A simple keyword check misses a clause that says the same thing in different words. A careful human reader would not make that mistake.

Where you have already seen it

  • Contract review software. Highlighting every "indemnification" or "liability" clause across an uploaded document automatically.
  • Due diligence in mergers. Reviewers scanning hundreds of acquired-company contracts for specific risky clause types, at speed.
  • Contract template checkers. Flagging a contract that is missing a clause type your company's policy requires, like a data-protection clause.

Remember this

  • Clause extraction sorts contract text into categories like termination, payment or confidentiality.
  • Keyword-based approaches work well until a clause is phrased without the expected keyword.
  • Missed clauses fail silently: the system shows no error, it never surfaces them at all.

What to learn next

Developer — Code and libraries.

This classifies five clauses by keyword, then shows exactly where that approach breaks: a real termination clause, phrased without the word "terminate".

Setup

Nothing to install. This uses only plain Python.

Keyword-based clause classification

clause_extract.py
CATEGORY_KEYWORDS = {
    "termination":     ["terminat", "end this agreement", "notice period"],
    "confidentiality": ["confidential", "non-disclosure", "secret"],
    "payment":         ["pay", "invoice", "fee", "compensation"],
    "liability":       ["liable", "liability", "damages", "indemnif"],
}

clauses = [
    "Either party may terminate this Agreement with thirty (30) days written notice.",
    "Each party shall keep the other's Confidential Information secret for five years.",
    "Client shall pay all invoices within thirty (30) days of receipt.",
    "Neither party shall be liable for indirect or consequential damages.",
    "The parties may bring this partnership to a close by mutual agreement.",
]

def classify(clause_text):
    text = clause_text.lower()
    hits = [category for category, keywords in CATEGORY_KEYWORDS.items()
            if any(kw in text for kw in keywords)]
    return hits or ["uncategorised"]

for clause in clauses:
    print(f"{classify(clause)}: {clause[:55]}...")
Output
['termination']: Either party may terminate this Agreement with thirty (...
['confidentiality']: Each party shall keep the other's Confidential Informat...
['payment']: Client shall pay all invoices within thirty (30) days o...
['liability']: Neither party shall be liable for indirect or consequen...
['uncategorised']: The parties may bring this partnership to a close by mu...

Four out of five clauses are correctly tagged. The fifth, a genuine termination clause written without the word "terminate", falls through completely. It does not get miscategorised; it gets ignored, which is the more dangerous failure of the two.

Line by line

CATEGORY_KEYWORDS is a dictionary of category to keyword list. Extending coverage means manually adding keywords, which does not scale to the enormous variety of legal phrasing lawyers actually use.

any(kw in text for kw in keywords) checks if any keyword appears anywhere in the clause, as a plain substring. This has no notion of meaning at all, only literal text matching.

hits or ["uncategorised"] falls back to a clear "uncategorised" label rather than silently returning an empty list. Always make a missed classification visible, so it can at least be reviewed by a person, rather than disappearing.

Common mistakes

Treating "uncategorised" results as noise to ignore. As shown above, uncategorised can mean "genuinely doesn't fit any category", or it can mean "a real match the keyword list happened to miss". These need different handling.

Assuming more keywords fixes the problem completely. You can add "bring...to a close" to the termination list, and the fifth clause fixes itself. Real contracts will keep finding new phrasings a growing keyword list cannot fully anticipate.

Ignoring that one clause can belong to multiple categories. A termination clause with financial penalties attached is both a termination clause and a payment clause. Forcing a single label loses that overlap.

Try it yourself

Add "bring this partnership to a close" to the "termination" keyword list, and rerun. Confirm the fifth clause now classifies correctly.

Then write one more clause, phrased differently again, that should also count as termination but uses none of the current keywords. This is the never-ending game keyword lists play, and it motivates the model-based approaches covered in later lessons.

What to learn next

Researcher — Mathematics and papers.

Framing clause extraction as a task

Clause extraction is usually decomposed into two sub-tasks, sometimes handled jointly and sometimes separately:

text
segmentation:    document -> [clause_1, ..., clause_n]
classification:  clause_i -> category (single-label or multi-label)

The developer block's regex-based segmentation, from the previous lesson, plus a rule-based classifier here, is the simplest possible pipeline. Production systems typically replace the classification step with a fine-tuned transformer, and often the segmentation step too, when clause boundaries are not reliably numbered.

The CUAD benchmark

Hendrycks et al. (2021) frame clause extraction as span extraction with multi-label categories. For each of 41 predefined clause types (governing law, non-compete, IP ownership, and others), locate the specific span of text, if any, that instantiates it. This differs from the developer block's whole-clause classification: it works at the sub-clause span level, and allows a contract zero, one, or several spans per category.

Baseline transformer models fine-tuned specifically on CUAD substantially outperform both keyword rules and general-purpose extractive QA applied zero-shot. This task benefits meaningfully from domain-specific supervision, not from general capability alone.

Multi-label classification and its metrics

Because a clause genuinely can belong to multiple categories simultaneously, evaluation uses multi-label metrics rather than simple accuracy:

text
precision_c = TP_c / (TP_c + FP_c)     for each category c
recall_c    = TP_c / (TP_c + FN_c)
  • TP_c, FP_c, FN_c are true positives, false positives and false negatives for category c specifically.
  • Macro-averaging these across categories weights rare clause types equally with common ones; micro-averaging instead weights every individual clause instance equally.

The choice between macro and micro averaging materially changes reported performance on CUAD specifically, since clause type frequency is highly imbalanced: some categories appear in nearly every contract, others in only a handful.

Why recall matters more here than in general NLP

In most classification settings, precision and recall are balanced concerns. In legal clause extraction, a missed clause typically costs far more than a spurious one. A lawyer reviewing a flagged-but-irrelevant clause loses a few minutes; a missed liability clause can mean genuine legal exposure. This asymmetry pushes production systems toward deliberately recall-favouring thresholds, trading more false positives for fewer misses.

Complexity

Rule-based classification, as in the developer block, is O(n * k) for n clauses and k keywords, negligible in practice. Transformer-based classification costs one forward pass per clause, O(n) forward passes total, dominant in wall-clock time but still modest for a single contract of ordinary length.

Key references

  • Hendrycks, D. et al. (2021). CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review. arXiv:2103.06268
  • Chalkidis, I. et al. (2019). Neural Contract Element Extraction Revisited. Workshop on Document Intelligence.
  • Leivaditi, S., Rossi, J. & Kanoulas, E. (2020). A Benchmark for Lease Contract Review. arXiv:2010.10386

Current state and open problems

Transformer-based clause classification, fine-tuned on domain data such as CUAD, is standard practice in commercial contract-review tools today, comfortably outperforming rule-based approaches on paraphrase robustness specifically.

The open problem is coverage of the long tail: clause types and phrasings not represented in any existing benchmark, across jurisdictions and languages CUAD-style English-language datasets do not cover. Extending reliable clause extraction beyond a small number of well-resourced legal systems remains substantially unsolved, mirroring the broader multilingual gap discussed in this course's NLP sections.

What to learn next