Extracting clauses from contracts
Clause extraction sorts contract clauses into categories like termination or payment, and a simple keyword approach reliably misses clauses that say the same thing in different words.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Clause extraction sorts each part of a contract into a category, like "termination" or "payment". A reader can then find what matters without reading everything.
A librarian sorts a returned pile of books by genre, glancing at each cover. Most books sort themselves instantly. A book with an unusual, misleading cover can end up on the wrong shelf.
Clause extraction is that sorting job, done on a contract. Given a clause's text, decide what category it belongs to: termination, confidentiality, payment, liability, and more. Simple approaches get most clauses right, and confidently misplace the rest.
Why it exists
A 40-page contract might have one clause a lawyer actually needs right now, say, the termination terms. Reading all 40 pages to find it wastes time that adds up fast across hundreds of contracts.
Clause extraction lets software jump straight to the relevant category. Instead of "read the whole contract", the task becomes "show me every termination clause across these 200 supplier agreements".
That single capability is the foundation nearly every other legal AI tool in this section builds on top of.
How it works
Clause: "Either party may terminate this Agreement
with thirty (30) days written notice."
Check for category keywords:
"terminat..." -> found!
Category: TERMINATION Clause: "The parties may bring this partnership
to a close by mutual agreement."
Check for category keywords:
"terminat..." -> not found
(no obvious keyword matches any category)
Category: UNKNOWN <- this is actually a termination clause tooThe second example is the real lesson here. A simple keyword check misses a clause that says the same thing in different words. A careful human reader would not make that mistake.
Where you have already seen it
- Contract review software. Highlighting every "indemnification" or "liability" clause across an uploaded document automatically.
- Due diligence in mergers. Reviewers scanning hundreds of acquired-company contracts for specific risky clause types, at speed.
- Contract template checkers. Flagging a contract that is missing a clause type your company's policy requires, like a data-protection clause.
Remember this
- Clause extraction sorts contract text into categories like termination, payment or confidentiality.
- Keyword-based approaches work well until a clause is phrased without the expected keyword.
- Missed clauses fail silently: the system shows no error, it never surfaces them at all.
What to learn next
- Legal documents as data — the clause-splitting step this technique builds directly on.
- Named entity recognition — a related extraction task, pulling out names and dates instead of clause types.
- Text classification — the general technique behind smarter, model-based clause categorisation.
Developer — Code and libraries.
This classifies five clauses by keyword, then shows exactly where that approach breaks: a real termination clause, phrased without the word "terminate".
Setup
Nothing to install. This uses only plain Python.
Keyword-based clause classification
CATEGORY_KEYWORDS = {
"termination": ["terminat", "end this agreement", "notice period"],
"confidentiality": ["confidential", "non-disclosure", "secret"],
"payment": ["pay", "invoice", "fee", "compensation"],
"liability": ["liable", "liability", "damages", "indemnif"],
}
clauses = [
"Either party may terminate this Agreement with thirty (30) days written notice.",
"Each party shall keep the other's Confidential Information secret for five years.",
"Client shall pay all invoices within thirty (30) days of receipt.",
"Neither party shall be liable for indirect or consequential damages.",
"The parties may bring this partnership to a close by mutual agreement.",
]
def classify(clause_text):
text = clause_text.lower()
hits = [category for category, keywords in CATEGORY_KEYWORDS.items()
if any(kw in text for kw in keywords)]
return hits or ["uncategorised"]
for clause in clauses:
print(f"{classify(clause)}: {clause[:55]}...")['termination']: Either party may terminate this Agreement with thirty (... ['confidentiality']: Each party shall keep the other's Confidential Informat... ['payment']: Client shall pay all invoices within thirty (30) days o... ['liability']: Neither party shall be liable for indirect or consequen... ['uncategorised']: The parties may bring this partnership to a close by mu...
Four out of five clauses are correctly tagged. The fifth, a genuine termination clause written without the word "terminate", falls through completely. It does not get miscategorised; it gets ignored, which is the more dangerous failure of the two.
Line by line
CATEGORY_KEYWORDS is a dictionary of category to keyword list. Extending coverage means manually adding keywords, which does not scale to the enormous variety of legal phrasing lawyers actually use.
any(kw in text for kw in keywords) checks if any keyword appears anywhere in the clause, as a plain substring. This has no notion of meaning at all, only literal text matching.
hits or ["uncategorised"] falls back to a clear "uncategorised" label rather than silently returning an empty list. Always make a missed classification visible, so it can at least be reviewed by a person, rather than disappearing.
Common mistakes
Treating "uncategorised" results as noise to ignore. As shown above, uncategorised can mean "genuinely doesn't fit any category", or it can mean "a real match the keyword list happened to miss". These need different handling.
Assuming more keywords fixes the problem completely. You can add "bring...to a close" to the termination list, and the fifth clause fixes itself. Real contracts will keep finding new phrasings a growing keyword list cannot fully anticipate.
Ignoring that one clause can belong to multiple categories. A termination clause with financial penalties attached is both a termination clause and a payment clause. Forcing a single label loses that overlap.
Try it yourself
Add "bring this partnership to a close" to the "termination" keyword list, and rerun. Confirm the fifth clause now classifies correctly.
Then write one more clause, phrased differently again, that should also count as termination but uses none of the current keywords. This is the never-ending game keyword lists play, and it motivates the model-based approaches covered in later lessons.
What to learn next
- Named entity recognition — a technique for pulling out specific facts, like dates and party names, from within a clause.
- Text classification — training a model to catch paraphrased clauses this keyword approach misses.
- E-discovery and technology-assisted review — the same classify-and-rank idea, applied to entire documents instead of single clauses.
Researcher — Mathematics and papers.
Framing clause extraction as a task
Clause extraction is usually decomposed into two sub-tasks, sometimes handled jointly and sometimes separately:
segmentation: document -> [clause_1, ..., clause_n]
classification: clause_i -> category (single-label or multi-label)The developer block's regex-based segmentation, from the previous lesson, plus a rule-based classifier here, is the simplest possible pipeline. Production systems typically replace the classification step with a fine-tuned transformer, and often the segmentation step too, when clause boundaries are not reliably numbered.
The CUAD benchmark
Hendrycks et al. (2021) frame clause extraction as span extraction with multi-label categories. For each of 41 predefined clause types (governing law, non-compete, IP ownership, and others), locate the specific span of text, if any, that instantiates it. This differs from the developer block's whole-clause classification: it works at the sub-clause span level, and allows a contract zero, one, or several spans per category.
Baseline transformer models fine-tuned specifically on CUAD substantially outperform both keyword rules and general-purpose extractive QA applied zero-shot. This task benefits meaningfully from domain-specific supervision, not from general capability alone.
Multi-label classification and its metrics
Because a clause genuinely can belong to multiple categories simultaneously, evaluation uses multi-label metrics rather than simple accuracy:
precision_c = TP_c / (TP_c + FP_c) for each category c
recall_c = TP_c / (TP_c + FN_c)TP_c,FP_c,FN_care true positives, false positives and false negatives for categorycspecifically.- Macro-averaging these across categories weights rare clause types equally with common ones; micro-averaging instead weights every individual clause instance equally.
The choice between macro and micro averaging materially changes reported performance on CUAD specifically, since clause type frequency is highly imbalanced: some categories appear in nearly every contract, others in only a handful.
Why recall matters more here than in general NLP
In most classification settings, precision and recall are balanced concerns. In legal clause extraction, a missed clause typically costs far more than a spurious one. A lawyer reviewing a flagged-but-irrelevant clause loses a few minutes; a missed liability clause can mean genuine legal exposure. This asymmetry pushes production systems toward deliberately recall-favouring thresholds, trading more false positives for fewer misses.
Complexity
Rule-based classification, as in the developer block, is O(n * k) for n clauses and k keywords, negligible in practice. Transformer-based classification costs one forward pass per clause, O(n) forward passes total, dominant in wall-clock time but still modest for a single contract of ordinary length.
Key references
- Hendrycks, D. et al. (2021). CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review. arXiv:2103.06268
- Chalkidis, I. et al. (2019). Neural Contract Element Extraction Revisited. Workshop on Document Intelligence.
- Leivaditi, S., Rossi, J. & Kanoulas, E. (2020). A Benchmark for Lease Contract Review. arXiv:2010.10386
Current state and open problems
Transformer-based clause classification, fine-tuned on domain data such as CUAD, is standard practice in commercial contract-review tools today, comfortably outperforming rule-based approaches on paraphrase robustness specifically.
The open problem is coverage of the long tail: clause types and phrasings not represented in any existing benchmark, across jurisdictions and languages CUAD-style English-language datasets do not cover. Extending reliable clause extraction beyond a small number of well-resourced legal systems remains substantially unsolved, mirroring the broader multilingual gap discussed in this course's NLP sections.
What to learn next
- Text classification — the general supervised-learning technique this task specialises.
- E-discovery and technology-assisted review — where recall-favouring thresholds matter equally, at document scale.
- Evaluating a legal AI system — the precision/recall trade-off discussed above, made concrete.