Legal and Compliance AI

Legal documents as data

Legal documents are numbered and hierarchical rather than free-flowing prose, and that numbering is the most reliable structure available for processing them.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Legal documents are built from numbered clauses, not flowing paragraphs. That numbering is the most useful structure a computer can grab onto.

Assembling flat-pack furniture, you do not read the instruction sheet like a novel, start to finish. You jump straight to step 7 when you need it, because every step carries its own number.

A contract works the same way. "Section 4.2" means something specific and findable. Treating a contract like ordinary prose throws away that structure, the exact thing that makes legal text usable at all.

Why it exists

Regular NLP tools were mostly built and tested on news articles, reviews, and social media posts. That kind of text flows as sentences and paragraphs, with no formal numbering a program can rely on.

Legal text is different on purpose. Clauses are numbered so a lawyer can say "see clause 4.2" and any reader can find it instantly. That numbering exists for humans, and it turns out to help a computer too.

A general text tool treats a contract as one long blob of prose. That throws away this ready-made structure before any real analysis begins.

How it works

  1. Definitions
     "Confidential Information" means any non-public information...

  2. Term
     This Agreement begins on the Effective Date...

  3. Termination
     Either party may terminate this Agreement with...

     3.1. Termination for Cause
          Either party may terminate immediately if...

  4. Confidentiality
     Each party shall keep the other's Confidential Information...

A program can split this into exact clauses and sub-clauses, using the numbering alone. This first step needs no understanding of English at all. Everything downstream, search, extraction, comparison, builds on getting it right.

Where you have already seen it

  • Contract review software. Tools that jump straight to "the termination clause" of an uploaded contract rely on exactly this kind of structural split.
  • Legal search engines. Searching case law by section number, rather than by page, depends on this same numbered structure.
  • "Ctrl+F doesn't work" complaints. When a scanned, unstructured PDF loses its clause numbers, both humans and tools struggle with it equally.

Remember this

  • Legal documents are numbered and hierarchical, unlike ordinary prose.
  • That numbering is a reliable structure a program can use, before any language understanding is needed.
  • Losing the numbering, for example in a badly scanned PDF, makes a document much harder to work with.

What to learn next

Developer — Code and libraries.

This splits a short contract into its numbered clauses and sub-clauses, using only the standard library.

Setup

Nothing to install. This uses only the re module from the standard library.

Splitting a contract into its numbered clauses

clause_split.py
import re

contract = """
1. Definitions
"Confidential Information" means any non-public information disclosed by either party.

2. Term
This Agreement begins on the Effective Date and continues for twelve (12) months.

3. Termination
Either party may terminate this Agreement with thirty (30) days written notice.

3.1. Termination for Cause
Either party may terminate immediately if the other party breaches a material term.

4. Confidentiality
Each party shall keep the other's Confidential Information secret for five (5) years.
"""

# Legal text is numbered, not free-flowing. That numbering is the only
# structure we get for free, so we split on it directly.
clause_pattern = re.compile(r"^(\d+(?:\.\d+)?)\.\s+(.+)$", re.MULTILINE)
matches = list(clause_pattern.finditer(contract))

clauses = []
for i, m in enumerate(matches):
    number, heading = m.group(1), m.group(2)
    start = m.end()
    end = matches[i + 1].start() if i + 1 < len(matches) else len(contract)
    body = contract[start:end].strip()
    clauses.append({"number": number, "heading": heading, "body": body})

for c in clauses:
    prefix = "  -> " if "." in c["number"] else ""
    print(f"{prefix}[{c['number']}] {c['heading']}")
    print(f"      {c['body'][:55]}...")
Output
[1] Definitions
      "Confidential Information" means any non-public informa...
[2] Term
      This Agreement begins on the Effective Date and continu...
[3] Termination
      Either party may terminate this Agreement with thirty (...
  -> [3.1] Termination for Cause
      Either party may terminate immediately if the other par...
[4] Confidentiality
      Each party shall keep the other's Confidential Informat...

Notice 3.1 prints with an arrow, correctly recognised as a sub-clause of 3. This split needed no understanding of what any clause actually says, only its number.

Line by line

r"^(\d+(?:\.\d+)?)\.\s+(.+)$" matches a number like 3 or 3.1 at the start of a line, followed by a period, then the heading text. The (?:\.\d+)? part makes the sub-clause number optional, so it matches both top-level and nested clauses.

re.MULTILINE makes ^ and $ match the start and end of each line, not only the start and end of the whole string. Without it, this pattern would only ever match once, at the very beginning of the contract.

matches[i + 1].start() finds where the next clause begins, so each clause's body runs right up to it. The final clause instead runs to the end of the whole text.

Common mistakes

Assuming every contract uses this exact numbering style. Real contracts vary: some use (a), (i), or Article 4 instead of 4.. A production system needs several patterns, tried in order, or a more flexible parser.

Splitting on every period in the text. A naive text.split(".") would also break on "twelve (12) months." mid-sentence, and on abbreviations. Anchoring the pattern to the start of a line, as done here, avoids this.

Forgetting nested clauses reference their parent. 3.1 is meaningless without knowing it sits under 3. Keep the hierarchy, not only a flat list, if the downstream use needs it.

Try it yourself

Add a 5.1.1 sub-sub-clause to the contract text, one level deeper than 3.1. Rerun the script and see whether the current pattern still detects it correctly.

It should, since (?:\.\d+)? only allows one extra level. Extending the pattern to handle arbitrary depth, using (?:\.\d+)* instead, is a small, worthwhile exercise.

What to learn next

Researcher — Mathematics and papers.

Document structure as a weak-supervision signal

Formally, a legal document can be modelled as a tree T = (V, E), where each node v corresponds to a clause or sub-clause, labelled with a hierarchical identifier (3, 3.1, 3.1.1), and edges encode parent-child nesting inferred directly from identifier structure.

This structure is a form of weak supervision: it costs nothing to extract, requires no manual labelling, and yet provides a strong prior for downstream tasks. Clause boundaries found this way are far more reliable than boundaries found by generic sentence or paragraph segmentation, which was tuned on prose without this kind of formal numbering.

Corpora and standard formats

CUAD (Contract Understanding Atticus Dataset, Hendrycks et al., 2021) provides 510 commercial contracts with expert-annotated clause categories, and is the standard benchmark for clause-level extraction and classification tasks in this domain.

Akoma Ntoso and LegalDocML are XML standards for representing legal documents with explicit, machine-readable hierarchical structure. Several national parliaments and legal publishers use them directly. The regex-based recovery in the developer block is only necessary when a document was never produced in a structured format to begin with.

Why generic NLP pipelines underperform here

Chalkidis et al. (2022), in the LexGLUE benchmark paper, document a consistent gap. Models pretrained on general web text underperform on legal classification and extraction tasks, compared to models pretrained on legal corpora specifically (LegalBERT, CaseLawBERT), even at matched parameter count.

Two structural reasons are commonly cited. First, legal sentences are long, frequently exceeding 100 words with heavy subordinate clause nesting, straining sentence-segmentation and context-window assumptions tuned on shorter, ordinary prose. Second, legal vocabulary includes terms of art whose meaning is standardised by centuries of case law, not by ordinary usage, which general-purpose tokenizers and embeddings do not specially account for.

Complexity

Regex-based clause splitting, as in the developer block, is O(n) in document length, dominated by a single linear scan. Building the full hierarchical tree from flat clause identifiers is O(n log n) if identifiers are sorted first, or O(n) with a stack-based parse that tracks nesting depth as it goes.

Key references

  • Hendrycks, D. et al. (2021). CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review. arXiv:2103.06268
  • Chalkidis, I. et al. (2020). LEGAL-BERT: The Muppets straight out of Law School. arXiv:2010.02559
  • Chalkidis, I. et al. (2022). LexGLUE: A Benchmark Dataset for Legal Language Understanding in English. arXiv:2110.00976

Current state and open problems

Structural extraction from well-formatted contracts is largely a solved engineering problem, with mature commercial tools handling the common numbering conventions reliably.

The unsolved part is the long tail. Contracts scanned with inconsistent OCR quality, documents mixing numbering conventions in one file, jurisdictions with structural conventions unlike the section-and-subsection style used above. General-purpose document-layout models increasingly handle this tail, at higher compute cost than the pattern-matching shown here, and with their own new failure modes.

What to learn next