Legal and Compliance AI

Comparing contract versions

Redlining finds the exact words that changed between two contract drafts, so a reviewer checks only what moved instead of reading both versions in full.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Redlining finds exactly which words changed between two versions of a contract, so a reviewer checks only what moved.

A "spot the difference" puzzle gives you two nearly identical pictures side by side. You do not study every pixel. Your eyes jump straight to the handful of spots that actually changed.

Comparing two contract drafts is the same puzzle, at a larger scale. Redlining is the process, and technique, of finding exactly which words changed, so a reviewer can check only those spots.

Why it exists

Contract negotiation means sending a draft back and forth, both sides editing it each round. After five rounds, the current draft can look nothing like the one you last carefully read in full.

Re-reading the entire contract every round, hunting for what changed by eye, is slow and genuinely error-prone. A single quietly-changed number, thirty days becoming ninety, is easy to miss buried in ten pages of otherwise identical text.

Redlining automates the "spot the difference" step, so a human reviewer's attention goes exactly where it is needed.

How it works

  Version 1: "...deliver the goods within thirty (30) days..."
  Version 2: "...deliver the goods within fifteen (15) days..."

  Redlined:  "...deliver the goods within
              [-thirty (30)-]{+fifteen (15)+} days..."

              [-removed-]   {+added+}

Everything unchanged stays as plain text. Everything removed and everything added gets marked right where it happened. No side-by-side read of two full documents is needed.

Where you have already seen it

  • Track changes in Word. The same visual idea, showing insertions and deletions inline, that most people have used for essays or reports.
  • Contract lifecycle management software. Legal teams reviewing vendor contracts rely on automatic redlining to spot changes across negotiation rounds quickly.
  • Code review tools. A code "diff" is the exact same technique, applied to source code instead of legal prose.

Remember this

  • Redlining finds exactly what changed between two document versions, word by word.
  • It saves a reviewer from re-reading an entire document to catch one changed number.
  • The same underlying technique powers Word's track changes and a programmer's code diff.

What to learn next

Developer — Code and libraries.

This compares two versions of a short contract, first line by line, then word by word for a specific changed line, using only the standard library.

Setup

Nothing to install. This uses only the difflib module from the standard library.

Line-level and word-level diffs

redline_demo.py
import difflib

version_1 = """The Vendor shall deliver the goods within thirty (30) days of order confirmation.
Payment is due within sixty (60) days of delivery.
Either party may terminate this Agreement with ninety (90) days written notice."""

version_2 = """The Vendor shall deliver the goods within fifteen (15) days of order confirmation.
Payment is due within sixty (60) days of delivery.
Either party may terminate this Agreement with thirty (30) days written notice.
This Agreement is governed by the laws of the State of Delaware."""

lines_1 = version_1.splitlines()
lines_2 = version_2.splitlines()

diff = difflib.unified_diff(lines_1, lines_2, lineterm="", n=0)
for line in diff:
    if line.startswith(("+++", "---", "@@")):
        continue
    tag = "ADDED  " if line.startswith("+") else "REMOVED" if line.startswith("-") else ""
    print(f"{tag} {line[1:].strip()}")

# Now zoom into exactly what changed on the first line, word by word.
print("\nword-level view of the delivery-time change:")
old_words, new_words = lines_1[0].split(), lines_2[0].split()
sm = difflib.SequenceMatcher(None, old_words, new_words)
out = []
for op, i1, i2, j1, j2 in sm.get_opcodes():
    if op == "equal":
        out.append(" ".join(old_words[i1:i2]))
    elif op == "replace":
        out.append(f"[-{' '.join(old_words[i1:i2])}-]{{+{' '.join(new_words[j1:j2])}+}}")
    elif op == "delete":
        out.append(f"[-{' '.join(old_words[i1:i2])}-]")
    elif op == "insert":
        out.append(f"{{+{' '.join(new_words[j1:j2])}+}}")
print(" ".join(out))
Output
REMOVED The Vendor shall deliver the goods within thirty (30) days of order confirmation.
ADDED   The Vendor shall deliver the goods within fifteen (15) days of order confirmation.
REMOVED Either party may terminate this Agreement with ninety (90) days written notice.
ADDED   Either party may terminate this Agreement with thirty (30) days written notice.
ADDED   This Agreement is governed by the laws of the State of Delaware.

word-level view of the delivery-time change:
The Vendor shall deliver the goods within [-thirty (30)-]{+fifteen (15)+} days of order confirmation.

The line-level diff correctly flags three changes: a shorter delivery window, a shorter notice period, and one entirely new clause. The word-level view then zooms into exactly which words changed within one specific line, the way Word's track changes marks up text inline.

Line by line

difflib.unified_diff(..., n=0) shows zero lines of unchanged context around each change, useful here for a clean, compact demonstration. Real tools usually show a few surrounding lines, so a reviewer sees each change in context.

SequenceMatcher(None, old_words, new_words) finds the longest matching blocks between two lists, here lists of words, then reports the gaps between them as insertions, deletions or replacements.

get_opcodes() returns exactly what kind of change happened at each position: "equal", "replace", "delete" or "insert", which is what lets the code format each kind differently.

Common mistakes

Diffing at the character level for legal text. A single retyped sentence with the same meaning can look like a huge character-level change, when word-level or sentence-level diffing would show it as one small, meaningful edit.

Ignoring line-order changes. A clause moved from position 3 to position 7, unchanged otherwise, can appear as a confusing tangle of removals and additions in a naive diff, rather than as one clean "moved" event.

Treating an empty diff as "nothing to review". Some changes are formatting-only, whitespace or line breaks, and genuinely do not need review. Others silently reformat text in a way that quietly changes its meaning; always sanity-check what a diff tool is actually comparing.

Try it yourself

Add a fourth, unchanged sentence to both versions, identical in both. Confirm the diff correctly shows zero changes for that line.

Then swap the order of two unchanged sentences in version_2 only. Rerun and see how the diff represents a pure reordering, worth understanding before trusting a diff tool on a real contract.

What to learn next

Researcher — Mathematics and papers.

The longest common subsequence formulation

Text diffing is fundamentally a longest common subsequence (LCS) problem: given two sequences, find the longest subsequence common to both, in order, though not necessarily contiguous. Everything not in that common subsequence is either a deletion or an insertion.

text
LCS(A, B) via dynamic programming:
  dp[i][j] = dp[i-1][j-1] + 1              if A[i] == B[j]
  dp[i][j] = max(dp[i-1][j], dp[i][j-1])   otherwise
  • dp[i][j] holds the LCS length of the first i elements of A and first j elements of B.
  • This runs in O(|A| * |B|) time and space with the naive dynamic-programming table, the classical algorithm behind most general-purpose diff tools.

Python's difflib.SequenceMatcher uses a different algorithm, based on Ratcliff and Obershelp's "gestalt pattern matching", which is not guaranteed to find the true minimal-edit LCS in every case, but runs faster in typical practice and produces good results for text.

Plain line-based or word-based diffing, as in the developer block, ignores document structure entirely. Structure-aware diffing instead diffs at the clause level first, using the numbered-clause structure from the legal-documents-as-data lesson, then diffs word-by-word only within matched clause pairs.

This matters because clause reordering is common in contract negotiation, moving a clause from section 5 to section 8 without changing its wording. Naive line-based diffing represents this as a large deletion plus a large insertion. Clause-aware diffing can instead show it as a single "moved" operation, far more useful to a reviewer trying to understand what actually happened between drafts.

Semantic diffing

A further refinement compares meaning rather than exact wording, using the sentence-embedding techniques covered earlier in this course, to flag a clause as "materially changed" only when its meaning shifts, not its phrasing alone.

text
materially_changed(clause_old, clause_new) =
    cosine_similarity(embed(clause_old), embed(clause_new)) < threshold

This trades false negatives (missing a meaning-changing edit that happens to reuse similar words) against false positives (flagging a harmless rewording as material), and the right threshold is genuinely task- and domain-dependent, with no universally correct value.

Complexity

Character or word-level LCS diffing costs O(n * m) time and space, for documents of length n and m, using the classical dynamic-programming algorithm. SequenceMatcher's heuristic trades a small amount of optimality for meaningfully better average-case speed on real text.

Clause-level diffing first requires aligning clauses between versions, an assignment problem solvable with the same edit-distance machinery at the clause level, before word-level diffing runs within each matched pair, adding a cheap preprocessing pass ahead of the more expensive fine-grained diff.

Key references

  • Ratcliff, J. & Obershelp, D. (1988). Pattern Matching: The Gestalt Approach. Dr. Dobb's Journal — the algorithm behind Python's difflib.
  • Myers, E. (1986). An O(ND) Difference Algorithm and Its Variations. Algorithmica — the classical algorithm behind Unix diff and Git.
  • Hendrycks, D. et al. (2021). CUAD. arXiv:2103.06268 — clause-level structure this comparison approach commonly relies on.

Current state and open problems

Line and word-level diffing is a mature, well-understood technology, and most commercial contract-comparison tools handle it reliably as a baseline capability.

The open problem is semantic diffing at production quality: reliably distinguishing a harmless rewording from a materially different obligation, at scale, across the enormous variety of ways two lawyers can phrase the same underlying commitment. This remains an active area, closely related to the regulatory-change-monitoring problem covered next, where the same "did the meaning actually change" question reappears in a different setting.

What to learn next