Tracking regulatory change
Regulatory change monitoring compares old and new rule text to flag what actually changed, and a purely word-overlap approach can mistake a harmless rewording for a real change.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Regulatory change monitoring checks whether a rule applying to your company actually changed. It checks meaning, not only wording.
A notice board at work lists the parking rules. You read it once, and mostly stop noticing it. If the wording changes six months later, you would not know, unless someone specifically flags it.
Companies have hundreds of regulations like that notice board, except missing an update can mean a real fine. Regulatory change monitoring is software reading that notice board constantly, on a company's behalf.
Why it exists
Regulators update rules constantly: a reporting deadline shortens, a new disclosure gets required, a threshold changes. A large company can be subject to thousands of rules, across many regulators and jurisdictions, all changing independently.
A compliance team cannot re-read every regulation every single day, checking for changes by memory. That does not scale past a handful of rules.
Automated monitoring compares each new version of a regulation's text against the version last seen. It flags exactly what is different. That is the same idea as redlining a contract, applied to law instead.
How it works
Old rule: "A firm must report a data breach to the
regulator within 72 hours."
New rule: "A firm must notify the regulator of any
data breach no later than 72 hours after
discovery."
Compare meaning, not just words:
Similarity score: 0.50 -> flagged for human review
A human checks it: same 72-hour obligation, reworded.
Not a real change. But the tool was right to ask.The tool did not decide on its own that this was safe to ignore. It flagged a borderline case for a person to check. That is the honest, correct behaviour for a tool this imperfect.
Where you have already seen it
- Compliance software. Products that alert a bank's compliance team when a relevant regulator publishes an update to a rule.
- RegTech news feeds. Automated summaries flagging "what changed" in a new regulatory publication, instead of the full, unchanged document.
- Internal policy tracking. Large companies tracking whether their internal policies still match current external regulation, after each regulatory update.
Remember this
- Regulatory change monitoring compares old and new rule text, and flags what looks different.
- Comparing meaning is harder than comparing exact words, and current tools get this wrong sometimes.
- A flagged, uncertain case going to a human reviewer is the safe, correct outcome, not a failure.
What to learn next
- Comparing contract versions — the same underlying comparison technique, applied to contracts instead of regulations.
- Legal documents as data — the structure this kind of monitoring relies on to compare rule to rule cleanly.
- BERTScore — a meaning-aware comparison technique that improves on the word-overlap approach shown here.
Developer — Code and libraries.
This compares old and new versions of three rules using word overlap, and shows exactly where that approach succeeds, and where it genuinely gets fooled.
Setup
pip install scikit-learnComparing old and new rule text
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
old_rules = [
"A firm must report a data breach to the regulator within 72 hours.",
"Customer funds must be held in a segregated account.",
"Annual audits are required for firms with over 50 employees.",
]
new_rules = [
"A firm must notify the regulator of any data breach no later than 72 hours after discovery.",
"Customer funds must be held in a segregated account with a licensed custodian.",
"Firms must publish a diversity report every year.",
]
vectorizer = TfidfVectorizer().fit(old_rules + new_rules)
sim_matrix = cosine_similarity(vectorizer.transform(old_rules), vectorizer.transform(new_rules))
for i, old_rule in enumerate(old_rules):
j = sim_matrix[i].argmax()
score = sim_matrix[i, j]
if score > 0.6:
verdict = "reworded, same obligation"
elif score > 0.25:
verdict = "obligation likely changed -- review"
else:
verdict = "no close match -- may be a new or removed rule"
print(f"[{verdict}] (similarity={score:.2f})")
print(f" old: {old_rule}")
print(f" new: {new_rules[j]}\n")[obligation likely changed -- review] (similarity=0.50) old: A firm must report a data breach to the regulator within 72 hours. new: A firm must notify the regulator of any data breach no later than 72 hours after discovery. [reworded, same obligation] (similarity=0.81) old: Customer funds must be held in a segregated account. new: Customer funds must be held in a segregated account with a licensed custodian. [no close match -- may be a new or removed rule] (similarity=0.09) old: Annual audits are required for firms with over 50 employees. new: Firms must publish a diversity report every year.
Look closely at the first result. Those two sentences describe the same 72-hour obligation, worded differently: "report... within" versus "notify... no later than". This word-overlap approach only scored them 0.50, landing them in the "review" bucket instead of "reworded". The tool is not wrong to flag it, but it is not confident either, and that honesty is the useful part.
Line by line
TfidfVectorizer().fit(old_rules + new_rules) builds one shared vocabulary across both old and new text, so both sets get compared in the same space, on equal footing.
The three-way threshold (0.6, 0.25) is a judgement call, not a law of nature. Different thresholds trade false alarms against missed changes differently, and the right balance depends on how costly each kind of mistake is for your specific use case.
sim_matrix[i].argmax() finds each old rule's single closest match among the new rules, assuming a roughly one-to-one relationship between old and new rule sets. A rule that was split into two, or two rules merged into one, breaks this assumption.
Common mistakes
Trusting a word-overlap score as a true measure of meaning. As shown above, this method can score a harmless rewording surprisingly low, purely because different words were chosen to say the same thing.
Setting the "no change" threshold too high. A too-strict threshold sends every rewording to a human reviewer, defeating the point of automating the easy cases and burying reviewers in low-value alerts.
Assuming every old rule has exactly one new-rule match. Regulators split rules apart and merge them together across revisions. A rigid one-to-one comparison misses both cases entirely.
Try it yourself
Replace TfidfVectorizer with an embedding-based similarity method, of the kind covered in the BERTScore lesson, and rerun the first comparison.
Check whether a meaning-aware method scores the 72-hour rewording noticeably higher than 0.50. If it does, that is direct, hands-on evidence for why production regulatory-monitoring tools tend to move past word overlap.
What to learn next
- BERTScore — the meaning-aware alternative this lesson's "try it yourself" points toward.
- Comparing contract versions — the closely related technique for a different kind of document.
- Evaluating a legal AI system — measuring how often a monitoring system like this gets the "changed or not" call right.
Researcher — Mathematics and papers.
Change classification as a three-way decision problem
Formally, given an old rule text r_old and its best-matching new rule text r_new, the task is a three-way classification:
class(r_old, r_new) in { unchanged/reworded, materially changed, no match }The developer block implements this with a similarity score and two fixed thresholds, the simplest possible instance of this framework. Production systems typically replace the fixed thresholds with a learned classifier, trained on a labelled dataset of rule-pairs where legal experts have already determined which category applies.
Why lexical similarity is a weak proxy for legal materiality
TF-IDF cosine similarity measures word overlap, weighted by how distinctive each word is across the corpus. It has no model of legal meaning, obligation structure, or which specific word changes matter versus which do not.
This produces two characteristic failure modes, both visible in the developer block. False positives: a synonym-heavy rewording, like "report" becoming "notify", scores lower than its true semantic similarity warrants, because the surface words differ more than the underlying obligation does. False negatives: two rules that share substantial vocabulary but differ in one legally load-bearing word, for example a numeric threshold or a modal verb like "may" versus "must", can score deceptively high, since one changed word contributes little to overall lexical overlap.
Toward materiality-aware comparison
Better approaches specifically target the second failure mode, since it is the more dangerous one: a system that misses a real obligation change is worse than one that over-flags harmless rewording. Techniques include:
- Entailment-based comparison: using the natural language inference techniques from earlier in this course, checking whether the old rule and new rule mutually entail each other, rather than relying on surface overlap.
- Numeric and modal-verb extraction: specifically parsing out thresholds, deadlines and obligation strength ("must" versus "may" versus "should") as structured fields, then comparing those fields directly rather than comparing raw text.
- LLM-based comparison: prompting a capable model to explain, in natural language, exactly what changed between two rule texts, then having a human or a second model verify that explanation, an application of the LLM-as-a-judge technique from earlier in this course.
Complexity
TF-IDF-based comparison, as in the developer block, is cheap: O(V) per document pair for vocabulary size V, computed once and reused across all comparisons via a single similarity matrix. Entailment-based and LLM-based comparison cost one model forward pass per rule pair, meaningfully more expensive at the scale of thousands of monitored regulations across multiple jurisdictions, updated continuously.
Key references
- Chalkidis, I. et al. (2022). LexGLUE: A Benchmark Dataset for Legal Language Understanding in English. arXiv:2110.00976
- Zheng, L. et al. (2021). When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset. arXiv:2104.08671
- Hendrycks, D. et al. (2021). CUAD. arXiv:2103.06268 — the clause-comparison infrastructure this task builds on.
Current state and open problems
Commercial RegTech platforms increasingly combine lexical, embedding-based and LLM-based comparison in a tiered pipeline. A cheap lexical or embedding filter narrows a large corpus of updates to a manageable set. A more expensive method then makes the final materiality call on that smaller set.
The open problem is the same one raised repeatedly across this section. Distinguishing rewording from real change requires genuine legal understanding, not only linguistic similarity. No current automated technique matches an experienced compliance lawyer's judgement reliably enough to remove the human reviewer entirely. The realistic goal production systems actually target is reducing reviewer workload, not eliminating review.
What to learn next
- Natural language inference — the entailment-checking technique that improves on raw word overlap here.
- Comparing contract versions — the closely related comparison task for contracts.
- Evaluating a legal AI system — measuring a monitoring system's real-world reliability rigorously.