Legal and Compliance AI

E-discovery and technology-assisted review

Technology-assisted review ranks a huge pile of documents by predicted relevance, so reviewers find the important ones early instead of reading everything in random order.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Technology-assisted review ranks a huge pile of documents by how likely each one is to matter. Reviewers read the important ones first.

A hiring manager gets ten thousand resumes for one job. Reading them in the order they arrived wastes time, since good and weak candidates are scattered randomly through the pile.

A quick first pass, sorting resumes from most to least promising, means the best candidates get seen early. TAR, technology-assisted review, does exactly this for lawsuits, ranking millions of documents by predicted relevance.

Why it exists

A lawsuit can involve millions of emails and files that need review before a court deadline. This process is called e-discovery. It means finding every document relevant to a legal case, buried inside everything a company has ever written.

Reading every single document in arrival order is far too slow, and far too expensive, for cases of this size. Someone has to be paid to read each one.

TAR trains a model on a small batch of documents a lawyer already reviewed and labelled. It then ranks everything else, so the most promising documents get read first. The ranking keeps improving as more documents get labelled.

How it works

  A lawyer labels a small starter batch:
    "Merger agreement draft"        -> RELEVANT
    "Office closed for holiday"     -> not relevant
    "Due diligence report attached" -> RELEVANT

  A model learns the pattern, then ranks 100,000
  unread documents by predicted relevance:

    0.87  "Final merger terms attached"
    0.81  "Indemnification clause redline"
    0.34  "Team lunch reminder"
    0.12  "Parking garage maintenance notice"

Reviewers start at the top of this ranked list, not at document number one. The most promising material surfaces first, instead of being buried somewhere in the middle of an unsorted pile.

Where you have already seen it

  • Lawsuits and regulatory investigations. Large law firms use TAR routinely for major litigation, where manual review of everything is not realistic.
  • Corporate compliance audits. Companies searching internal communications for policy violations, across years of email, use the same ranking approach.
  • Freedom of information requests. Government agencies sorting through large document sets to find what is actually responsive to a request.

Remember this

  • TAR ranks documents by predicted relevance, so reviewers see the most promising ones first.
  • It learns from a small batch of documents a human already labelled.
  • The goal is finding relevant documents faster, not replacing human review entirely.

What to learn next

Developer — Code and libraries.

This trains a tiny relevance classifier on eight labelled example documents, then ranks five new, unlabelled ones by predicted relevance.

Setup

bash
pip install scikit-learn

Ranking unreviewed documents by predicted relevance

tar_demo.py
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression

# A human reviewer already labelled these: 1 = responsive to the case, 0 = not.
docs = [
    ("Please approve the Q3 marketing budget by Friday.", 0),
    ("Attached is the signed merger agreement draft v3.", 1),
    ("Reminder: the office is closed on Monday for the holiday.", 0),
    ("Can we discuss the acquisition price before the board meeting?", 1),
    ("Lunch order for the team meeting is confirmed.", 0),
    ("The due diligence report on the target company is attached.", 1),
    ("Happy birthday! Cake in the break room at 3pm.", 0),
    ("Legal has flagged a risk in the merger indemnification clause.", 1),
]
texts, labels = zip(*docs)

vectorizer = TfidfVectorizer().fit(texts)
clf = LogisticRegression().fit(vectorizer.transform(texts), labels)

# A new, unreviewed batch. In a real case this could be 100,000 documents.
unseen = [
    "Draft indemnification language for the acquisition contract.",
    "Team standup notes from Tuesday.",
    "Updated valuation model for the target company.",
    "Parking garage will be repaved next week.",
    "Board wants final numbers on the merger before signing.",
]

scores = clf.predict_proba(vectorizer.transform(unseen))[:, 1]
ranked = sorted(zip(unseen, scores), key=lambda x: -x[1])

print("unreviewed documents, ranked by predicted relevance:\n")
for doc, score in ranked:
    print(f"  {score:.3f}  {doc}")
Output
unreviewed documents, ranked by predicted relevance:

  0.565  Board wants final numbers on the merger before signing.
  0.527  Draft indemnification language for the acquisition contract.
  0.511  Updated valuation model for the target company.
  0.496  Parking garage will be repaved next week.
  0.454  Team standup notes from Tuesday.

The three merger-related documents rank above the two unrelated ones. With only 8 training examples, the model's confidence is modest, all scores sit close to 0.5, but the ranking direction is right. Real TAR systems train on thousands of labelled seed documents, producing far more confident, more separated scores.

Line by line

TfidfVectorizer().fit(texts) learns which words are distinctive across this document set, down-weighting common words that appear everywhere and carry little signal.

predict_proba(...)[:, 1] returns the predicted probability of the "relevant" class, not a hard yes/no label. Ranking by this probability is what lets reviewers work top-down through the pile.

Sorting with key=lambda x: -x[1] puts the highest-scoring, most-likely-relevant document first, since Python's sort is ascending by default.

Common mistakes

Treating the raw score as a real probability. With this little training data, 0.565 does not mean "56.5% chance of being relevant" in any calibrated sense. It only reliably tells you the ranking is more relevant than a lower-scoring document.

Stopping review once early documents look unpromising. TAR ranks, it does not filter perfectly. A relevant document can still land lower in the ranking than it should, especially early in the process with few labelled examples.

Never updating the model as review continues. Real TAR workflows are iterative: reviewers label more documents as they go, and the model retrains on the growing labelled set, steadily improving the ranking.

Try it yourself

Add two more labelled training examples, one relevant and one not, that look similar to the two lowest-ranked unseen documents above. Retrain and rerank.

Watch how much the ranking shifts with only a few more labelled examples. This sensitivity to training set size is exactly why real TAR workflows treat labelling as an ongoing, iterative process, not a one-time setup step.

What to learn next

Researcher — Mathematics and papers.

The TAR workflow as active learning

TAR is a specific application of active learning: a human labels a small seed set, a classifier ranks the remainder, the human reviews and labels more documents from that ranking (often prioritising uncertain or high-scoring cases), and the classifier retrains, repeating until a stopping criterion is met.

text
repeat:
    train classifier on labelled set L
    rank unlabelled set U by predicted relevance
    human labels a batch from U, moves it into L
until stopping criterion met

Two dominant protocols exist in practice. TAR 1.0 uses a single, static seed set with limited iteration, closer to the simplified example in the developer block. TAR 2.0 uses continuous active learning (CAL), retraining after every batch. It has been the dominant approach since roughly 2015, generally needing less reviewer effort to reach a given recall target.

The stopping problem

Unlike ordinary classification, TAR has no natural test set. The entire unlabelled pool is the actual population of interest. The goal is a legally defensible statement like "we have found at least 80% of the relevant documents", not only "our held-out accuracy is high".

Standard stopping methods include:

  • Sample-based estimation: draw a random sample from the unreviewed pool, estimate remaining recall using this sample directly, and stop once the estimated recall crosses the target threshold, with a confidence interval attached.
  • Elusion testing: after review is declared complete, sample the documents that were never reviewed, and check the observed relevance rate there is acceptably low.

Grossman & Cormack (2011, 2014) established that well-implemented TAR workflows can match or exceed manual review's recall, at a fraction of the reviewer-hours. Their empirical work substantially shaped e-discovery case law, and has been cited in numerous US federal court rulings approving TAR as a defensible review method.

Unlike most machine learning deployments, a TAR workflow's output can be challenged in court, and its methodology may need to be disclosed to opposing counsel. This creates requirements beyond pure predictive performance: documented process (what training data, what stopping criterion), reproducibility, and statistically defensible recall estimates, rather than only a good-looking ranked list.

Complexity

Training cost per round is standard supervised learning cost, O(|L| * d) for a linear model over |L| labelled documents and d features, negligible relative to human review time even at large scale. Ranking the full unlabelled pool costs O(|U| * d), a single inference pass, which is why TAR scales to millions of documents where manual review cannot.

Key references

  • Grossman, M. & Cormack, G. (2011). Technology-Assisted Review in E-Discovery Can Be More Effective and More Efficient Than Exhaustive Manual Review. Richmond Journal of Law and Technology.
  • Cormack, G. & Grossman, M. (2014). Evaluation of Machine-Learning Protocols for Technology-Assisted Review in Electronic Discovery. SIGIR.
  • Da Silva Moore v. Publicis Groupe (2012) — the first US federal court decision explicitly approving TAR as a defensible discovery method.

Current state and open problems

TAR 2.0 continuous active learning is now well-established practice in large-scale e-discovery, with a substantial body of supportive case law and empirical validation behind it.

The open problem is generalising this defensibility standard beyond English-language, text-heavy document sets. Multilingual e-discovery, and discovery involving images, audio, or structured data rather than prose, does not yet have the same maturity of tooling or the same depth of case law establishing what "defensible" review looks like for those formats.

What to learn next