Scoping an ML Project

Designing for a human reviewer

Most deployed models do not replace people — they sort work so confident cases run automatically and uncertain ones reach a human, and that band is a design decision.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A human-in-the-loop system lets the model decide the easy cases and routes the uncertain ones to a person.

Airport security works this way. The scanner clears most bags without anyone touching them. A few bags look ambiguous, and those — only those — go to the officer at the table. Nobody expects the scanner to open bags, and nobody hand-searches every suitcase.

The scanner is not failing when it asks for help. Asking for help on the right bags is the design.

Why it exists

The all-or-nothing framing — "the model replaces the reviewer" — sets an impossible bar. A model wrong 5% of the time cannot be trusted with life-changing decisions alone. The same model, allowed to say "not sure, you check", becomes deployable today.

This works because model mistakes are not spread evenly. They pile up in the cases where the model itself is hesitant. Skim off the hesitant cases and the remaining automatic decisions are far more reliable than the average suggests.

How it works

The model outputs a confidence score. Two lines cut it into three zones:

confidence:  0 ──── 0.2 ────────── 0.8 ──── 1
             reject  │  ask a human  │  approve
              (auto) │  (the band)   │  (auto)

Widen the review band — the ask-a-human zone — and automated decisions get safer, but humans get more work. Narrow it and automation rises along with risk. That width is a dial the business owns, and it can move weekly without retraining anything.

Two details decide whether the design works in practice:

  • The human must see why the case was flagged, or reviews become guesswork.
  • Track whether humans start approving everything without reading. A rubber-stamp reviewer converts your careful design back into full automation, silently.

A real example you have seen

UPI apps approve most payments instantly. Occasionally one pauses — "confirm this is really you", or a call from the bank. That pause is a review band decision: the model's confidence dipped, so a human (you, or an agent) entered the loop.

Remember this

  • The model handles confident cases; the review band routes uncertain ones to people.
  • Band width is a business dial: automation rate on one side, safety on the other.
  • The reviewer's attention is a resource the design must protect, not assume.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7 and numpy 1.26.

A mediocre model becomes a deployable system

This claims model is wrong 16% of the time — nowhere near trustworthy alone. Watch the band change what it is worth.

review_band.py
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

# Insurance claims: approve, reject, or send to a human for review.
X, y = make_classification(n_samples=8000, n_features=10, n_informative=8,
                           flip_y=0.02, class_sep=1.0, random_state=3)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0)
p = LogisticRegression(max_iter=1000).fit(X_tr, y_tr).predict_proba(X_te)[:, 1]
pred = (p >= 0.5).astype(int)

print(f"no human at all: 100% automated, {(pred != y_te).mean():.1%} of decisions wrong\n")
print("review band     automated   errors in automated   claims to humans")
for low, high in [(0.35, 0.65), (0.2, 0.8), (0.05, 0.95)]:
    auto = (p <= low) | (p >= high)          # model is confident either way
    err_auto = (pred[auto] != y_te[auto]).mean()
    print(f"{low:.2f} to {high:.2f}    {auto.mean():9.0%}   {err_auto:19.1%}   {(~auto).sum():16d}")
Output
no human at all: 100% automated, 16.4% of decisions wrong

review band     automated   errors in automated   claims to humans
0.35 to 0.65          84%                 10.7%                330
0.20 to 0.80          61%                  8.1%                771
0.05 to 0.95          26%                  4.4%               1476

The walkthrough

Errors are concentrated where confidence is low. Full automation is wrong 16.4% of the time. Automate only the confident 26% and those decisions are wrong 4.4% of the time — a fourfold improvement, from the same model, with zero retraining.

Each row is a staffing plan. 771 claims to humans is roughly ten reviewers; 330 is four. The band is where ML capability and headcount budgets meet in one table. Bring this table to the meeting, not an AUC.

The trade never disappears. No band gives 100% automation and low error — with this model. A better model shifts the whole table upward; the band then re-splits the improvement between fewer humans and fewer errors.

Confidence is not truth. This works because the model's probability estimates are roughly honest. A badly calibrated model is confidently wrong, and confidently wrong cases sail through the band. Check calibration before trusting any band.

Common mistakes

Sending humans a score with no reasons. A reviewer shown "0.55, claim #8412" can only guess. Show the top factors behind the score — explaining a model to a stakeholder covers how.

Measuring the model but not the loop. The deployed system is model + humans. Track end-to-end error and review turnaround, or the loop degrades invisibly — reviewers agreeing with the model 99.8% of the time are a warning sign, not a success.

Feeding reviewed cases back as training labels without care. Humans mostly see the band, so their labels cover the uncertain region only. Retraining on band-only labels distorts the model everywhere else.

A band nobody re-tunes. Traffic shifts, fraudsters adapt, and last quarter's 0.2–0.8 band drifts off its intended workload. Re-derive the cut-offs from fresh data on a schedule.

Try it yourself

Add a cost line to each row: a wrong automated decision costs ₹2,000 and a human review costs ₹80. Find the band with the lowest total cost. Then double the review cost and watch the optimal band narrow.

What to learn next

Researcher — Mathematics and papers.

Rejection as a first-class action

Classification with a reject option adds an action $r$ with fixed cost $c_r$ to the label set. Under 0-1 loss, the Bayes-optimal rule is Chow's rule:

$$ \text{reject} \iff \max_k P(Y = k \mid X = x) < 1 - c_r $$

Where:

  • $c_r \in [0, \tfrac{1}{2})$ — the cost of deferring, in units of a misclassification.
  • $P(Y = k \mid X = x)$ — the true posterior; in practice, the model's calibrated estimate.

Chow (1970), On Optimum Recognition Error and Reject Tradeoff (IEEE Trans. Information Theory), also characterises the error-reject curve: error rate as a function of rejection rate, whose slope is governed by the posterior's distribution near the decision boundary. The developer table above is a three-point sample of exactly this curve.

From rejection to deferral

Rejection assumes the fallback is perfect. Learning to defer models the human as a second, imperfect classifier with their own error pattern, and trains the deferral rule jointly:

  • Madras, Pitassi, Zemel (2018), Predict Responsibly (NeurIPS) — defer where the human is likely to outperform the model, not where the model is uncertain.
  • Mozannar and Sontag (2020), Consistent Estimators for Learning to Defer (ICML) — a consistent surrogate loss for the joint objective.
  • The distinction matters when human and model errors are decorrelated: optimal deferral can send confident model cases to the human because the human is better on that region.

Failure modes of the loop, empirically

  • Automation bias: reviewers over-accept algorithmic suggestions; Green and Chen (2019), Disparate Interactions (FAT*), measure it in risk-assessment settings and find compliance varies by defendant demographics — the loop can add bias the model lacked.
  • Selective labels: outcomes observed only for one side of the human decision (Lakkaraju et al., 2017, KDD), biasing retraining and evaluation.
  • Distribution shift induced by the band: retraining on band-region labels is covariate-shifted by construction; importance weighting or dedicated exploration traffic corrects it.

Calibration dependency

Chow's rule consumes posteriors, so its guarantees degrade with calibration error. Modern networks are systematically overconfident (Guo et al., 2017, On Calibration of Modern Neural Networks, ICML); a band on uncalibrated scores rejects the wrong cases. Post-hoc calibration (temperature scaling, isotonic regression) is therefore a prerequisite for principled band placement, and expected calibration error belongs on the monitoring dashboard next to the automation rate.

What to learn next