Agentic Frameworks

Human-in-the-loop agents

How to build agents that pause and ask a person before doing anything risky — approval gates, escalation, and the fatigue trap that defeats them.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. The honest part
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A human-in-the-loop agent is an agent that must get a person's yes before its risky actions.

Think of a driving school car. The learner holds the steering wheel and does the actual driving. But the instructor's seat has its own brake pedal.

The instructor does not press it every metre. It exists for the few moments that matter.

Why it exists

An agent — a model that takes actions on its own, see AI agents — has two awkward properties.

First, it is confident even when wrong. A model states a wrong refund amount in the same calm tone as a right one — the same trait behind hallucination.

Second, some actions cannot be taken back. A sent email is sent. A deleted file is gone. Money that left the account has left.

Put those together and the rule writes itself: let the agent do freely what is reversible, and put a person's yes in front of what is not. It is how a shop owner treats a new employee — count stock alone from day one, but the owner signs the cheques for the first six months.

How it works

 agent proposes an action
        |
        v
   risky? ──no──→ do it
        |
       yes
        v
 show a person the exact action
        |
   ┌────┴─────┐
 approve    reject
   |           |
   v           v
 do it    the "no" goes back to the agent,
          which rethinks and tries again

Four common patterns, in plain words:

  • Approve before act — the agent waits for a yes on risky steps. The driving instructor's brake.
  • Edit then approve — the person fixes the draft before releasing it. Correcting the address on a parcel before it ships.
  • Ask when unsure — the agent itself escalates when out of its depth, like a junior calling the senior doctor.
  • Review after — everything is logged and a person audits later. This one works only for mistakes that can be undone.

Where you have already seen it

Your UPI app is exactly this design. Checking your balance needs no PIN — reversible, harmless. Sending money demands the PIN every single time — irreversible, gated. And AI coding assistants ask "run this command?" before touching your files.

The honest part

The weakness of this design is not the machine. It is that people get tired. Show someone twenty approval boxes an hour and by the afternoon they click yes without reading — and the gate becomes theatre. Fewer, better questions beat many careless ones. This part is genuinely hard, and pretending otherwise is how these systems fail.

Remember this

  • Gate the irreversible; let the reversible flow.
  • A rejection is feedback to the agent, not the end of the task.
  • Too many questions create rubber-stamping — the gate that guards nothing.

What to learn next

Developer — Code and libraries.

Setup

bash
python3 --version   # standard library only — nothing to install

Framework-free, so the pattern is visible. The same shape sits inside every framework's "interrupt" or "human input" feature.

An approval gate in thirty lines

approval_gate.py
READ_ONLY = {"look_up_order"}            # harmless: run without asking
IRREVERSIBLE = {"send_email", "refund"}  # cannot be undone: need a yes first

# A scripted human, so this file prints the same thing on every run.
# In a real app this is a button in a UI, or Python's input().
DECISIONS = iter([False, True])

def human_approves(action):
    decision = next(DECISIONS)
    verdict = "approved" if decision else "REJECTED"
    print(f"ask human -> {action['tool']}{action['args']}: {verdict}")
    return decision

# Pretend the model proposed this plan. The second step has a wrong amount.
plan = [
    {"tool": "look_up_order", "args": ("OD-1002",)},
    {"tool": "refund", "args": ("OD-1002", 2400)},
    {"tool": "refund", "args": ("OD-1002", 799)},   # the corrected retry
]

for action in plan:
    if action["tool"] in IRREVERSIBLE and not human_approves(action):
        print("   blocked. The rejection goes back to the agent as feedback.")
        continue
    print(f"   ran {action['tool']}{action['args']}")
Output
   ran look_up_order('OD-1002',)
ask human -> refund('OD-1002', 2400): REJECTED
   blocked. The rejection goes back to the agent as feedback.
ask human -> refund('OD-1002', 799): approved
   ran refund('OD-1002', 799)

What the output demonstrates

The lookup ran silently. Read-only tools never bother the human. Classify tools by blast radius — what breaks if this runs wrongly — not by which feature they belong to.

The wrong refund was caught at the gate. The human saw the exact tool and the exact arguments, and rejected. In a real agent loop, that rejection message is appended to the conversation, and the model proposes the corrected 799.

The person sees args, not a vague intent. Approving "refund('OD-1002', 2400)" is meaningful. Approving "handle the customer's issue" is a blank cheque. Bind every approval to the exact arguments — and if the agent changes the arguments afterwards, that is a new approval.

The scripted DECISIONS iterator exists so this page's output is reproducible. Swap it for input("approve? [y/n] ") and the program gains a real human and loses its determinism — say so in your tests, and script the approver there too.

The part this demo hides: the pause

Between "ask" and "answer" in production sit minutes, hours, or a weekend. The process may restart before the yes arrives. So the pending action, the surrounding plan, and the conversation must be saved somewhere durable and reloaded on resume — which is agent memory doing safety work. As of 2026 this save-pause-resume machinery is a first-class feature in the major frameworks: LangGraph interrupts a graph at a checkpoint and resumes it later, and CrewAI and AutoGen ship equivalent human-input hooks. The concept is identical everywhere; check your framework's current docs for the call names.

Common mistakes

Gating everything. Twenty prompts an hour trains the human to click yes blind. Reserve the gate for the irreversible few, and batch related approvals into one review.

Gating by tool name when the risk is in the arguments. A 50-rupee refund and a 50,000-rupee refund are different risks behind one tool. Gate on thresholds: auto-approve under a limit, ask above it.

Losing state across the pause. Process restarts, plan gone, and the agent re-runs step one — including steps that already happened. Persist the plan and mark completed steps. Idempotent tools limit the damage — see designing good tools.

No audit trail. When something goes wrong you must be able to answer: what was proposed, who approved, what ran, in what order. Log all four at the gate, where they cannot be skipped.

Try it yourself

Add a threshold rule: refunds under 500 rupees run without asking, larger ones go to the gate. Rerun and predict the printout before you look. Then add a log list recording every proposal, verdict, and execution — you have built an audit trail in five lines.

What to learn next

Researcher — Mathematics and papers.

When to ask: a value-of-information view

Interrupting has a cost; so does an unchecked error. Gate an action when

$$ p_{\text{err}} \cdot C_{\text{err}} > C_{\text{ask}} $$

where $p_{\text{err}}$ is the probability this action is wrong, $C_{\text{err}}$ the cost if it executes wrongly (including reversal cost, which is infinite for the truly irreversible), and $C_{\text{ask}}$ the cost of interruption — human time plus latency plus the marginal fatigue that degrades every future review. The inequality explains both design rules from the developer block: irreversibility inflates $C_{\text{err}}$ (gate it), and frequent asking inflates effective $C_{\text{ask}}$ through fatigue (gate less, better). The practical difficulty is that $p_{\text{err}}$ requires calibrated self-assessment, and model confidence is poorly calibrated on exactly the out-of-distribution inputs where errors cluster.

The human factors literature got here first

  • Sheridan and Verplank (1978) defined a ten-level scale of automation, from "human does everything" to "machine acts and ignores the human". Approve-before-act sits near the middle; most agent products slide up the scale as trust accumulates.
  • Parasuraman and Riley (1997), Humans and Automation: Use, Misuse, Disuse, Abuse (Human Factors 39(2)), documented automation bias: operators over-trust automation and under-monitor it. Rubber-stamping approvals is this phenomenon, predicted decades before LLMs.
  • Orseau and Armstrong (2016), Safely Interruptible Agents (UAI), treat interruption at the learning level: an RL agent repeatedly interrupted can learn to avoid the interruption itself, and they derive conditions under which optimal behaviour is unaffected by interrupts. Today's LLM agents are not trained against their own approval gates, so the failure mode is mostly theoretical for this stack — worth knowing as agents begin to learn from their own deployment traces.

Empirical grounding

Yao et al. (2024), τ-bench (arxiv.org/abs/2406.12045), evaluate agents conversing with simulated users under domain policy constraints — the setting where escalation and confirmation actually matter — and find reliability under repeated trials (pass^k) far below single-run success. The gap between "worked once in the demo" and "safe to leave alone" is the quantitative justification for oversight, and measuring that gap is the subject of evaluating agents.

Regulatory pressure

The EU AI Act (Regulation 2024/1689, obligations phasing in through 2026-27) makes human oversight an explicit legal requirement for high-risk systems: Article 14 requires that such systems be designed so natural persons can effectively oversee them, including the ability to intervene or interrupt. Comparable expectations appear in sectoral guidance elsewhere. Oversight design is becoming compliance work, not only good engineering.

Open problems

Meaningful review at scale — an approval interface that keeps $C_{\text{ask}}$ low without collapsing into rubber-stamping — is an unsolved interface problem. Binding approvals cryptographically to exact arguments, so the approved action and the executed action cannot drift, is engineering that few frameworks enforce yet. And scalable oversight — supervising agents whose work volume exceeds available human attention — remains open at every level, from UI design to alignment research.

What to learn next