Human-in-the-loop does not scale. Here is what does.
One reviewer approving every agent action works at ten actions a day and collapses at ten thousand. These controls preserve human accountability without making a human the bottleneck — or the rubber stamp.
Past a few dozen approvals a day, humans stop evaluating and start clicking. The control still exists on paper; it stopped existing in practice.
When everything needs approval, risk-tiering disappears and the dangerous action waits in the same queue as the trivial one.
Incidents produce action storms. HITL throughput is fixed — so the control fails precisely when it matters most.
Classify actions by consequence (read / write / irreversible / consequential). Autonomy for the bottom tiers, mandatory human decision for the top.
Deterministic rules that block, allow, or escalate every tool call in milliseconds — the same decision at action #1 and action #100,000.
Humans supervise streams and intervene by exception, backed by a kill-switch — accountability without per-action queues.
Every decision reviewable after the fact, which is what boards and regulators actually require.
For credit, trade, medical, and legal-effect decisions, human accountability is architecturally enforced (CP.10 HEAR Doctrine) — the human decides, the agent executes.