Intelligent Operations · Control design
Autonomous vs Human-in-the-Loop Operations
Nobody argues against keeping a human in the loop. The argument worth having is which decision that human is actually reviewing, and whether they can still say no on the four-hundredth one.
Key takeaways
- Autonomy is a property of an action, not of a workflow. Most real workflows contain both kinds, and forcing one setting onto all of them makes the workflow either useless or dangerous.
- Four properties decide the level: reversibility, value ceiling, verifiability, and who sees the failure first.
- A reviewer handling hundreds of items a day is a rubber stamp with a salary. If you cannot shrink the queue, you have not built a control — you have moved the risk into someone’s inbox.
- The most underused setting is not review-before-execute. It is execute after a veto window on reversible actions.
- Autonomy should be earned per action type on recorded evidence, and withdrawn automatically when the correction rate rises.
The short answer
Set autonomy per action, not per workflow. An action can run unattended when it is reversible, bounded in value, verifiable after the fact, and visible to you before it is visible to a customer. Anything that fails one of those tests needs a human decision with real stopping power — and anything that fails none of them should not be in a review queue at all, because a reviewer with a full queue approves everything.
Five levels, not two
The question is usually framed as a switch: autonomous or supervised. In practice there are five settings, and the two in the middle are where most operational work belongs.
Suggest
The agent drafts and a human executes. Useful for building confidence and little else — throughput is bounded by the human, and the human owns the outcome.
Draft and hold
The agent prepares the action completely — the payment, the email, the ledger entry — and it waits. The reviewer sees the exact thing that will happen, not a summary of it.
Execute with a veto window
The action is scheduled, announced, and executed after a delay unless someone stops it. Legitimate where the action is reversible or the delay is tolerable, and the most underused setting on this list.
Execute and notify
The action runs immediately; a human is told and can correct afterwards. Correct for high-volume reversible work where a wrong action costs a correction rather than an incident.
Unattended
No human in the path. Appropriate for reversible, bounded, well-instrumented actions, and for nothing else.
The distinction that matters
Levels 01 and 02 put a human before the action. Levels 03 and 04 put one after it with the power to undo. That is a different control, not a weaker one — and for reversible work it is usually the better trade, because it does not have to queue behind a person to be useful.
What decides the level
Four properties of the action, evaluated with its actual parameter values rather than its name. “Send email” is not one risk class: sending it to a colleague and sending it to a customer list are not the same action.
- 01ReversibilityCan this be undone by the same system, without asking anyone? A draft can. A posted journal entry can be reversed, but it leaves a trail. A wire transfer and a sent email cannot be undone at all.
- 02Value ceilingThe most expensive plausible version of this action, not the median one. Ceilings belong in policy, expressed in currency and count, per caller and per window.
- 03VerifiabilityWhether the result can be checked afterwards against something authoritative — a purchase order, a contract, a ledger balance. Unverifiable output needs the human in front, because there is nothing to catch it behind.
- 04Blast visibilityWho notices the mistake first. If the answer is a customer, a regulator or an auditor, the action does not qualify for unattended execution however well it usually goes.
| Action | Reversible | Verifiable against | Level it earns |
|---|---|---|---|
| Draft a reply for a human to send | Yes | Not needed | Unattended |
| Update a CRM field from a signed contract | Yes | The contract | Execute and notify |
| Post an invoice matched to a PO and receipt | With a trail | Three-way match | Veto window |
| Issue a refund above the tier ceiling | Partly | The order | Draft and hold |
| Send an external email to a customer list | No | Nothing | Draft and hold |
| Change a vendor’s bank details | No | Out-of-band contact only | Draft and hold |
Scroll the table horizontally on narrow screens.
The test people skip
Reversibility is a property of your systems, not of the action in the abstract. If nobody has exercised the undo path in the last quarter, treat the action as irreversible — an untested rollback is a plan, not a control.
How approval theatre creeps in
Approval theatre is the state where a control exists, is documented, is audited, and changes nothing, because the person operating it cannot evaluate what they are approving. It is not a people problem; it arrives by four predictable routes.
- Queue volume grows until median review time falls to a few seconds. The approval becomes a click, and the record of it becomes misleading evidence.
- The reviewer sees a summary the agent wrote rather than the action itself, so what is being approved is a description of the risk instead of the risk.
- Everything is routed for review because nobody wanted to argue about thresholds, which teaches the reviewer that approvals are routine.
- Timeouts auto-approve to keep work moving, converting the control into a delay.
- ›Median approval time under five seconds
- ›Approval rate above 99% with no corrections recorded
- ›One reviewer covering every action type
- ›Auto-approve on timeout
- ›The reviewer sees a natural-language summary
- ›A queue a reviewer can actually read each day
- ›A correction rate you can quote from data
- ›Routing by action type to whoever owns that system
- ›Fail closed on timeout, with an explicit expiry
- ›The reviewer sees the exact parameters that will execute
The uncomfortable implication is that adding review to more actions can reduce safety. Attention is a fixed budget, and spreading it across everything spends it on the routine cases.
Designing the handoff
When an action does need a person, the quality of the control is decided almost entirely by what the review request contains.
Review request — the minimum useful payload
action transfer.create
caller agent:ap-clerk (tenant 4471, run 8f21c9)
parameters amount: 48,200.00 USD
beneficiary: Northwind Supply Ltd
account: ****4417 [CHANGED 2 days ago]
references: INV-20871, PO-4402
why stopped value above tier ceiling (25,000) AND beneficiary account changed
already done three-way match: PASS duplicate check: PASS vendor master: FAIL
if approved executes immediately, irreversible
if ignored expires in 4h -> denied, workflow parks, owner notified
cost of delay invoice due in 6 days, early-payment discount lost after today- 01The exact action, not a paraphraseApproving a summary is approving the agent’s account of its own intention.
- 02Why this one stoppedA reviewer who cannot see which rule fired cannot tell a routine case from the case the rule was written for.
- 03What was already checkedAutomated evidence — matches, duplicate checks, policy results — tells the reviewer what is left for them to judge.
- 04What happens if they do nothingExpiry must be explicit and it must fail closed. Silence is not consent, and a timeout that approves is a control with an off switch.
- 05The cost of not actingReviewers systematically over-weight the risk of approving and under-weight the cost of the delay. Stating both makes the real trade visible.
A detail that pays for itself
Route by action type, not by seniority. The person who can judge a changed bank account is whoever owns the vendor master, and that is rarely the person with the highest approval limit.
Earning autonomy
Autonomy levels should move on evidence, in both directions, per action type. The mechanism is simple enough to write down as four rules.
Start at draft and hold
Every new action type begins with a human before execution, whatever the demo suggested.
Promote on a clean run
A meaningful streak of approvals with zero reviewer corrections promotes the action one level. Reversible actions may reach a veto window; irreversible ones stop below it.
Demote on correction
Any reviewer correction, or any reversal found after the fact, drops the action back a level immediately and automatically. The demotion is not a meeting.
Reset on change
A new model, a new tool version, a changed schema or a new integration resets the streak. The evidence was about the previous configuration.
Thresholds are yours, the shape is not
How long a streak has to be depends on your volume and on what a wrong action costs: a team doing 50 invoices a month cannot use the same number as one doing 50,000. What is not negotiable is that the promotion rule is written down, and that demotion happens without anyone having to agree to it.
The ceiling does not move. No amount of clean history makes an irreversible, externally visible, unverifiable action a candidate for unattended execution — that is not a confidence problem, and confidence is not the thing that would have to change.
What to measure
- ›Percentage of workflows automated
- ›Number of approvals processed
- ›Model accuracy on an evaluation set
- ›Hours saved, estimated
- ›Autonomous completion rate per action type
- ›Correction rate, and what reviewers changed
- ›Median and 95th-percentile approval latency
- ›Rubber-stamp rate: approvals granted in under five seconds
- ›Post-hoc reversals per thousand actions
Correction rate is the one that says whether the loop is doing anything at all. A review step with a correction rate near zero is either perfectly calibrated or purely decorative, and the way to tell them apart is to look at what reviewers changed on the occasions they did intervene.
The full measurement set for an agentic process, including the denominators that make these numbers comparable, is in measuring agentic process quality. Where the thresholds themselves come from is designing approval thresholds.
The product
Approvals, veto windows and durable audit around real business systems
BarzelOps runs governed workflows across HubSpot, Xero, Gmail, Google Drive and Slack with human-in-the-loop approvals and a durable audit trail. Action-level policy, step-up enforcement and the four deterministic outcomes live in BarzelVault.
Key terms
- Human-in-the-loop
- A control in which a person decides before the action executes.
- Human-on-the-loop
- A control in which the action executes and a person supervises, with the power to stop or reverse it.
- Veto window
- A deliberate delay between scheduling an action and executing it, during which a human can cancel.
- Approval theatre
- A review step that is documented and audited but cannot change outcomes, usually because queue volume exceeds reviewer attention.
- Autonomy ratchet
- A rule that promotes an action type on recorded evidence and demotes it automatically on correction.
- Blast visibility
- Whether a failed action is noticed first internally or by a customer, regulator or auditor.
Frequently asked questions
Should AI agents ever run without a human in the loop?
Yes — for actions that are reversible, bounded in value, verifiable against a system of record, and internally visible when they fail. That describes a large share of operational work. It does not describe payments, external communications, or changes to who gets paid.
What is the difference between human-in-the-loop and human-on-the-loop?
In the loop, the person decides before execution. On the loop, the action executes and the person supervises with the power to stop or reverse it. On-the-loop scales; in-the-loop does not.
How do you stop approvals becoming a rubber stamp?
Cap the queue, route by action type to whoever owns that system, show the exact parameters rather than a summary, fail closed on timeout, and measure the share of approvals granted in under five seconds.
When should an action be promoted to unattended execution?
After a recorded streak of clean approvals for that specific action type, with automatic demotion on any correction or reversal — and never for irreversible, externally visible, unverifiable actions.
Does a veto window count as human oversight?
It does if the window is long enough to act within, the notification reaches someone accountable, and cancelling actually works. It does not if the delay is a formality nobody watches.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.
- 01 · EU AI Act explorer EU AI Act — consolidated text and timeline ↗ Obligations determining required documentation and human oversight.
- 02 · NIST AI Risk Management Framework ↗ The govern / map / measure / manage structure these controls map onto.
- 03 · ISO ISO/IEC 42001 — AI management systems ↗ Certifiable management-system standard for AI governance programmes.
- 04 · OWASP GenAI Security Project OWASP GenAI LLM Top 10 (2026) ↗ Current consensus list of LLM application risks, including excessive agency.
- 05 · Reference definition Principle of least privilege ↗ The 1975 Saltzer and Schroeder formulation this all descends from.
- 06 · Google Google SRE Book — Service Level Objectives ↗ Where the SLI/SLO/error-budget vocabulary used here comes from.
Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.
Reference documentation
Want the specification rather than the argument?
Which steps BarzelOps gates by default, which it lets run, and the durable record behind both.
Free tier · 100 calls/day · on the MCPize marketplace
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.