AI Agent Governance · Design
Designing Approval Thresholds
Human-in-the-loop fails in a specific, predictable way: too many approvals, so reviewers stop reading. The design problem is not whether to have approvals — it is where the line goes.
Key takeaways
- The failure mode is approval fatigue, not under-approval. A queue nobody reads is a control that exists only on paper.
- Four axes decide it: reversibility, financial materiality, external visibility, and deviation from observed norm.
- Set the numbers from real traffic. Guessed thresholds are simultaneously too tight for routine work and too loose for the unusual case.
- Give the reviewer the decision, not the log. Show what will happen, what is unusual about it, and the two buttons — nothing else.
- Approval rate is the health metric. Track it per agent; a rising rate means policy drift or a genuinely changed workload, and both need attention.
The short answer
Require human approval where an action is irreversible, financially material, externally visible, or operating outside its observed norm — and nowhere else. Approval on everything trains reviewers to click through, which is worse than no approval because it manufactures an audit trail of unread consent. Set thresholds from two weeks of observed traffic, not from intuition, and measure the approval rate: above roughly 5% of calls, you are building rubber-stamping.
The failure nobody plans for
Teams introduce human-in-the-loop as the safety answer, route every consequential-looking action to a queue, and are pleased with the design. Six weeks later the queue has hundreds of items a day, one person approves in batches without opening them, and the organisation now has documented consent for actions no human actually assessed.
That is strictly worse than having no approval step. It consumes real time, it slows the workflow enough to invite bypass, and it produces evidence that a human agreed — which will be read literally in an incident review.
Say the quiet part
Every approval you add reduces the attention available for the ones that matter. Approval is a scarce resource being allocated, not a safety feature being layered on. Spend it where reversal is expensive.
The four axes
An action earns a human when it scores on at least one of these. Most actions score on none, and should execute.
Reversibility
Can this be undone, by whom, and at what cost? Sending an email is irreversible in practice. Updating a draft record is trivially reversible. This axis alone resolves most cases.
Financial materiality
Does money move, or does a commitment get made? Set a currency threshold and apply it to the parameter value, not the tool name — the same payment tool handles $12 and $120,000.
External visibility
Does a third party see the result? Customer emails, public posts, filings and partner API calls all leave your blast radius and your ability to quietly correct.
Deviation from norm
Is this outside what this agent normally does? Volume, hour, destination, record count and frequency all work as signals. This is the axis that catches a compromised session doing individually-permitted things.
| Action | Reversible | Material | External | Verdict |
|---|---|---|---|---|
| Read a customer record | n/a | No | No | Execute |
| Draft a reply for review | Yes | No | No | Execute |
| Update an internal status field | Yes | No | No | Execute |
| Send an email to a customer | No | No | Yes | Approve |
| Issue a refund under $100 | Partly | Low | Yes | Auto, with monitor |
| Issue a refund over $1,000 | Partly | Yes | Yes | Approve |
| Delete records in bulk | No | Maybe | No | Approve, always |
| Export 50,000 rows externally | No | No | Yes | Deny, not approve |
Scroll the table horizontally on narrow screens.
The last row matters. Some actions should not be approvable at all. Routing a plainly out-of-scope request to a human converts a clean denial into a social-engineering opportunity.
Setting the numbers
Thresholds picked in a meeting are wrong in both directions at once. Derive them.
- 01Run in observe-only for two weeksEverything permitted, everything recorded. You now have the actual distribution of tools, values, volumes and hours per agent.
- 02Read the percentiles, not the meanSet the deviation threshold near the 95th percentile of observed behaviour. The mean tells you nothing about the tail, and the tail is the risk.
- 03Set money thresholds from your existing delegation policyYour organisation already has spend authority limits for humans. Reuse them — inventing a separate scale for agents invites the argument that agents are special, which they are not.
- 04Target an approval rate under 5% of callsIf policy produces more, the boundary is in the wrong place. Tighten scope so routine work never enters the queue, rather than asking reviewers to absorb it.
- 05Review monthly for the first quarterApproval rate, denial rate and time-to-decision. All three move as agents change, and stale thresholds fail silently.
What the reviewer sees
Approval quality is a design problem. A reviewer given a JSON payload will approve it; a reviewer given a decision will make one.
- ›Raw tool call JSON
- ›The full prompt and conversation
- ›Every field, undifferentiated
- ›A generic “agent requests permission”
- ›Plain-language statement of what will happen
- ›The specific value that crossed the threshold
- ›What is unusual about it versus this agent’s norm
- ›Reverse cost, and a deadline for deciding
One more detail
Record the approver, the timestamp and the reason against the action itself — not in a ticket that references it. An auditor asking “who authorised this?” should get one answer from one place, which is the whole point of putting the approval in the call path. See how to audit AI agent actions.
Measuring whether it works
Four numbers, reviewed monthly. Each one has a specific failure it detects.
Approval rate
Approvals as a share of all calls. Rising means scope is too broad or the workload genuinely changed. Above 5% means you are manufacturing rubber-stamping.
Rejection rate within approvals
If reviewers reject almost nothing, the threshold is too low and the queue is theatre. A healthy queue has real rejections in it.
Time to decision
Long waits are where bypass pressure comes from. If the median is hours, someone will get approval delegated away entirely.
Denial rate
Blocked calls per agent. A spike is the earliest signal of an injection attempt or a misconfigured agent — and it needs an owner who reads it.
Report these per agent rather than in aggregate. One agent with a broken scope will hide inside a fleet-level average, and the fleet-level average is what most dashboards show.
Where approval sits relative to the other gates
Gate four is the approval decision — deliberately after scope and parameter bounds, so plainly out-of-scope requests are denied outright rather than escalated to a human.
Diagram of an AI agent action decision flow showing human approval as the fourth of five gates, after caller identity, tool scope and parameter value checks, and before trajectory analysis. Out-of-scope actions are denied at earlier gates rather than escalated for approval.
Original diagram by Real Biz Digital. Reuse it anywhere with a link back to this article.
The product
Approval as a state in the call path, not a ticket
BarzelVault holds approval thresholds, step-up challenges and the named-approver record against the action itself — so the authorisation and the act are one record.
Key terms
- Approval threshold
- The rule that determines whether an action executes automatically or waits for a named human decision.
- Approval fatigue
- The degradation of review quality caused by volume, resulting in consent recorded without assessment.
- Step-up
- Interrupting an in-flight action to require additional authorisation before it proceeds.
- Reverse cost
- The effort and expense required to undo an action, which is the primary input to whether it needs approval.
- Deviation threshold
- A bound set relative to an agent’s observed normal behaviour rather than an absolute value.
Frequently asked questions
Which AI agent actions need human approval?
Those that are irreversible, financially material, externally visible, or outside the agent’s observed normal behaviour. Actions meeting none of those should execute automatically, because every unnecessary approval reduces attention for the ones that matter.
What is approval fatigue?
The degradation of review quality caused by approval volume. Reviewers begin batch-approving without assessment, which produces an audit trail of consent that no human actually gave — worse than having no approval step.
What is a healthy approval rate?
Below roughly 5% of calls, with real rejections occurring inside that. A higher rate means the agent’s scope is too broad; the fix is tightening scope so routine work never enters the queue.
Should every risky action be approvable?
No. Some should be denied outright. Routing a plainly out-of-scope request to a human turns a clean refusal into a social-engineering opportunity and adds queue volume that degrades other decisions.
How do we set the thresholds?
Run in observe-only mode for two weeks, then set deviation thresholds near the 95th percentile of observed behaviour and reuse your existing human delegation-of-authority limits for money. Review monthly for the first quarter.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.
- 01 · NIST AI Risk Management Framework ↗ The govern / map / measure / manage structure these controls map onto.
- 02 · EU AI Act explorer EU AI Act — consolidated text and timeline ↗ Obligations determining required documentation and human oversight.
- 03 · OWASP GenAI Security Project OWASP GenAI LLM Top 10 (2026) ↗ Current consensus list of LLM application risks, including excessive agency.
- 04 · ISO ISO/IEC 42001 — AI management systems ↗ Certifiable management-system standard for AI governance programmes.
- 05 · COSO COSO Internal Control — Integrated Framework ↗ The control framework auditors map financial process evidence to.
Last reviewed 19 August 2026. External links open in a new tab; we do not control their content.
Reference documentation
Want the specification rather than the argument?
Where the gate actually sits in a running workflow, and the approval tools that resolve it.
Free tier · 100 calls/day · on the MCPize marketplace
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.