5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

AI Agent Governance · Design

Designing Approval Thresholds

Human-in-the-loop fails in a specific, predictable way: too many approvals, so reviewers stop reading. The design problem is not whether to have approvals — it is where the line goes.

By Mark Alex, Founder Published 19 Aug 2026 9 min read

Key takeaways

  • The failure mode is approval fatigue, not under-approval. A queue nobody reads is a control that exists only on paper.
  • Four axes decide it: reversibility, financial materiality, external visibility, and deviation from observed norm.
  • Set the numbers from real traffic. Guessed thresholds are simultaneously too tight for routine work and too loose for the unusual case.
  • Give the reviewer the decision, not the log. Show what will happen, what is unusual about it, and the two buttons — nothing else.
  • Approval rate is the health metric. Track it per agent; a rising rate means policy drift or a genuinely changed workload, and both need attention.

The short answer

Require human approval where an action is irreversible, financially material, externally visible, or operating outside its observed norm — and nowhere else. Approval on everything trains reviewers to click through, which is worse than no approval because it manufactures an audit trail of unread consent. Set thresholds from two weeks of observed traffic, not from intuition, and measure the approval rate: above roughly 5% of calls, you are building rubber-stamping.

The failure nobody plans for

Teams introduce human-in-the-loop as the safety answer, route every consequential-looking action to a queue, and are pleased with the design. Six weeks later the queue has hundreds of items a day, one person approves in batches without opening them, and the organisation now has documented consent for actions no human actually assessed.

That is strictly worse than having no approval step. It consumes real time, it slows the workflow enough to invite bypass, and it produces evidence that a human agreed — which will be read literally in an incident review.

Say the quiet part

Every approval you add reduces the attention available for the ones that matter. Approval is a scarce resource being allocated, not a safety feature being layered on. Spend it where reversal is expensive.

The four axes

An action earns a human when it scores on at least one of these. Most actions score on none, and should execute.

Axis 01

Reversibility

Can this be undone, by whom, and at what cost? Sending an email is irreversible in practice. Updating a draft record is trivially reversible. This axis alone resolves most cases.

Axis 02

Financial materiality

Does money move, or does a commitment get made? Set a currency threshold and apply it to the parameter value, not the tool name — the same payment tool handles $12 and $120,000.

Axis 03

External visibility

Does a third party see the result? Customer emails, public posts, filings and partner API calls all leave your blast radius and your ability to quietly correct.

Axis 04

Deviation from norm

Is this outside what this agent normally does? Volume, hour, destination, record count and frequency all work as signals. This is the axis that catches a compromised session doing individually-permitted things.

ActionReversibleMaterialExternalVerdict
Read a customer recordn/aNoNoExecute
Draft a reply for reviewYesNoNoExecute
Update an internal status fieldYesNoNoExecute
Send an email to a customerNoNoYesApprove
Issue a refund under $100PartlyLowYesAuto, with monitor
Issue a refund over $1,000PartlyYesYesApprove
Delete records in bulkNoMaybeNoApprove, always
Export 50,000 rows externallyNoNoYesDeny, not approve

Scroll the table horizontally on narrow screens.

The last row matters. Some actions should not be approvable at all. Routing a plainly out-of-scope request to a human converts a clean denial into a social-engineering opportunity.

Setting the numbers

Thresholds picked in a meeting are wrong in both directions at once. Derive them.

  1. 01Run in observe-only for two weeksEverything permitted, everything recorded. You now have the actual distribution of tools, values, volumes and hours per agent.
  2. 02Read the percentiles, not the meanSet the deviation threshold near the 95th percentile of observed behaviour. The mean tells you nothing about the tail, and the tail is the risk.
  3. 03Set money thresholds from your existing delegation policyYour organisation already has spend authority limits for humans. Reuse them — inventing a separate scale for agents invites the argument that agents are special, which they are not.
  4. 04Target an approval rate under 5% of callsIf policy produces more, the boundary is in the wrong place. Tighten scope so routine work never enters the queue, rather than asking reviewers to absorb it.
  5. 05Review monthly for the first quarterApproval rate, denial rate and time-to-decision. All three move as agents change, and stale thresholds fail silently.

What the reviewer sees

Approval quality is a design problem. A reviewer given a JSON payload will approve it; a reviewer given a decision will make one.

Do not show
  • ›Raw tool call JSON
  • ›The full prompt and conversation
  • ›Every field, undifferentiated
  • ›A generic “agent requests permission”
Show
  • ›Plain-language statement of what will happen
  • ›The specific value that crossed the threshold
  • ›What is unusual about it versus this agent’s norm
  • ›Reverse cost, and a deadline for deciding

One more detail

Record the approver, the timestamp and the reason against the action itself — not in a ticket that references it. An auditor asking “who authorised this?” should get one answer from one place, which is the whole point of putting the approval in the call path. See how to audit AI agent actions.

Measuring whether it works

Four numbers, reviewed monthly. Each one has a specific failure it detects.

Approval rate

Approvals as a share of all calls. Rising means scope is too broad or the workload genuinely changed. Above 5% means you are manufacturing rubber-stamping.

Rejection rate within approvals

If reviewers reject almost nothing, the threshold is too low and the queue is theatre. A healthy queue has real rejections in it.

Time to decision

Long waits are where bypass pressure comes from. If the median is hours, someone will get approval delegated away entirely.

Denial rate

Blocked calls per agent. A spike is the earliest signal of an injection attempt or a misconfigured agent — and it needs an owner who reads it.

Report these per agent rather than in aggregate. One agent with a broken scope will hide inside a fleet-level average, and the fleet-level average is what most dashboards show.

Diagram · free to reuse with attribution

Where approval sits relative to the other gates

Gate four is the approval decision — deliberately after scope and parameter bounds, so plainly out-of-scope requests are denied outright rather than escalated to a human.

Original diagram by Real Biz Digital. Reuse it anywhere with a link back to this article.

The product

Approval as a state in the call path, not a ticket

BarzelVault holds approval thresholds, step-up challenges and the named-approver record against the action itself — so the authorisation and the act are one record.

Key terms

Approval threshold
The rule that determines whether an action executes automatically or waits for a named human decision.
Approval fatigue
The degradation of review quality caused by volume, resulting in consent recorded without assessment.
Step-up
Interrupting an in-flight action to require additional authorisation before it proceeds.
Reverse cost
The effort and expense required to undo an action, which is the primary input to whether it needs approval.
Deviation threshold
A bound set relative to an agent’s observed normal behaviour rather than an absolute value.

Frequently asked questions

Which AI agent actions need human approval?

Those that are irreversible, financially material, externally visible, or outside the agent’s observed normal behaviour. Actions meeting none of those should execute automatically, because every unnecessary approval reduces attention for the ones that matter.

What is approval fatigue?

The degradation of review quality caused by approval volume. Reviewers begin batch-approving without assessment, which produces an audit trail of consent that no human actually gave — worse than having no approval step.

What is a healthy approval rate?

Below roughly 5% of calls, with real rejections occurring inside that. A higher rate means the agent’s scope is too broad; the fix is tightening scope so routine work never enters the queue.

Should every risky action be approvable?

No. Some should be denied outright. Routing a plainly out-of-scope request to a human turns a clean refusal into a social-engineering opportunity and adds queue volume that degrades other decisions.

How do we set the thresholds?

Run in observe-only mode for two weeks, then set deviation thresholds near the 95th percentile of observed behaviour and reuse your existing human delegation-of-authority limits for money. Review monthly for the first quarter.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.

  1. 01 · NIST AI Risk Management Framework ↗ The govern / map / measure / manage structure these controls map onto.
  2. 02 · EU AI Act explorer EU AI Act — consolidated text and timeline ↗ Obligations determining required documentation and human oversight.
  3. 03 · OWASP GenAI Security Project OWASP GenAI LLM Top 10 (2026) ↗ Current consensus list of LLM application risks, including excessive agency.
  4. 04 · ISO ISO/IEC 42001 — AI management systems ↗ Certifiable management-system standard for AI governance programmes.
  5. 05 · COSO COSO Internal Control — Integrated Framework ↗ The control framework auditors map financial process evidence to.

Last reviewed 19 August 2026. External links open in a new tab; we do not control their content.

Reference documentation

Want the specification rather than the argument?

Where the gate actually sits in a running workflow, and the approval tools that resolve it.

Free tier · 100 calls/day · on the MCPize marketplace

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.