What should an AI approval threshold be based on?
Value is the obvious axis and the least sufficient on its own.
Value
Set it from data, not intuition. Take the distribution of action values over the previous twelve months and place the threshold so that a few percent of actions — representing a large share of total value — fall above it. That gives you a handful of decisions a day covering most of the money. A threshold set by feel is typically either so low it blocks normal work and gets raised to irrelevance within a month, or so high nothing reaches it.
Counterparty
New counterparties, counterparties whose bank details changed recently, entities outside the usual jurisdictions. This dimension catches what value misses, because the amounts involved look entirely ordinary. Most fraud that gets through automated controls gets through here.
Operation type
Reversals, corrections, credit notes, deletions, permission changes, anything issued outside the normal channel. These deserve a lower threshold precisely because they are rare — rare operations are the least-tested paths in any automation, and the least familiar to whoever reviews them.
Behaviour
The most valuable and the most often omitted. Not "this action is large" but "this agent has taken forty actions in ten minutes where it normally takes four an hour". Rate anomalies surface a malformed input or a runaway loop dozens of actions before a value threshold would, because each individual action looks fine.
Three response levels
| Level | Behaviour | Applies to |
|---|---|---|
| Pass and record | Executes; full entry in the audit trail | The large majority of actions |
| Pass and flag | Executes; notification for review | Deviations without high risk — a new counterparty at low value |
| Hold for approval | Does not execute until a person decides | Above value threshold, reversals above a floor, rate anomalies |
The middle level is what makes the design survive contact with production. Without it, every deviation routes to the hold queue; the queue grows; approvers start confirming in batches to clear it; and the control stops working while continuing to exist. That failure mode is worse than having no control, because it manufactures an audit trail of approvals that did not involve judgement — which is exactly what a regulator will probe.
What does an approver actually need to see?
This is the part that separates a defensible control from a decorative one, and it is almost always underbuilt.
An approval prompt that shows a tool name and a JSON blob does not produce an informed decision. The approver cannot evaluate it, so they approve on trust — and the record shows an approval that, examined later, nobody can defend. The surface has to render what the action will actually do, in the vocabulary of the business rather than the system: not invoke: post_credit_note but "issue a €14,200 credit note to Meridian Supplies against invoice INV-2291".
Record, at minimum:
- Who approved — a named person, not a role or a service account.
- What they were shown — the rendered content, or a hash of it. If the approver saw only a total and the problem was in the line items, that has to be establishable.
- When, precisely enough to bind the decision to the action.
- Which policy and version triggered the hold. Policies change; the evidence must point to the one in force at the time.
- The decision and any stated reason, especially on refusal.
The second item is the one that matters in a dispute. "Approved at 14:32" says nothing about whether the approval was meaningful. "Approved at 14:32, having been shown these five fields" is evidence.
Why is approving AI actions now a legal question?
Under UK GDPR Article 22A, whether a decision counts as solely automated turns on meaningful human involvement — and a decision with meaningful involvement falls outside the regime entirely. Batch approval, an approver who cannot see enough to disagree, and involvement with no recorded basis are the three patterns least likely to qualify. See the full DUAA analysis.
Colorado's replacement statute requires meaningful human review in similar terms from January 2027, California's ADMT rules grant a right to appeal, and Denmark's NIS 2-loven § 7 places approval of security measures on the management body personally. Different vocabularies, same underlying demand: show that a person actually decided.
The threshold must live outside the agent
If the model decides whether its own action needs approval, there is no control — and anything that can influence the model's context can influence the threshold, including content the agent reads. The determination has to happen in a layer the agent passes through and cannot reconfigure.
This is also the practical limit of the Model Context Protocol's MRTR mechanism, introduced in the 2026-07-28 specification. MRTR gives a server a clean way to pause and ask a human. It does not decide whether to ask, what to show, or how to evidence the answer. Those three remain the implementer's problem — see the MCP security analysis.
Operating it
- Thresholds are configuration with change history, not constants in code. If changing one requires a deployment, nobody will, and it will be stale within a quarter.
- Build a documented emergency path. A control with no override gets overridden informally — switched off during month-end close. A designed bypass with elevated logging is safer than an undocumented one.
- Review the hold rate monthly. A rising rate means either the threshold is wrong or the agent's behaviour has drifted. Both need investigating, and the second is the one you want to catch early.
- Segment by team. A threshold that makes sense for a sales operation is meaningless for a department issuing three high-value invoices a month.
Frequently asked questions
Where should the threshold be set?
From twelve months of action values, so a few percent of actions covering most of the value fall above it.
Why is a middle response level necessary?
Without it the hold queue grows until approvers clear it in batches, and the control stops functioning while appearing to exist.
What must be recorded?
Who approved, what they were shown, when, which policy version triggered the hold, and the decision with any reason.
Can the agent decide whether approval is needed?
No. That places the threshold inside the model's context, where a prompt can move it.
Does MRTR solve this?
It provides the mechanism to ask. Deciding when to ask, what to show and how to evidence it remains yours.
Related
- AI agent governance — the complete guide
- How to scope AI agent permissions
- How to build an audit trail for AI agents
- Human-in-the-loop vs autonomous agents
BarzelVault evaluates thresholds across all four dimensions before execution, renders the approval surface in business terms, and records what the approver was shown alongside the policy version in force.
In practice
Permission before the action. Evidence after it.
The duties on this page attach to the moment an automated system acts: who permitted it, on which data, under which policy version, and what a person saw before approving. Barzel enforces that decision before execution and writes the record an auditor, a regulator or a data subject can be shown.
430 days leftEU AI Act high-risk obligations (Annex III) apply from 2 December 2027
BarzelVault
The AI action firewall: decide what an agent may do before it does it.
- Approval thresholds and policy checks enforced before execution; human approvals that expire and escalate.
- Cryptographically signed audit receipts: trigger, inputs, policy version, approver, outcome.
- Credential isolation, spend and action limits, and an emergency kill switch.
Free tier: 10,000 calls a monthPaid plans from $199 a monthLive on MCPize
Barzel Central Gateway
The AI governance control plane: one inventory and one policy layer across every MCP server and agent.
- Registers and synchronises every tool; enforces identity, policy, region, cost and health per tool.
- Identity mapping through OIDC, Entra ID, Okta, SAML and SPIFFE, with credential brokerage.
- Trace and SIEM export (W3C trace context, OTLP) for the security team and the regulator.
Free tier: 1,000 calls a monthPaid plans from $10 a monthLive on MCPize
Enterprise: written quote by email within two business days. No sales call.
Sources
- Data (Use and Access) Act 2025, s.80; UK GDPR Articles 22A–22C.
- Colorado SB 26-189 (effective 1 January 2027); CPPA ADMT regulations.
- LOV nr 434 af 06/05/2025 (NIS 2-loven), § 7.
- Model Context Protocol, 2026-07-28 specification, SEP-2322 (MRTR).
- FINRA, 2026 Annual Regulatory Oversight Report.
This article is for information and does not constitute legal advice.