Practical guide · controls
How to Control What an AI Agent Is Allowed to Do
Most guidance on this is architectural. This one is practical: five levers, three rules to write first, and what to do in the week you have rather than the quarter you do not.
The short answer
To control what an AI agent may do, use five levers in order: narrow the credential its tools hold, entitle it to specific tools rather than whole servers, bound the parameter values those tools accept, require human approval above a consequence threshold, and cap its call and spend rate per identity. Each is enforced outside the model. The first lever is the only one that shrinks the worst case rather than reducing the chance of reaching it. Three rules to write on Monday, and the test that distinguishes a control from a preference.
Key takeaways
- 01Start with the credential. Everything else is cheaper against a smaller worst case.
- 02Entitle per tool, never per server. A server grant hands over the mutating tools with the read-only ones.
- 03Your first three rules should be a ceiling, an allow-list and a rate cap. That is a week, and it closes most of the gap.
- 04Reject out-of-range values — do not clamp them. The attempt is the signal you want recorded.
- 05A control you can test by attempting the thing it should refuse. If you cannot write that test, it is a preference.
- 06Nothing here depends on the model behaving. That is the whole point.
Five levers, in order
The order is not arbitrary. Each lever makes the next one cheaper, and doing them out of sequence mostly wastes effort — there is little point tuning an approval threshold on a tool whose credential should have been narrowed first.
Narrow the credential
Establish what the tool’s credential can do on the backing system at its widest — not what the tool intends to use it for. Then reduce it. A read-only integration that holds a read-write key is one prompt away from being a write tool.
This is the only lever that makes the worst case smaller rather than less likely, and it is the one that takes calendar time rather than engineering time, because it needs the backing system’s owner to agree. Start it first for that reason alone.
Entitle per tool
Grant the specific tools the agent needs, not the servers they live on. Default entitlement is nobody. A server-level grant includes every mutating tool on that server, plus every tool added to it in future.
Filter discovery by identity too, so tools the agent is not entitled to never appear in its list. A tool the model never learns exists cannot be talked into attempting it.
Bound the parameters
For every mutating tool: a maximum, an allow-list, a row cap, a date range, an enumeration. This is the largest gap in most estates, because permitting a tool feels like it should be enough and it is not. A transfer tool with no ceiling is a different tool at 10 and at 10 million.
Reject out-of-range values rather than clamping them. A clamp silently rewrites the request and deletes the evidence that something tried to exceed the bound.
Require approval above a consequence threshold
For the small set of irreversible or externally-visible actions — usually under a dozen tools — hold above a threshold and show a human the tool, the values and the caller.
Trigger on consequence, not category. “All payments” generates volume that destroys review quality; “payments above £5,000 to a new destination” generates a handful.
Cap the rate, per identity
Calls per minute and spend per day, per agent. This bounds the damage rate of everything above it, and it catches the most common real incident: a retry loop making four hundred calls in thirty seconds.
Per identity, never per server. A per-server limit cannot tell you which agent consumed it, and one runaway agent takes everyone else’s headroom with it.
Your first three rules
A week of work, and between them they close most of the practical gap. Write these before anything sophisticated.
- A ceiling on the highest-magnitude mutating tool
- Find the tool whose parameters permit the largest single effect — usually a transfer, a bulk update or a delete. Give it a maximum that reflects what the agent is actually for. If the schema permits an unbounded amount, the tool is currently as dangerous as its maximum, which is to say unboundedly.
- An allow-list on the most externally-visible destination
- Whatever the agent can send to, pay or publish: constrain the recipients to a set established outside the agent’s control. This is the single rule that defeats the whole injection class, because an attacker cannot direct the outcome anywhere useful.
- A per-identity rate and spend cap
- Pick a number about five times your observed peak, then tighten it once you have a fortnight of data. An imperfect cap in place on Monday beats a perfect one specified in a document.
A note on the second rule, because it is the one people defer as too disruptive. The allow-list does not have to be static or manually curated — for payments it can be the customer’s own registered payment methods, for email the addresses already on the account record. The constraint is that it is established through a path the agent does not control, not that a human types it.
Control or preference?
One test, and it settles most arguments. Can you verify it by attempting the thing it should refuse, and see the refusal in a log?
If yes, it is a control. If the answer involves how the model has been instructed, how unlikely the case is, or what the agent has been trained to avoid, it is a preference. Preferences are worth having and they reduce how often things go wrong. They are not what you point at when somebody asks what would have stopped it.
| Mechanism | Control or preference? | Why |
|---|---|---|
| Credential scoped to read-only | Control | Attempt a write with the credential directly; the backing system refuses |
| Parameter ceiling at the enforcement point | Control | Submit one increment over; the call is rejected and the attempt is logged |
| Tool absent from the entitlement list | Control | Request it anyway; refused, and it never appeared in the tool list |
| System prompt saying not to exceed £5,000 | Preference | Cannot be tested by attempting it — the outcome depends on wording and context |
| Model fine-tuned to be cautious | Preference | Reduces frequency, provides no bound |
| A second model reviewing the action | Preference | Persuadable by the same content, and non-deterministic |
| SIEM alert on completed payments | Neither — detection | Tells you afterwards; useful, and not a bound |
The distinction is not academic. When an incident review asks what prevented a wider outcome, only the first three rows are answers. The others are contributing factors.
What to do in week one
Assumes one agent, some tools, and no existing enforcement point. Five days, and each day’s output is independently useful.
List what it can actually reach
Call tools/list for every server the agent is configured against and write down every tool, flagging the mutating ones. Do not use the documentation. Most teams find more than they expected.
The output is a list, not a system. That is fine — every subsequent day operates on this list.
Find the widest credential
For each server, establish what its credential permits at its widest on the backing system. Sort by blast radius. Start the conversation to narrow the worst one; it will not finish this week and that is expected.
Cut the tool list
Remove entitlement to everything the agent has not used and does not need. If you have no usage data, ask what the agent is for and remove anything outside that. Over-removal is recoverable in minutes; over-permission is not.
Write the three rules
Ceiling, allow-list, rate cap, at whatever enforcement point you have. If you have none, this is the day you find out — and putting one agent behind one inline point is a day’s work, not a project.
Try to break it
Submit a value over the ceiling. Send to a destination off the allow-list. Burst past the rate cap. Plant an instruction in a document the agent reads telling it to use a tool it is not entitled to. Every one should be refused and every refusal should be visible.
This is the day that tells you whether the week was real. Run the last test with someone else watching.
Built on this thinking
Levers three, four and five are one product
BarzelVault holds parameter bounds, consequence-based approval and per-identity ceilings as a pre-execution decision point with four deterministic outcomes — so the three rules you write on Thursday are enforced outside the model and every refusal is recorded. Barzel Central Gateway supplies identity and per-tool entitlement in front of it.
Five mistakes worth naming
Starting with approval workflows
The slowest, most expensive control, reached for first because it feels responsible. Narrow the credential instead — it costs nothing at runtime and eliminates cases entirely.
Server-level entitlement
One yes covering every tool the server has now and every tool it gains later.
Clamping instead of rejecting
Silently rewrites the request and destroys the signal that something tried to exceed a bound.
Rate limits per server
Cannot attribute a runaway to an agent, and lets one agent consume everyone’s headroom.
Treating the prompt as a control
No wording reliably beats an instruction the model finds in content it has just read.
What these levers do not give you
All five bound what an agent can do. None of them make what it does correct. An agent operating entirely within a narrow credential, an approved tool list, sensible ceilings and a reasonable rate can still issue the wrong refund to the right customer for a plausible-sounding reason. That is a process-quality problem and it needs measurement rather than enforcement.
There is also a real cost to over-tightening. Bounds set too narrowly produce refusals on legitimate work, and the pressure that follows is towards blanket widening rather than careful adjustment. Set a bound you can defend, watch the refusals for a fortnight, and tune from evidence — a rule nobody can justify gets removed entirely rather than corrected.
And this is a starting sequence, not a finished posture. It covers one agent. An estate needs the same controls applied consistently across many, which is a different problem — consistency at scale is what an enforcement point exists for, and it is where per-server configuration stops working.
Frequently asked questions
How do you control what an AI agent is allowed to do?
Five levers in order: narrow the credential its tools hold, entitle it to specific tools rather than whole servers, bound the parameter values those tools accept, require human approval above a consequence threshold, and cap call and spend rate per identity. Each is enforced outside the model.
Which control should you implement first?
Narrowing the credential. It is the only lever that makes the worst case smaller rather than reducing the probability of reaching it, and every other control is cheaper to get right against a smaller worst case. It also takes calendar time because it needs the backing system owner to agree.
What are the first three rules to write?
A ceiling on the highest-magnitude mutating tool, an allow-list on the most externally-visible destination, and a per-identity rate and spend cap. Roughly a week of work, and between them they close most of the practical gap.
How do you tell a control from a preference?
Ask whether you can verify it by attempting the thing it should refuse and seeing the refusal in a log. If the answer depends on how the model was instructed or how unlikely the case is, it is a preference — worth having, but not what stopped anything.
Should out-of-range parameter values be rejected or clamped?
Rejected. A clamp silently rewrites the request and deletes the evidence that something tried to exceed the bound, which is exactly the signal you want recorded — it tells you either the policy is wrong or the agent is being steered.
Why entitle per tool rather than per server?
Because a server-level grant hands over every mutating tool alongside the read-only ones, and it silently extends to any tool added to that server later. Tool-level grants are the minimum granularity at which approval means something specific.
Is a system prompt a way to control agent actions?
No. Instructions in a system prompt compete on equal terms with instructions the model finds in content it has just read, and no wording reliably wins. Prompt text reduces frequency; it does not provide a bound.
What should rate limits be set on?
Per identity, never per server. A per-server limit cannot attribute a runaway to a specific agent, and one agent in a retry loop consumes every other agent’s headroom. Start at about five times observed peak and tighten after a fortnight of data.
What can you achieve in one week?
Enumerate what the agent can actually reach from tools/list rather than documentation; identify the widest credential and start narrowing it; cut entitlement to what it needs; write the ceiling, allow-list and rate cap; then spend Friday trying to break all of it.
Do these controls make an agent’s decisions correct?
No. They bound what a wrong decision can do. An agent inside a narrow credential, an approved tool list and sensible ceilings can still issue the wrong refund to the right customer — which is a process-quality problem needing measurement rather than enforcement.
Glossary
- Control lever
- A mechanism that bounds what an agent can do, enforced outside the model so it does not depend on the agent’s reasoning.
- Credential scope
- The full set of actions a tool’s credential permits on its backing system, which defines the worst case for everything that tool can be induced to do.
- Parameter bound
- A constraint on the values a tool will accept — a ceiling, an allow-list, a row cap — evaluated at call time.
- Consequence threshold
- The parameter value at which an otherwise-routine action becomes one requiring a human decision.
- Control versus preference
- A control can be verified by attempting the action it should refuse; a preference merely makes the action less likely.
Sources and further reading
The control families here restate established security practice, cited below. The five-lever ordering, the first-three-rules recommendation and the control-versus-preference test are our own operating judgement.
- 01 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
- 02 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
- 03 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
- 04 · NISTNIST SP 800-207 — Zero Trust Architecture ↗The policy decision point / policy enforcement point split this architecture borrows directly.
- 05 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 06 · OWASPOWASP Application Security Verification Standard ↗Input-validation, authorization and logging requirements restated here in MCP terms.
- 07 · NISTNIST AI Risk Management Framework ↗Govern-map-measure-manage; the vocabulary most enterprise AI risk programmes are written against.
- 08 · ISOISO/IEC 42001 — AI management systems ↗The management-system standard auditors increasingly map AI governance evidence against.
Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). How to Control What an AI Agent Is Allowed to Do. Real Biz Digital. https://realbizdigital.net/insights/control-ai-agent-actions/
Try the mechanics on a live server
To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
BarzelVault is the pre-execution decision point, sold as a running product
Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Dev | Free | 10,000 policy decisions/mo · 9 tools, 4 outcomes, hash-chained audit | A first regulated workflow: one agent, one high-consequence system |
| Team | $199/mo | 75,000 decisions/mo · approval workflow, spend and action limits | Several agents acting on money, records or customer-visible systems |
| Business | $799/mo | 750,000 decisions/mo · exact HTTPS execution, credential isolation, emergency controls | Enterprise-wide pre-execution enforcement with audit obligations |
| Enterprise | $3,999/mo | 5,000,000 decisions/mo · everything in Business, scaled | Group-wide rollout across many teams and systems |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.