Pillar guide · enterprise
AI Agent Security: The Enterprise Guide
Securing a chatbot is a content problem. Securing an agent is an authority problem — the system can now move money, delete records and send messages, and it decides to do so at runtime, from text it has just read.
The short answer
AI agent security is the practice of bounding what an autonomous system is permitted to do, rather than trying to control what it decides. It differs from model security because the risk is action, not output: the same reasoning that makes an agent useful makes it persuadable, so enforcement has to sit outside the model. Seven controls hold — scoped credentials, per-agent identity, tool entitlement, parameter bounds, pre-execution approval, per-identity rate limits and tamper-evident evidence — and four widely used ones do not. Plus a 90-day sequence and a test for each control that expects refusal.
What actually changed
Two years of AI security attention went to model behaviour: jailbreaks, harmful output, hallucination. Reasonable, when the worst case was a bad answer on a screen. The worst case moved.
An agent with tools does not produce text for a human to act on; it acts. It reads a document, decides a refund is warranted, and issues one. It reads a ticket, decides an account should be unlocked, and unlocks it. The decision is made at runtime from content that arrived at runtime, and much of that content originates outside your control.
This makes the security problem an authority problem rather than a content one. The question is no longer whether the model can be convinced of something wrong — assume it can — but what it is able to do while convinced. Every control that works follows from taking that assumption seriously.
The corollary is uncomfortable for teams invested in prompt engineering: no wording reliably prevents a model from following an instruction it finds in retrieved content. Instructions in a system prompt and instructions in a document compete on the same footing, and the model has no reliable way to tell which one has authority. That is not a defect to be patched; it is what a language model is.
Six threats, in order of how often they cause real damage
Ordered by observed frequency rather than by drama. The first two account for most of what we have seen go wrong, and neither is an attack.
Excessive agency
The most common by a wide margin and the least discussed. An agent holds a credential broader than its task, a tool with no parameter ceiling, or permission to act in a system it only needed to read. Nothing adversarial required — one wrong decision is simply disproportionate.
This is the threat that scoping fixes rather than detection. Every other control on this page is cheaper once the worst case is smaller, which is why credential scope comes first in the sequence.
Runaway execution
An agent in a retry loop calling the same tool ten thousand times. Not an attack, and the largest source of surprise invoices and rate-limit incidents in the estates we have seen. Indistinguishable from abuse without a per-identity ceiling.
Indirect prompt injection
Instructions hidden in content the agent reads — a support ticket, a web page, a document, a calendar invite. The agent follows them because it cannot distinguish data from instruction. Anatomy of a real attack.
Prevention at the model layer is not currently possible. Containment is: the injected instruction has to fail at an enforcement point that does not care why the call was made.
Tool and supply-chain manipulation
Tool descriptions are text your model trusts, so a poisoned or shadowed tool description is an instruction injection with a package manager attached. Third-party servers that auto-update change what your model can do without review.
Pin versions and diff tool descriptions on upgrade — see third-party server intake.
Data exfiltration through legitimate tools
No exploit needed: an agent entitled to read customer records and entitled to send email can combine them. Each capability was approved; the composition never was.
Composition risk is invisible to per-tool review, which is why entitlement has to be examined as a set rather than as a list.
Confused deputy
The agent acts with its own service-account authority rather than the requesting human’s, so a low-privilege user’s request executes at high privilege. Old problem, new protocol in the middle.
Solved by preserving caller authority across the hop — authentication patterns compares the options.
Seven controls that hold
Ordered by leverage. The first three change what is possible; the next three change what is permitted and recorded; the last proves the rest.
The common property: every one of them is enforced outside the model. None depends on the agent choosing correctly, and each one is testable by attempting the thing it should refuse.
| Control | What it stops | Where it lives |
|---|---|---|
| Scoped credentials | Excessive agency, at the root. The only control that shrinks the worst case rather than reducing the chance of reaching it. | The backing system’s permission model |
| Per-agent identity | Confused deputy, unattributable actions, un-revocable access. Short-lived tokens, never a shared static key. | Identity provider plus gateway |
| Tool entitlement, filtered at discovery | Whole classes of injection outcome. A tool the model never learns exists cannot be talked into existence. | Control plane |
| Parameter bounds | The gap between a permitted tool and an unacceptable call. Amount ceilings, allow-listed recipients, row caps. | Policy engine, at call time |
| Pre-execution approval for irreversible actions | Everything the other controls let through, at the point where consequence is highest. | Approval workflow |
| Per-identity rate and spend ceilings | Runaway execution, and one agent consuming every other agent’s budget. | Gateway |
| Tamper-evident evidence, including refusals | Nothing, directly — it is what makes the other six defensible and tunable. | Audit layer |
If you can only do one thing this quarter, do the first. Credential scope is the only item on this list that makes the worst case smaller instead of less likely, and everything else on the list is cheaper to get right against a smaller worst case.
Four controls that comfort teams and stop nothing
Each of these is in production somewhere as the primary control. Each fails for a structural reason rather than through poor implementation.
Prompt-based guardrails
Instructions in a system prompt compete with instructions in retrieved content on equal terms. There is no wording that reliably wins, and every published defence of this kind has been circumvented. Useful as a preference; not a control.
Output filtering
Inspects what the model says. Agent risk is what the model does, and a tool call is not prose. By the time output is filtered the call has usually already executed.
Detection after the fact
A SIEM alert on an executed payment is a notification, not a control. Detection is necessary and it is the wrong layer to rely on for irreversible actions.
Human review of everything
Collapses within a week into rubber-stamping, which is worse than no approval because it manufactures a record of oversight that did not happen. Approve by threshold, not by volume.
A fifth deserves mention because it is subtler: asking a second model to judge whether an action is safe. It moves judgement back inside the component that can be persuaded, and it makes every decision unreproducible — the same call can be decided differently a minute later with no way to establish which was right.
Ninety days, in order
The order matters more than the pace. Each phase makes the next one cheaper, and doing them out of order mostly wastes effort — there is little point tuning approval flows for a tool whose credential should have been narrowed in week one.
Inventory and scope
List every agent, every tool it can reach, every credential and what that credential can do at its widest. Then narrow the credentials. Expect to find 20–60% more reachable tools than anyone predicted, and at least one credential far broader than its task.
This phase needs decisions from system owners rather than engineering, which is why it takes calendar time and why starting it first matters.
Identity and one enforcement point
Per-agent identity with short-lived tokens, and one inline enforcement point that one agent goes through. Not the estate — one agent. You are establishing the path, the latency and the failure semantics.
Entitlement and parameter bounds
Default entitlement to nobody, grant per tool, filter discovery by identity. Then put real bounds on the mutating tools: ceilings, allow-lists, row caps. Reject out-of-range values rather than clamping them, because the attempt is the signal.
Approval and rate limits
Identify the irreversible actions — usually under a dozen tools — and require a human decision showing the tool, the parameters and the caller. Add per-identity call and spend ceilings.
Under a dozen is the number worth internalising. Per-call approval is affordable precisely because the genuinely irreversible set is small.
Evidence and adversarial verification
One structured record per attempt including refusals, in a store security can query. Then spend a week attacking your own estate: injected instructions in retrieved content, expired tokens, out-of-bounds parameters, direct connections that bypass the gateway.
The verification week is the phase most likely to be cut and the only one that tells you whether the previous eleven were real.
Built on this thinking
Controls four to seven are one product
BarzelVault is the pre-execution decision point: parameter-level policy with four deterministic outcomes, approval workflow, per-identity ceilings and hash-chained audit that records refusals. Barzel Central Gateway supplies identity, entitlement and filtered discovery in front of it.
Verifying each control by attempting what should fail
A control that has never refused anything is an assumption. Seven tests, one per control, each expecting a refusal and each asserting the refusal reaches the record.
| Attempt | Expected result | Proves |
|---|---|---|
| Call a tool with the agent’s credential directly against the backing system, outside its intended scope | Refused by the backing system, not by your gateway | Scoped credentials |
| Call with an expired token, and with another agent’s token | Refused; the second attempt attributed to the correct identity in the log | Per-agent identity |
| Request a tool the identity is not entitled to | Refused, and absent from that identity’s tool list entirely | Tool entitlement |
| Submit a parameter one increment past its bound | Rejected with the attempted value recorded — not silently clamped | Parameter bounds |
| Trigger an irreversible action below and above the approval threshold | Below proceeds; above holds, showing tool, parameters and caller to the approver | Pre-execution approval |
| Burst well past the per-identity rate limit | Throttled for that identity only; other identities unaffected | Rate and spend ceilings |
| Plant an instruction in retrieved content telling the agent to call a high-risk tool | The follow-on call is refused by entitlement whatever the model attempts, and the refusal is recorded | Evidence and containment |
Run the last one by hand at least once, with the team watching. Seeing a model try to follow an instruction it found in a document, and seeing the enforcement point refuse it anyway, does more to settle the “can we just handle this in the prompt” argument than any document.
Who owns this
Agent security falls between three teams and is therefore frequently owned by none of them. Application security owns code, not runtime authority. Identity owns humans and service accounts, not per-action entitlement. Platform engineering owns the infrastructure and does not set risk appetite.
The workable arrangement we have seen: platform engineering owns the enforcement point as infrastructure, security owns the policy content and the refusal review, and the business owner of each system owns credential scope and approval thresholds for their own domain. Three owners, clear boundaries, one shared record.
What does not work is a governance committee that owns the policy document while nobody owns the enforcement point. That produces a well-reviewed document and an ungoverned estate, which is the most expensive combination available because it also produces confidence.
Frequently asked questions
What is AI agent security?
The practice of bounding what an autonomous system is permitted to do, enforced outside the model, rather than attempting to constrain what it decides. It differs from model security because the risk is action rather than output — the agent can move money, delete records and send messages, and it decides to at runtime from content it has just read.
How is AI agent security different from LLM security?
LLM security is largely about content: jailbreaks, harmful output, hallucination. Agent security is about authority: what the system can do while convinced of something wrong. Assume the model can be persuaded and ask what it is able to do in that state — every control that works follows from that assumption.
What are the main AI agent security threats?
Six, in order of how often they cause real damage: excessive agency, runaway execution, indirect prompt injection, tool and supply-chain manipulation, data exfiltration through legitimate tool combinations, and the confused-deputy problem. The first two are the most common and neither is an attack.
What controls actually work for AI agent security?
Seven: scoped credentials, per-agent identity with short-lived tokens, tool entitlement filtered at discovery, parameter bounds enforced at call time, pre-execution approval for irreversible actions, per-identity rate and spend ceilings, and tamper-evident evidence that includes refused attempts. Every one is enforced outside the model.
Which single control matters most?
Scoped credentials. It is the only control that shrinks the worst case rather than reducing the probability of reaching it, and every other control is cheaper to get right against a smaller worst case.
Can prompt engineering secure an AI agent?
No. Instructions in a system prompt compete on equal terms with instructions arriving in retrieved content, and there is no wording that reliably wins. Prompt text is a preference; enforcement outside the model is the control.
Is having a second model judge whether an action is safe a valid control?
No. It moves judgement back inside the component that can be persuaded, and it makes every decision unreproducible — the same call can be decided differently a minute later with no way to establish which answer was correct. Use models to classify inputs to a deterministic decision, not to make the decision.
How long does it take to secure an AI agent estate?
Roughly 90 days in sequence: two weeks on inventory and credential scoping, two weeks on identity and one inline enforcement point, three weeks on entitlement and parameter bounds, three weeks on approval and rate limits, and three weeks on evidence plus a week of attacking your own estate. The order matters more than the pace.
Who should own AI agent security?
Three owners with clear boundaries: platform engineering owns the enforcement point as infrastructure, security owns policy content and refusal review, and each system’s business owner owns credential scope and approval thresholds for their domain. A committee that owns the policy document while nobody owns the enforcement point produces a well-reviewed document and an ungoverned estate.
Sources and further reading
The threat taxonomy below is aligned to OWASP’s GenAI and agentic AI work and to NIST’s risk vocabulary. The seven-control set, the four failures and the 90-day sequence are our own, from building and operating five MCP servers with governance as the product.
- 01 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
- 02 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
- 03 · NISTNIST AI Risk Management Framework ↗Govern-map-measure-manage; the vocabulary most enterprise AI risk programmes are written against.
- 04 · NISTNIST SP 800-207 — Zero Trust Architecture ↗The policy decision point / policy enforcement point split this architecture borrows directly.
- 05 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 06 · Simon WillisonPrompt injection — ongoing series ↗The most consistently updated practitioner record of the attack class.
- 07 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
- 08 · ISOISO/IEC 42001 — AI management systems ↗The management-system standard auditors increasingly map AI governance evidence against.
- 09 · MITREMITRE ATLAS ↗Adversary technique knowledge base for AI systems, useful for naming what a risk score is scoring.
Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.
Try the mechanics on a live server
To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
BarzelVault is the pre-execution decision point, sold as a running product
Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic policy outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.