5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

Runtime · controls

AI Agent Runtime Security: Protecting Agents While They Act

You can review an agent’s code, its prompts and its tool list, and still know almost nothing about what it will do — because every input that determines its behaviour arrives after the review is finished.

By Mark Alex, FounderPublished 24 Aug 2026Updated 2 Sep 202613 min

The short answer

AI agent runtime security is the set of controls that apply while an agent executes, rather than at build or review time. It matters because an agent’s behaviour is determined by inputs that do not exist until runtime — retrieved documents, tool results, user requests. There are five control points inside a single tool call: identity validation, entitlement, parameter evaluation, execution containment and result inspection. Circuit breakers and kill switches bound the damage a loop can do. Plus the fail-closed decision, which you should make before an outage makes it for you.

Why build-time review is not enough

A code review answers what the software can do. For most software that is nearly the whole question, because the inputs are constrained and the control flow is written down.

An agent inverts the ratio. Its control flow is chosen at runtime by a model, on the basis of content that arrived at runtime, much of which came from outside your organisation. You can review the tool list, the prompts and the framework code and still not know whether the agent will issue a refund this afternoon, because the document that will convince it does not exist yet.

This does not make review worthless — it is where credential scope, tool surface and parameter schemas get fixed, and those are the highest-leverage decisions available. But review is the wrong place to look for behavioural guarantees. The guarantee has to be enforced at the moment of action, which means at runtime, by something that is not the agent.

Five control points inside one tool call

A single tool call passes five places where a decision can be made. Most estates enforce at one or two of them and are surprised by what gets through the others.

01 · Point 01

Identity, at every request

Validate the token on each call, not once at connection. Long-lived sessions are how a revoked agent keeps working for another six hours, and connection-time-only validation is the default in more clients than you would expect.

02 · Point 02

Entitlement, before the tool is even visible

Filter what the identity may see, then refuse what it may not call. Two separate enforcements: one at discovery, one at invocation. Discovery filtering shrinks the attack surface; invocation enforcement is what actually holds.

03 · Point 03

Parameter evaluation, at call time

The values are the risk. Ceilings, allow-lists, row caps, date ranges, enumerated values — evaluated against policy that can see the identity and the environment as well as the value. Reject rather than clamp: the attempt is the signal you want recorded.

04 · Point 04

Execution containment

The server runs as a non-root user in its own container, with egress restricted to the systems it legitimately calls and a timeout on every call. Egress restriction is the item most often skipped and the one that most limits damage.

A contained execution cannot be used to move data somewhere else, whatever the agent is persuaded to attempt. It is the difference between a bad call and an exfiltration.

05 · Point 05

Result inspection

What comes back enters the model’s context and competes with your instructions. Cap response size, strip or neutralise control markup, label the boundary between data and instruction, and scan for the classes of sensitive data you are obliged to protect.

This does not solve indirect prompt injection — nothing at this layer does — but it removes the cheapest version and it is where data-loss prevention has to live.

Loops: the incident that is not an attack

The most common runtime failure in the estates we have seen has no adversary in it. An agent calls a tool, the result is not what it expected, it tries again with a slight variation, and the variation does not help. Thirty seconds later it has made four hundred calls.

This is worse than it sounds for three reasons. Each call may have a side effect, so four hundred calls can mean four hundred emails or four hundred partial writes. It consumes a shared budget, so one agent’s loop starves every well-behaved agent. And it is indistinguishable from abuse without per-identity accounting, so nobody notices until the invoice or the upstream rate-limit response arrives.

Three mechanisms, in the order they should be added.

MechanismTriggerBehaviour
Per-identity rate ceilingCalls per minute and per day, per agent identityRefuse further calls for that identity only. Other agents unaffected — this is why the ceiling is per identity and not per server.
Circuit breakerError or refusal rate above a threshold in a short windowStop calling that tool or route entirely, then re-admit gradually. For mutating tools, probe with a read tool rather than half-opening on a call with side effects.
Repeat-call detectorN near-identical calls from one identity within a windowRefuse and alert. This is the mechanism that catches the loop specifically, rather than catching it as a volume side effect.

Add a spend ceiling alongside the call ceiling wherever calls cost money. Setting agent spend ceilings covers the numbers; the important design point is that the ceiling has to refuse rather than notify.

Kill switches, and what makes one real

Every agent estate needs a way to stop an agent immediately. Most have one in principle. Four properties separate the ones that work from the ones that get discovered to be theoretical during an incident.

Out-of-band

It must not require the agent, its framework or its orchestrator to cooperate. If stopping the agent means the agent has to notice a flag, it is a request rather than a switch.

Granular and global

Stop one agent, one tool class, or everything. An estate-wide switch is the only one people build and the least likely to be used, because using it stops the business too.

Reachable by whoever is on call

Not by the team that built the agent, during business hours, via a deploy. A runbook entry with a command, tested.

Tested on a schedule

An untested kill switch is a hypothesis. Exercise it quarterly against a real agent in production and record how long it took.

The number worth measuring is time-to-stop: from decision to the agent making no further calls. Under a minute is achievable; teams that have never tested usually discover theirs is measured in tens of minutes, most of it spent finding who has access.

Built on this thinking

Five control points, one enforcement layer

BarzelVault covers points two, three and five — entitlement, parameter evaluation and result inspection — with execution-resilience controls for loops and circuit breaking, and hash-chained audit across all of it. Barzel Central Gateway validates identity per request and holds the per-identity ceilings.

Fail-closed or fail-open: decide before the outage

If the enforcement point is unavailable, do calls proceed or stop? Both answers are defensible and the wrong time to choose is at 3am while a queue backs up.

Fail-closed means unavailability stops agent work. Correct for anything irreversible: a refused payment is an inconvenience, an unauthorised one is an incident. It also means your enforcement point is now a production dependency with the availability requirement that implies.

Fail-open means work continues ungoverned. Defensible for read-only, low-consequence traffic where blocking causes more harm than the unlikely bad call. Indefensible for the X-class actions.

The workable answer is neither, uniformly — it is per risk class. Read-only tools fail open. Bounded writes fail closed with a short grace window for in-flight calls. Irreversible actions fail closed with no grace, ever. Write the table down, and state what happens to a pending approval when the enforcement point restarts, because that is the case nobody specifies and the one that produces a duplicated payment.

Risk classOn enforcement-point failureRationale
Read, non-sensitiveFail openBlocking costs more than the residual risk. Log the ungoverned window.
Read, sensitiveFail closed, short graceDisclosure is not reversible either. Grace only for in-flight calls.
Bounded writeFail closed, short graceReversible, but reversal costs real effort and someone’s afternoon.
High-reach or hard-to-reverse writeFail closed, no graceThe whole reason the enforcement point exists.
Irreversible or externally visibleFail closed, no grace, alertNever proceeds unevaluated. A pending approval that survives a restart must not auto-approve.

Testing runtime controls

Runtime controls are only real if they have refused something. Five exercises, run on a schedule rather than once.

  • 01Revoke a token mid-session and confirm the next call fails. This is the test that catches connection-time-only validation, and it fails more often than teams expect.
  • 02Drive a deliberate loop — a tool that always returns an unhelpful result — and confirm the rate ceiling and repeat detector both fire, and that other agents keep working.
  • 03Kill the enforcement point mid-call for each risk class and confirm the behaviour matches the table you wrote.
  • 04Restart with an approval pending and confirm it neither auto-approves nor silently disappears.
  • 05Plant an instruction in retrieved content asking the agent to call a tool it is not entitled to, and confirm the refusal happens at the enforcement point and appears in the record.

Four mistakes worth naming

Validating identity at connection only

A revoked agent keeps working until its session ends, which on a long-lived connection can be hours.

Rate limits per server rather than per identity

One runaway agent consumes every other agent’s headroom, and the limit cannot tell you which agent caused it.

Uniform fail-closed everywhere

Sounds rigorous, blocks harmless reads during every deploy, and gets switched off after the second incident — usually globally.

A kill switch nobody has used

Untested means unknown. Measure time-to-stop once a quarter; the first measurement is always the surprising one.

Frequently asked questions

What is AI agent runtime security?

The set of controls that apply while an agent executes rather than at build or review time. It matters because an agent’s behaviour is determined by inputs that do not exist until runtime — retrieved documents, tool results, user requests — so review can fix capability but cannot give behavioural guarantees.

What are the runtime control points in an agent tool call?

Five: identity validation on every request, entitlement enforced both at discovery and at invocation, parameter evaluation at call time, execution containment with restricted egress and timeouts, and result inspection before the response re-enters the model’s context.

Why is build-time review insufficient for AI agents?

Because an agent’s control flow is chosen at runtime by a model, from content that arrived at runtime and often originated outside your organisation. Review fixes credential scope, tool surface and parameter schemas — the highest-leverage decisions — but it cannot tell you what the agent will do this afternoon.

What is an AI agent circuit breaker?

A mechanism that stops calls to a tool or route once error or refusal rate crosses a threshold in a short window, then re-admits traffic gradually. For mutating tools, probe recovery with a read tool rather than half-opening on a call that has side effects.

What makes an AI agent kill switch real rather than theoretical?

Four properties: it works out-of-band without the agent or its framework cooperating, it is granular as well as global, it is reachable by whoever is on call rather than by the building team during business hours, and it is exercised quarterly. Measure time-to-stop — the first measurement is usually the surprising one.

Should AI agent security fail closed or fail open?

Per risk class, not uniformly. Read-only non-sensitive traffic fails open with the ungoverned window logged. Sensitive reads and bounded writes fail closed with a short grace for in-flight calls. High-reach and irreversible actions fail closed with no grace. Uniform fail-closed blocks harmless reads during every deploy and gets switched off globally after the second incident.

What is the most common AI agent runtime incident?

A loop with no adversary in it: the agent calls a tool, gets an unexpected result, retries with a variation, and thirty seconds later has made hundreds of calls. Each may have a side effect, it starves every other agent’s budget, and without per-identity accounting nobody notices until the invoice arrives.

How do you test runtime controls?

Five exercises on a schedule: revoke a token mid-session and confirm the next call fails; drive a deliberate loop and confirm ceilings fire while other agents keep working; kill the enforcement point mid-call for each risk class; restart with an approval pending and confirm it neither auto-approves nor disappears; and plant an instruction in retrieved content asking for an unentitled tool.

Sources and further reading

Circuit-breaker and error-budget practice is standard SRE material, cited below. The five control points and the fail-closed-by-risk-class scheme are our own design, implemented across the Barzel servers.

  1. 01 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
  2. 02 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  3. 03 · NISTNIST SP 800-207 — Zero Trust Architecture ↗The policy decision point / policy enforcement point split this architecture borrows directly.
  4. 04 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
  5. 05 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  6. 06 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
  7. 07 · OWASPOWASP Application Security Verification Standard ↗Input-validation, authorization and logging requirements restated here in MCP terms.
  8. 08 · Simon WillisonPrompt injection — ongoing series ↗The most consistently updated practitioner record of the attack class.

Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.

Try the mechanics on a live server

To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

BarzelVault is the pre-execution decision point, sold as a running product

Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic policy outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.

DevFree10,000 calls/mo
Team$199/mo75,000 calls/mo
Business$799/mo750,000 calls/mo
Enterprise$3,999/mo5,000,000 calls/mo

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.