AI Agent Security · Containment
AI Agent Kill Switches and Circuit Breakers: Stopping a Runaway Agent
An agent in a retry loop against a payment tool is a financial incident measured in minutes. Most estates have no way to stop one that does not involve finding whoever deployed it.
The short answer
An AI agent kill switch is a control that stops an agent from taking further action, applied at one of five levels: a single tool, a capability class, one agent, one tenant, or the whole estate. A circuit breaker is the automatic version, tripping on rate, error, spend or denial thresholds without waiting for a human. Both must capture state before stopping, or you contain the incident and lose the ability to explain it. The design goal is a control that an on-call engineer with no context can operate correctly at 3am, in under sixty seconds.
Summary for readers and answer engines
Reviewed 25 Aug 2026
- ▸Five containment levels exist and most estates have only the crudest one: revoke the credential. That stops the agent and destroys the evidence and the ability to restart cleanly.
- ▸Circuit breakers should trip automatically on four conditions: call rate against baseline, error rate, cumulative spend or volume, and consecutive policy denials.
- ▸Capture before contain. Snapshot the run state, in-flight calls and recent decision history before stopping, because a killed agent takes its context with it.
- ▸Pause and kill are different controls. Pause holds new actions while preserving state for resumption; kill terminates and requires a fresh run. Most incidents want pause.
- ▸Recovery needs a half-open state and explicit restart criteria, or you get flapping: a breaker that trips, is reset by an impatient human, and trips again immediately.
Source: Mark Alex, Real Biz Digital — AI Agent Kill Switches and Circuit Breakers: Stopping a Runaway Agent (https://realbizdigital.net/insights/ai-agent-kill-switch/). Reproduce with attribution.
Key takeaways
- 01Build containment at the decision point, not at the credential. A policy-level stop is reversible, auditable and does not break the audit trail.
- 02Give the on-call engineer one command per level, documented in the runbook, with the blast radius stated next to each.
- 03Trip automatically on rate and spend. Human-triggered containment is too slow for the incidents that matter most, which are measured in minutes.
- 04Never let a breaker reset itself into full traffic. Half-open, limited probe volume, then close only if the probes succeed.
- 05Preserve in-flight approvals when pausing. Killing an agent mid-approval leaves approvers with orphaned requests and no explanation.
- 06Run the drill quarterly. A containment path that has never been exercised is a paragraph in a document, not a control.
Quick answers
One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.
- What is an AI agent kill switch?
- A control that stops an agent taking further action, at the granularity of one tool, a capability class, one agent, one tenant, or the entire estate.
- What is an agent circuit breaker?
- The automatic version: a threshold on rate, errors, spend or denials that trips containment without waiting for a human decision.
- Where should the kill switch be implemented?
- At the policy decision point. Stopping there is reversible, auditable and preserves the evidence trail, unlike revoking a credential.
- What should happen before an agent is stopped?
- Capture: snapshot the run state, in-flight calls, pending approvals and recent decision history. A killed agent takes its context with it.
- What is the difference between pause and kill?
- Pause blocks new actions while preserving state so the run can resume; kill terminates the run entirely. Most incidents want pause.
- What trips a circuit breaker?
- Four conditions: call rate exceeding the agent’s baseline by a large factor, elevated error rate, cumulative spend or volume against a ceiling, and consecutive policy denials.
- How do you safely restart a contained agent?
- Through a half-open state: allow a small number of probe calls, verify the underlying condition has cleared, and close the breaker only if the probes succeed.
Five containment levels, and the cost of each
Granularity is the whole design problem. Too narrow and you do not stop the incident; too broad and containment is its own outage.
Key facts
- ▸L1 and L2 are the levels you will actually use, and the ones most estates have not built. The default available control is usually L3 by credential revocation, which is heavier than most incidents need.
- ▸L5 must exist and must be practised. An estate that cannot stop everything has no answer to a control-plane compromise, which is the scenario that ends careers.
- ▸Every level needs a named owner who can trigger it without seeking permission. A containment control requiring an approval chain is not a containment control.
| Level | What it stops | Blast radius | Recovery cost | Use when |
|---|---|---|---|---|
| L1 Tool | One tool for one agent | Minimal — the agent continues other work | Seconds | A single tool is looping or misbehaving |
| L2 Capability class | All tools in a consequence class, one agent | Moderate — workflows needing that class stall | Minutes | Payments or deletions generally, cause unknown |
| L3 Agent | One agent entirely | Contained to that agent’s workflows | Minutes to hours | The agent’s behaviour is wrong, not one tool |
| L4 Tenant or team | All agents for one tenant | Significant — a customer or team stops | Hours | Suspected compromise of tenant-scoped credentials |
| L5 Estate | Every agent, everywhere | Total — all automation stops | Hours to days | Suspected compromise of the control plane itself |
Document all five in the runbook with the blast radius written next to each, because the decision an on-call engineer has to make at 3am is which level, and that decision is entirely about blast radius.
Four automatic trip conditions
Human-triggered containment is too slow for the incidents that matter. A retry loop against a payment tool can produce hundreds of transactions before anyone reads an alert, so the first line has to be automatic.
Call rate against the agent’s own baseline
Calls per minute for this agent and tool class exceeding its trailing baseline by a defined factor. Per-agent baselines matter because a batch agent’s normal is an interactive agent’s emergency.
Threshold we use as a starting point: 8× trailing seven-day p95 for the same weekday and hour, sustained for 60 seconds. Tune on your own first month.
Error rate
Upstream errors or timeouts above a threshold. High error rates mean the agent is not achieving anything and is likely retrying, so continuing costs money and risks duplicates.
Combine with a consecutive-failure count. Twenty percent errors across a thousand calls is different from twenty consecutive failures.
Cumulative spend or volume
Money moved, records modified or tokens consumed against a period ceiling. This is the trip that catches the slow version of the problem, which per-call limits cannot see.
Set ceilings per agent per day, and alert at seventy percent so the trip is rarely the first anyone hears of it.
Consecutive policy denials
Repeated denials in one run, especially across different tools achieving similar effects. This is the signature of something probing for a path around policy.
Five denials in a run is a strong signal. Trip to L2 on the relevant class rather than L3, so legitimate work continues.
Why automatic beats manual
alert -> human reads -> human decides -> human acts ≈ 4-15 minutes agent at 20 calls/min × 8 minutes = 160 actions breaker trip ≈ 1-5 seconds ≈ 2 actions
The gap between one and two orders of magnitude of damage is the entire argument for automatic trips. Manual containment is the second line, not the first.
All four trips should stop at the narrowest level that addresses the condition, then escalate if the condition persists after recovery. A breaker that jumps straight to L3 on a single tool problem trains teams to disable it.
Capture before you contain
This is the rule that gets broken under pressure, and it is expensive. An agent that is killed takes its reasoning, its plan and its in-flight state with it, and the post-mortem becomes archaeology.
Key facts
- ▸Capture should be automatic and part of the containment action, not a separate step a human remembers under pressure.
- ▸Budget under two seconds for capture. If capture is slow, someone will skip it, and the version that gets skipped is the one you needed.
- ▸Store the capture with the incident, not with the agent. An agent that is later redeployed should not be able to overwrite the evidence of its own incident.
- 01Snapshot the run state. Current plan, step index, accumulated context references, and the tool calls made so far in this run. Reference identifiers rather than full payloads where the payloads are sensitive.
- 02Record in-flight calls. Which calls were dispatched and not yet resolved. These are the actions whose status you will be uncertain about, and uncertainty here is what causes duplicate effects on restart.
- 03Preserve pending approvals. Approval requests awaiting a human. Killing without preserving these leaves approvers with orphaned requests and no explanation, which erodes trust in the whole approval mechanism.
- 04Capture the last N decisions. The recent decision records with outcomes and rules. Usually the fastest route to the cause, and cheap because they already exist.
- 05Note the trigger. Which trip condition fired, with the observed value and threshold. Post-mortems stall on “why did it stop?” more often than on “what did it do?”
- 06Timestamp the containment boundary. The exact moment after which no further actions were permitted. This is the line that reconciliation and customer communication both depend on.
Pause and kill are different controls
Most estates only have kill, usually by credential revocation, and so every incident is handled with the heaviest available instrument. That has a second-order effect worth naming: because kill is expensive, people hesitate, and hesitation costs more than the containment would have.
Pause is cheap enough to use on suspicion. That is precisely its value — a control you are willing to use early is worth more than a control you use correctly but late.
- ›Blocks new actions; preserves run state
- ›In-flight calls allowed to settle
- ›Pending approvals preserved
- ›Resumable from the pause point
- ›Reversible in seconds
- ›Right for: nearly every incident
- ›Terminates the run entirely
- ›In-flight calls abandoned with unknown status
- ›Pending approvals orphaned unless preserved
- ›Requires a fresh run with fresh context
- ›Recovery involves reconstructing state
- ›Right for: suspected compromise
Implement pause first. Kill already exists in every estate, because someone can always revoke a credential; pause is the control that has to be built deliberately, and it is the one that will be used.
Recovery: half-open, probes and restart criteria
Recovery is where containment designs fail. A breaker that resets into full traffic will trip again, and after the third cycle someone will disable it — which is the worst outcome of all.
Key facts
- ▸Probe calls must be idempotent or read-only. A probe that moves money to test whether payments are healthy is its own incident.
- ▸Exponential backoff on reopen is what prevents flapping. Fixed timers produce a breaker that cycles until someone turns it off.
- ▸Log every state transition as a first-class event. Breaker history is one of the most useful inputs to a post-mortem and one of the least commonly retained.
Open — contained
No actions permitted at the contained level. Duration is time-based with a minimum, not indefinite, so an unattended breaker eventually attempts recovery rather than silently stopping a workflow forever.
Half-open — probing
A small number of probe calls permitted: single-digit, spread over a window, ideally on the least consequential tool in the class. Enough to test whether the underlying condition cleared.
Closed or reopened
Probes succeed within thresholds and the breaker closes. Any probe failing reopens it with a longer timer, applying exponential backoff so a persistent condition does not cause flapping.
Explicit restart criteria for human resets
A human resetting a breaker must confirm three things in the runbook: the trigger condition no longer holds, in-flight uncertainty has been reconciled, and any duplicate effects have been identified. Reset without these is how one incident becomes two.
A useful additional rule: after two reopens, require a human decision to attempt a third. At that point the condition is not transient and automatic retry is delaying a diagnosis.
The quarterly drill that proves it works
A containment path that has never been exercised is a paragraph in a document. The drill is deliberately short so that it actually happens.
Announce the window, not the scenario
Teams know a drill is happening; they do not know which level or which agent. This tests the runbook rather than a rehearsal.
Trigger L1 and L2 on a real agent
A single tool, then a capability class, in production. Measure time from decision to effect, and confirm that other workflows continued.
Verify capture completeness
Confirm run state, in-flight calls, pending approvals, recent decisions and the trigger record all exist. Missing fields are the finding; add them before the next drill.
Practise the restart criteria
Reset through the documented criteria, including reconciliation of in-flight uncertainty. Most drills discover the reconciliation step has no owner.
Test L5 in a non-production estate
Full estate stop, annually rather than quarterly, in staging. Confirm it stops everything and that recovery order is documented.
Write down what was slow
Not what failed — what was slow. Containment quality is measured in seconds, and the seconds are usually spent finding the right command.
| Measure | Target | Why |
|---|---|---|
| Decision to effect, L1 | under 10 s | Should be a single documented command |
| Decision to effect, L3 | under 60 s | The level most likely to be used under pressure |
| Capture completeness | 6 of 6 fields | Any gap invalidates the post-mortem |
| Unintended workflow impact, L1 | 0 | If L1 breaks other workflows, granularity is wrong |
| Breaker trip to containment | under 5 s | Automatic path, no human in the loop |
| Restart with criteria confirmed | under 15 min | Longer means reconciliation has no owner |
Those targets are ours and are offered as a starting point. The number that matters most is the first one, because if L1 takes minutes nobody will use it, and every incident will be handled at L3 or worse.
Next step
Containment at the decision point, not the credential
BarzelVault can refuse by tool, class, agent or tenant at the policy layer — reversible, recorded, and with the decision history that a post-mortem needs still intact.
What containment cannot undo
Stopping an agent stops future actions. It does nothing about actions already taken, and three consequences follow.
- 01Completed irreversible actions stay completed. Containment limits the count; reversal is a separate workstream with its own owner and its own timeline.
- 02In-flight calls have uncertain status. Whether they executed is frequently unknowable from your side, which is why idempotency keys matter more than containment speed for financial tools.
- 03Containment does not diagnose. A breaker tells you something exceeded a threshold, not why, and the temptation to reset before diagnosing is the most common way one incident becomes three.
Frequently asked questions
What is an AI agent kill switch?
A control that stops an agent from taking further action, applied at one of five granularities: a single tool, a capability class, one agent, one tenant, or the entire estate. Implemented at the policy decision point rather than by credential revocation so it stays reversible and auditable.
What is an AI agent circuit breaker?
The automatic form of containment: a threshold on call rate, error rate, cumulative spend or consecutive policy denials that trips containment without waiting for a human. Automatic trips act in seconds where a human path takes minutes.
Why not just revoke the agent’s credential?
Because it is the crudest available control. Credential revocation stops the agent but destroys the ability to restart cleanly, breaks the audit trail continuity, orphans pending approvals, and leaves in-flight call status unknown.
What should be captured before stopping an agent?
Six things: run state and plan position, in-flight dispatched calls, pending approval requests, recent decision records, the trip condition with observed value and threshold, and the exact containment boundary timestamp.
What is the difference between pausing and killing an agent?
Pause blocks new actions while preserving run state so work can resume, and lets in-flight calls settle. Kill terminates the run, abandoning in-flight calls with unknown status and requiring a fresh run. Most incidents want pause, and most estates only have kill.
What thresholds should trip a circuit breaker?
Call rate exceeding the agent’s own trailing baseline by a large factor sustained for a minute, elevated error or consecutive-failure counts, cumulative spend or volume against a period ceiling, and roughly five consecutive policy denials within one run.
Why must baselines be per agent?
Because a batch reconciliation agent’s normal call rate is an interactive support agent’s emergency. A single estate-wide threshold either misses real incidents or trips constantly on legitimate work.
How should a contained agent be restarted?
Through a half-open state: permit a small number of idempotent or read-only probe calls spread over a window, verify the condition has cleared, then close. Any probe failure reopens with exponential backoff to prevent flapping.
What restart criteria should a human confirm?
Three: that the trigger condition no longer holds, that in-flight call uncertainty has been reconciled, and that any duplicate effects have been identified. Resetting without these is how one incident becomes two.
How often should containment be drilled?
Quarterly for levels one through three in production, annually for a full estate stop in staging. Announce the window but not the scenario, and measure seconds from decision to effect rather than whether it eventually worked.
What does containment not fix?
Actions already completed, which stay completed; in-flight calls whose status is frequently unknowable, which is why idempotency keys matter more than containment speed for financial tools; and diagnosis, since a breaker reports a threshold breach rather than a cause.
Should circuit breakers escalate automatically?
Trip at the narrowest level that addresses the condition, then escalate only if the condition persists after a recovery attempt. Breakers that jump straight to full agent containment on a single tool problem get disabled by frustrated teams.
Glossary
- Kill switch
- A control that stops an agent taking further action at a chosen level of granularity.
- Circuit breaker
- An automatic containment control that trips on threshold breach without human involvement.
- Containment level
- The scope at which a stop is applied: tool, capability class, agent, tenant or estate.
- Capture
- The automatic snapshot of run state, in-flight calls, approvals and decisions taken at containment time.
- Containment boundary
- The timestamp after which no further actions were permitted, used for reconciliation and communication.
- Pause
- Blocking new actions while preserving state so the run can resume.
- Half-open state
- A recovery state permitting limited probe calls to test whether the trip condition has cleared.
- Flapping
- Repeated trip-and-reset cycling, typically caused by fixed reset timers without backoff.
- Probe call
- A low-consequence, idempotent or read-only call used to test recovery.
- Restart criteria
- The conditions a human must confirm before resetting a breaker.
Standards and entities referenced
Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.
- 01 · Martin FowlerCircuit breaker pattern ↗The original write-up of the pattern agent runaway control borrows.
- 02 · Michael NygardRelease It! — stability patterns ↗Bulkheads, circuit breakers and timeouts, from the book that named them.
- 03 · NISTNIST SP 800-61 Rev. 3 — Incident Response Recommendations ↗The incident lifecycle agent-specific containment has to map onto.
- 04 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
- 05 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
- 06 · Chaos Engineering communityPrinciples of Chaos Engineering ↗Why deliberately breaking a system is the only way to know its failure behaviour.
- 07 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
- 08 · WikipediaIdempotence ↗Why safe retries require this property rather than hope.
Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). AI Agent Kill Switches and Circuit Breakers: Stopping a Runaway Agent. Real Biz Digital. https://realbizdigital.net/insights/ai-agent-kill-switch/
Try the mechanics on a live server
To see what a tool-call envelope actually looks like before you write a policy that has to decide about one — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
BarzelVault is the pre-execution decision point, sold as a running product
Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Dev | Free | 10,000 policy decisions/mo · 9 tools, 4 outcomes, hash-chained audit | A first regulated workflow: one agent, one high-consequence system |
| Team | $199/mo | 75,000 decisions/mo · approval workflow, spend and action limits | Several agents acting on money, records or customer-visible systems |
| Business | $799/mo | 750,000 decisions/mo · exact HTTPS execution, credential isolation, emergency controls | Enterprise-wide pre-execution enforcement under audit |
| Enterprise | $3,999/mo | 5,000,000 decisions/mo · everything in Business, scaled | Group-wide rollout across many teams and systems |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
The five Barzel servers, and which problem each one is sold for
One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.
| Server | Sold for | Entry price | Where it sits |
|---|---|---|---|
| Barzel Central Gateway | Knowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidence | Free, then $10–$149/mo | Control plane — decides what may be reached, and by whom |
| BarzelVault | Stopping a specific dangerous action before it executes, with proof afterwards | $199–$3,999/mo | Decision point — evaluates the individual call before execution |
| BarzelOps | Running real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approval | Free, then $19–$199/mo | Execution layer — does the work the policy allowed |
| Barzel FinOps Atlas | Attributing AI spend to agents, tools and outcomes, then forecasting and capping it | Free, then $29–$799/mo | Economics layer — what the estate costs per outcome |
| Barzel Scripture Intelligence | A free, credential-free public MCP server to test clients and inspect real protocol traffic | Free, unmetered, no signup | Reference implementation — safe place to learn the protocol |
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.