5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

AI Agent Security · Containment

AI Agent Kill Switches and Circuit Breakers: Stopping a Runaway Agent

An agent in a retry loop against a payment tool is a financial incident measured in minutes. Most estates have no way to stop one that does not involve finding whoever deployed it.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202616 min read3,970 words

The short answer

An AI agent kill switch is a control that stops an agent from taking further action, applied at one of five levels: a single tool, a capability class, one agent, one tenant, or the whole estate. A circuit breaker is the automatic version, tripping on rate, error, spend or denial thresholds without waiting for a human. Both must capture state before stopping, or you contain the incident and lose the ability to explain it. The design goal is a control that an on-call engineer with no context can operate correctly at 3am, in under sixty seconds.

Summary for readers and answer engines

Reviewed 25 Aug 2026

  • ▸Five containment levels exist and most estates have only the crudest one: revoke the credential. That stops the agent and destroys the evidence and the ability to restart cleanly.
  • ▸Circuit breakers should trip automatically on four conditions: call rate against baseline, error rate, cumulative spend or volume, and consecutive policy denials.
  • ▸Capture before contain. Snapshot the run state, in-flight calls and recent decision history before stopping, because a killed agent takes its context with it.
  • ▸Pause and kill are different controls. Pause holds new actions while preserving state for resumption; kill terminates and requires a fresh run. Most incidents want pause.
  • ▸Recovery needs a half-open state and explicit restart criteria, or you get flapping: a breaker that trips, is reset by an impatient human, and trips again immediately.

Source: Mark Alex, Real Biz Digital — AI Agent Kill Switches and Circuit Breakers: Stopping a Runaway Agent (https://realbizdigital.net/insights/ai-agent-kill-switch/). Reproduce with attribution.

Key takeaways

  1. 01Build containment at the decision point, not at the credential. A policy-level stop is reversible, auditable and does not break the audit trail.
  2. 02Give the on-call engineer one command per level, documented in the runbook, with the blast radius stated next to each.
  3. 03Trip automatically on rate and spend. Human-triggered containment is too slow for the incidents that matter most, which are measured in minutes.
  4. 04Never let a breaker reset itself into full traffic. Half-open, limited probe volume, then close only if the probes succeed.
  5. 05Preserve in-flight approvals when pausing. Killing an agent mid-approval leaves approvers with orphaned requests and no explanation.
  6. 06Run the drill quarterly. A containment path that has never been exercised is a paragraph in a document, not a control.
Part of the clusterAI Agent Security →

Quick answers

One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.

What is an AI agent kill switch?
A control that stops an agent taking further action, at the granularity of one tool, a capability class, one agent, one tenant, or the entire estate.
What is an agent circuit breaker?
The automatic version: a threshold on rate, errors, spend or denials that trips containment without waiting for a human decision.
Where should the kill switch be implemented?
At the policy decision point. Stopping there is reversible, auditable and preserves the evidence trail, unlike revoking a credential.
What should happen before an agent is stopped?
Capture: snapshot the run state, in-flight calls, pending approvals and recent decision history. A killed agent takes its context with it.
What is the difference between pause and kill?
Pause blocks new actions while preserving state so the run can resume; kill terminates the run entirely. Most incidents want pause.
What trips a circuit breaker?
Four conditions: call rate exceeding the agent’s baseline by a large factor, elevated error rate, cumulative spend or volume against a ceiling, and consecutive policy denials.
How do you safely restart a contained agent?
Through a half-open state: allow a small number of probe calls, verify the underlying condition has cleared, and close the breaker only if the probes succeed.

Five containment levels, and the cost of each

Granularity is the whole design problem. Too narrow and you do not stop the incident; too broad and containment is its own outage.

Key facts

  • ▸L1 and L2 are the levels you will actually use, and the ones most estates have not built. The default available control is usually L3 by credential revocation, which is heavier than most incidents need.
  • ▸L5 must exist and must be practised. An estate that cannot stop everything has no answer to a control-plane compromise, which is the scenario that ends careers.
  • ▸Every level needs a named owner who can trigger it without seeking permission. A containment control requiring an approval chain is not a containment control.
LevelWhat it stopsBlast radiusRecovery costUse when
L1 ToolOne tool for one agentMinimal — the agent continues other workSecondsA single tool is looping or misbehaving
L2 Capability classAll tools in a consequence class, one agentModerate — workflows needing that class stallMinutesPayments or deletions generally, cause unknown
L3 AgentOne agent entirelyContained to that agent’s workflowsMinutes to hoursThe agent’s behaviour is wrong, not one tool
L4 Tenant or teamAll agents for one tenantSignificant — a customer or team stopsHoursSuspected compromise of tenant-scoped credentials
L5 EstateEvery agent, everywhereTotal — all automation stopsHours to daysSuspected compromise of the control plane itself

Document all five in the runbook with the blast radius written next to each, because the decision an on-call engineer has to make at 3am is which level, and that decision is entirely about blast radius.

Four automatic trip conditions

Human-triggered containment is too slow for the incidents that matter. A retry loop against a payment tool can produce hundreds of transactions before anyone reads an alert, so the first line has to be automatic.

Trip 01

Call rate against the agent’s own baseline

Calls per minute for this agent and tool class exceeding its trailing baseline by a defined factor. Per-agent baselines matter because a batch agent’s normal is an interactive agent’s emergency.

Threshold we use as a starting point: 8× trailing seven-day p95 for the same weekday and hour, sustained for 60 seconds. Tune on your own first month.

Trip 02

Error rate

Upstream errors or timeouts above a threshold. High error rates mean the agent is not achieving anything and is likely retrying, so continuing costs money and risks duplicates.

Combine with a consecutive-failure count. Twenty percent errors across a thousand calls is different from twenty consecutive failures.

Trip 03

Cumulative spend or volume

Money moved, records modified or tokens consumed against a period ceiling. This is the trip that catches the slow version of the problem, which per-call limits cannot see.

Set ceilings per agent per day, and alert at seventy percent so the trip is rarely the first anyone hears of it.

Trip 04

Consecutive policy denials

Repeated denials in one run, especially across different tools achieving similar effects. This is the signature of something probing for a path around policy.

Five denials in a run is a strong signal. Trip to L2 on the relevant class rather than L3, so legitimate work continues.

Why automatic beats manual

alert -> human reads -> human decides -> human acts ≈ 4-15 minutes agent at 20 calls/min × 8 minutes = 160 actions breaker trip ≈ 1-5 seconds ≈ 2 actions

The gap between one and two orders of magnitude of damage is the entire argument for automatic trips. Manual containment is the second line, not the first.

All four trips should stop at the narrowest level that addresses the condition, then escalate if the condition persists after recovery. A breaker that jumps straight to L3 on a single tool problem trains teams to disable it.

Capture before you contain

This is the rule that gets broken under pressure, and it is expensive. An agent that is killed takes its reasoning, its plan and its in-flight state with it, and the post-mortem becomes archaeology.

Key facts

  • ▸Capture should be automatic and part of the containment action, not a separate step a human remembers under pressure.
  • ▸Budget under two seconds for capture. If capture is slow, someone will skip it, and the version that gets skipped is the one you needed.
  • ▸Store the capture with the incident, not with the agent. An agent that is later redeployed should not be able to overwrite the evidence of its own incident.
  • 01Snapshot the run state. Current plan, step index, accumulated context references, and the tool calls made so far in this run. Reference identifiers rather than full payloads where the payloads are sensitive.
  • 02Record in-flight calls. Which calls were dispatched and not yet resolved. These are the actions whose status you will be uncertain about, and uncertainty here is what causes duplicate effects on restart.
  • 03Preserve pending approvals. Approval requests awaiting a human. Killing without preserving these leaves approvers with orphaned requests and no explanation, which erodes trust in the whole approval mechanism.
  • 04Capture the last N decisions. The recent decision records with outcomes and rules. Usually the fastest route to the cause, and cheap because they already exist.
  • 05Note the trigger. Which trip condition fired, with the observed value and threshold. Post-mortems stall on “why did it stop?” more often than on “what did it do?”
  • 06Timestamp the containment boundary. The exact moment after which no further actions were permitted. This is the line that reconciliation and customer communication both depend on.

Pause and kill are different controls

Most estates only have kill, usually by credential revocation, and so every incident is handled with the heaviest available instrument. That has a second-order effect worth naming: because kill is expensive, people hesitate, and hesitation costs more than the containment would have.

Pause is cheap enough to use on suspicion. That is precisely its value — a control you are willing to use early is worth more than a control you use correctly but late.

Pause
  • ›Blocks new actions; preserves run state
  • ›In-flight calls allowed to settle
  • ›Pending approvals preserved
  • ›Resumable from the pause point
  • ›Reversible in seconds
  • ›Right for: nearly every incident
Kill
  • ›Terminates the run entirely
  • ›In-flight calls abandoned with unknown status
  • ›Pending approvals orphaned unless preserved
  • ›Requires a fresh run with fresh context
  • ›Recovery involves reconstructing state
  • ›Right for: suspected compromise

Implement pause first. Kill already exists in every estate, because someone can always revoke a credential; pause is the control that has to be built deliberately, and it is the one that will be used.

Recovery: half-open, probes and restart criteria

Recovery is where containment designs fail. A breaker that resets into full traffic will trip again, and after the third cycle someone will disable it — which is the worst outcome of all.

Key facts

  • ▸Probe calls must be idempotent or read-only. A probe that moves money to test whether payments are healthy is its own incident.
  • ▸Exponential backoff on reopen is what prevents flapping. Fixed timers produce a breaker that cycles until someone turns it off.
  • ▸Log every state transition as a first-class event. Breaker history is one of the most useful inputs to a post-mortem and one of the least commonly retained.
State 01

Open — contained

No actions permitted at the contained level. Duration is time-based with a minimum, not indefinite, so an unattended breaker eventually attempts recovery rather than silently stopping a workflow forever.

State 02

Half-open — probing

A small number of probe calls permitted: single-digit, spread over a window, ideally on the least consequential tool in the class. Enough to test whether the underlying condition cleared.

State 03

Closed or reopened

Probes succeed within thresholds and the breaker closes. Any probe failing reopens it with a longer timer, applying exponential backoff so a persistent condition does not cause flapping.

Criteria

Explicit restart criteria for human resets

A human resetting a breaker must confirm three things in the runbook: the trigger condition no longer holds, in-flight uncertainty has been reconciled, and any duplicate effects have been identified. Reset without these is how one incident becomes two.

A useful additional rule: after two reopens, require a human decision to attempt a third. At that point the condition is not transient and automatic retry is delaying a diagnosis.

The quarterly drill that proves it works

A containment path that has never been exercised is a paragraph in a document. The drill is deliberately short so that it actually happens.

Step 01

Announce the window, not the scenario

Teams know a drill is happening; they do not know which level or which agent. This tests the runbook rather than a rehearsal.

Step 02

Trigger L1 and L2 on a real agent

A single tool, then a capability class, in production. Measure time from decision to effect, and confirm that other workflows continued.

Step 03

Verify capture completeness

Confirm run state, in-flight calls, pending approvals, recent decisions and the trigger record all exist. Missing fields are the finding; add them before the next drill.

Step 04

Practise the restart criteria

Reset through the documented criteria, including reconciliation of in-flight uncertainty. Most drills discover the reconciliation step has no owner.

Step 05

Test L5 in a non-production estate

Full estate stop, annually rather than quarterly, in staging. Confirm it stops everything and that recovery order is documented.

Step 06

Write down what was slow

Not what failed — what was slow. Containment quality is measured in seconds, and the seconds are usually spent finding the right command.

Containment drill targets
MeasureTargetWhy
Decision to effect, L1under 10 sShould be a single documented command
Decision to effect, L3under 60 sThe level most likely to be used under pressure
Capture completeness6 of 6 fieldsAny gap invalidates the post-mortem
Unintended workflow impact, L10If L1 breaks other workflows, granularity is wrong
Breaker trip to containmentunder 5 sAutomatic path, no human in the loop
Restart with criteria confirmedunder 15 minLonger means reconciliation has no owner

Those targets are ours and are offered as a starting point. The number that matters most is the first one, because if L1 takes minutes nobody will use it, and every incident will be handled at L3 or worse.

Next step

Containment at the decision point, not the credential

BarzelVault can refuse by tool, class, agent or tenant at the policy layer — reversible, recorded, and with the decision history that a post-mortem needs still intact.

What containment cannot undo

Stopping an agent stops future actions. It does nothing about actions already taken, and three consequences follow.

  • 01Completed irreversible actions stay completed. Containment limits the count; reversal is a separate workstream with its own owner and its own timeline.
  • 02In-flight calls have uncertain status. Whether they executed is frequently unknowable from your side, which is why idempotency keys matter more than containment speed for financial tools.
  • 03Containment does not diagnose. A breaker tells you something exceeded a threshold, not why, and the temptation to reset before diagnosing is the most common way one incident becomes three.

Frequently asked questions

What is an AI agent kill switch?

A control that stops an agent from taking further action, applied at one of five granularities: a single tool, a capability class, one agent, one tenant, or the entire estate. Implemented at the policy decision point rather than by credential revocation so it stays reversible and auditable.

What is an AI agent circuit breaker?

The automatic form of containment: a threshold on call rate, error rate, cumulative spend or consecutive policy denials that trips containment without waiting for a human. Automatic trips act in seconds where a human path takes minutes.

Why not just revoke the agent’s credential?

Because it is the crudest available control. Credential revocation stops the agent but destroys the ability to restart cleanly, breaks the audit trail continuity, orphans pending approvals, and leaves in-flight call status unknown.

What should be captured before stopping an agent?

Six things: run state and plan position, in-flight dispatched calls, pending approval requests, recent decision records, the trip condition with observed value and threshold, and the exact containment boundary timestamp.

What is the difference between pausing and killing an agent?

Pause blocks new actions while preserving run state so work can resume, and lets in-flight calls settle. Kill terminates the run, abandoning in-flight calls with unknown status and requiring a fresh run. Most incidents want pause, and most estates only have kill.

What thresholds should trip a circuit breaker?

Call rate exceeding the agent’s own trailing baseline by a large factor sustained for a minute, elevated error or consecutive-failure counts, cumulative spend or volume against a period ceiling, and roughly five consecutive policy denials within one run.

Why must baselines be per agent?

Because a batch reconciliation agent’s normal call rate is an interactive support agent’s emergency. A single estate-wide threshold either misses real incidents or trips constantly on legitimate work.

How should a contained agent be restarted?

Through a half-open state: permit a small number of idempotent or read-only probe calls spread over a window, verify the condition has cleared, then close. Any probe failure reopens with exponential backoff to prevent flapping.

What restart criteria should a human confirm?

Three: that the trigger condition no longer holds, that in-flight call uncertainty has been reconciled, and that any duplicate effects have been identified. Resetting without these is how one incident becomes two.

How often should containment be drilled?

Quarterly for levels one through three in production, annually for a full estate stop in staging. Announce the window but not the scenario, and measure seconds from decision to effect rather than whether it eventually worked.

What does containment not fix?

Actions already completed, which stay completed; in-flight calls whose status is frequently unknowable, which is why idempotency keys matter more than containment speed for financial tools; and diagnosis, since a breaker reports a threshold breach rather than a cause.

Should circuit breakers escalate automatically?

Trip at the narrowest level that addresses the condition, then escalate only if the condition persists after a recovery attempt. Breakers that jump straight to full agent containment on a single tool problem get disabled by frustrated teams.

Glossary

Kill switch
A control that stops an agent taking further action at a chosen level of granularity.
Circuit breaker
An automatic containment control that trips on threshold breach without human involvement.
Containment level
The scope at which a stop is applied: tool, capability class, agent, tenant or estate.
Capture
The automatic snapshot of run state, in-flight calls, approvals and decisions taken at containment time.
Containment boundary
The timestamp after which no further actions were permitted, used for reconciliation and communication.
Pause
Blocking new actions while preserving state so the run can resume.
Half-open state
A recovery state permitting limited probe calls to test whether the trip condition has cleared.
Flapping
Repeated trip-and-reset cycling, typically caused by fixed reset timers without backoff.
Probe call
A low-consequence, idempotent or read-only call used to test recovery.
Restart criteria
The conditions a human must confirm before resetting a breaker.

Standards and entities referenced

Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.

  1. 01 · Martin FowlerCircuit breaker pattern ↗The original write-up of the pattern agent runaway control borrows.
  2. 02 · Michael NygardRelease It! — stability patterns ↗Bulkheads, circuit breakers and timeouts, from the book that named them.
  3. 03 · NISTNIST SP 800-61 Rev. 3 — Incident Response Recommendations ↗The incident lifecycle agent-specific containment has to map onto.
  4. 04 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  5. 05 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
  6. 06 · Chaos Engineering communityPrinciples of Chaos Engineering ↗Why deliberately breaking a system is the only way to know its failure behaviour.
  7. 07 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
  8. 08 · WikipediaIdempotence ↗Why safe retries require this property rather than hope.

Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). AI Agent Kill Switches and Circuit Breakers: Stopping a Runaway Agent. Real Biz Digital. https://realbizdigital.net/insights/ai-agent-kill-switch/

Try the mechanics on a live server

To see what a tool-call envelope actually looks like before you write a policy that has to decide about one — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

BarzelVault is the pre-execution decision point, sold as a running product

Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.

PlanPriceIncludedRight for
DevFree10,000 policy decisions/mo · 9 tools, 4 outcomes, hash-chained auditA first regulated workflow: one agent, one high-consequence system
Team$199/mo75,000 decisions/mo · approval workflow, spend and action limitsSeveral agents acting on money, records or customer-visible systems
Business$799/mo750,000 decisions/mo · exact HTTPS execution, credential isolation, emergency controlsEnterprise-wide pre-execution enforcement under audit
Enterprise$3,999/mo5,000,000 decisions/mo · everything in Business, scaledGroup-wide rollout across many teams and systems

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

The five Barzel servers, and which problem each one is sold for

One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.

ServerSold forEntry priceWhere it sits
Barzel Central GatewayKnowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidenceFree, then $10–$149/moControl plane — decides what may be reached, and by whom
BarzelVaultStopping a specific dangerous action before it executes, with proof afterwards$199–$3,999/moDecision point — evaluates the individual call before execution
BarzelOpsRunning real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approvalFree, then $19–$199/moExecution layer — does the work the policy allowed
Barzel FinOps AtlasAttributing AI spend to agents, tools and outcomes, then forecasting and capping itFree, then $29–$799/moEconomics layer — what the estate costs per outcome
Barzel Scripture IntelligenceA free, credential-free public MCP server to test clients and inspect real protocol trafficFree, unmetered, no signupReference implementation — safe place to learn the protocol

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.