5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

Incident response · runbook

AI Agent Incident Response: Containing a Compromised Agent

Conventional incident response assumes you can isolate a host and read what happened. An agent’s reasoning left no artefact, its context is gone the moment you kill it, and the damage was done through entirely legitimate calls.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202616 min2,958 words

The short answer

AI agent incident response differs from conventional incident response in three ways: there is nothing to patch because the agent was not exploited but persuaded, containment means removing authority rather than isolating a host, and the evidence lives in the context window rather than on disk. The first ten minutes should capture context and revoke authority, in that order — killing the agent first destroys the record that explains what happened. A ten-minute runbook, six incident types, and the blast-radius questions that determine notification.

Key takeaways

  1. 01Capture before you contain. Killing the agent destroys the context window, which is the only place its reasoning exists.
  2. 02Containment is authority removal, not host isolation. Revoke the credential; the agent process is not the problem.
  3. 03There is nothing to patch. The agent was not exploited — it was persuaded, and the same persuasion works on a fresh instance.
  4. 04Reconstruct blast radius from the evidence record, not from the agent. Ask it what it did and it will confabulate plausibly.
  5. 05Assume every read the agent performed while compromised is disclosed, and every write is suspect.
  6. 06The recovery question is not “is it fixed?” but “what bound would have made this uninteresting?”

Three ways this differs from normal IR

Standard incident response has a shape: detect, contain the host, eradicate the vulnerability, recover, learn. Each stage assumes something that does not hold for agents.

There is nothing to eradicate. The agent was not exploited. No vulnerability was used, no code executed that should not have. It read something and did what it was asked. Patching is not available as a step, and a rebuilt agent is exactly as susceptible as the one you removed.

Isolating the host does not contain it. The damage vector is the credential, not the process. An agent whose container is stopped but whose token remains valid is uncontained if anything else can present that token. Containment means revoking authority.

The evidence is volatile in an unusual way. Conventional forensics reads disk and memory. The agent’s reasoning exists only in its context window — the conversation, the retrieved documents, the tool responses — and that is gone the instant you terminate the session. Which creates a genuine conflict with the instinct to kill it immediately.

Key facts

  • ▸An agent incident is usually persuasion, not exploitation. There is no patch.
  • ▸Containment is credential revocation, not process termination.
  • ▸Terminating the agent destroys the only record of why it acted.

The first ten minutes

Order matters and it is counter-intuitive. Capture precedes containment, because containment destroys evidence and the capture takes seconds.

0–1 min

Snapshot the context, do not terminate

Dump the agent’s full conversation, retrieved content and recent tool responses to durable storage. Automate this in advance as a single command; nobody composes an export script during an incident.

This is the step teams skip and regret. Every subsequent question — what convinced it, which document, whether other agents read the same source — is answerable only from this snapshot.

1–3 min

Revoke the credential and the entitlements

Not the process. Revoke the token, disable the identity, withdraw entitlements at the enforcement point. If revocation takes longer than two minutes, that is a finding for afterwards; note it and continue.

3–5 min

Freeze pending approvals and in-flight workflows

Anything held awaiting a human must not auto-approve, and partially applied workflows must stop rather than complete. Record which workflows are mid-flight and at which step — that list determines the recovery work.

5–7 min

Pull the action list from the evidence record

Every call by that identity in the relevant window: succeeded, refused, and with what parameters. From the audit store, never from the agent.

Do not ask the agent what it did. It will produce a fluent, plausible, partly-invented account, and you will spend the next hour disproving it.

7–10 min

Check for lateral spread

Did this agent delegate to others? Did other agents read the same source content? A poisoned document reaches every agent with access to it, and the second agent is usually discovered late.

Query by the source document, not only by the agent. This is where a single-agent incident becomes an estate incident, and where chain ids earn their keep.

Six incident types

Different containment, different recovery. Identifying which one you have in the first ten minutes changes what you do next.

TypeSignatureContainmentRecovery difficulty
Injected instructionRefusals spike on one agent; an unusual source document in contextRevoke; quarantine the source contentModerate — find and remove the source
Runaway loopCall rate per identity spikes; near-identical parametersRate ceiling should have caught it; revoke and inspectLow — usually no adversary
Over-entitlement discoveredAn agent did something legitimate and unintendedNarrow the entitlement, not the agentLow technically, awkward organisationally
Credential compromiseCalls from unexpected sources, times or IPs with a valid tokenRevoke and rotate; check every use of that credentialHigh — conventional IR applies
Poisoned tool or serverBehaviour changed after a version bump; description diff shows an editPin back; isolate the server; audit every call since the upgradeHigh — supply-chain scope
Chain escalationAn agent achieved something outside its scope via delegationBreak the chain; revoke the intermediate agent’s delegation rightsHigh — requires chain reconstruction

The third row is the most common in practice and the least like a security incident. Nothing was attacked; an agent did exactly what its entitlements permitted and the outcome was unwanted. Treat it as an incident anyway — the response is a control change, and routing it through the incident process is what ensures the change actually happens.

Reconstructing blast radius

The question everyone asks first and it is answerable only from the record. Six queries, in the order that matters for notification decisions.

  • 01What did it read, and was any of it personal or regulated data? Assume every successful read during the compromise window is disclosed. This is the query that determines whether you have a notification obligation, so run it first.
  • 02What did it write, and where? Every mutating call, with parameters. Each is either intended or not, and each unintended one needs a compensating action.
  • 03What left the organisation? Emails, messages, published documents, payments, share links. Irreversible effects and the ones a counterparty saw.
  • 04What was refused? The refusals tell you what the attacker was trying to achieve, which is often more informative than what succeeded. They also demonstrate which controls held.
  • 05Which other agents touched the same source? Query by document, not by agent. Lateral spread through shared content is the agent-specific propagation path.
  • 06What is mid-flight? Partially applied workflows leave the world in an inconsistent state that will not resolve itself — see partial application.

A note on the first query. “Assume disclosed” is the defensible posture: you cannot establish what a model retained, summarised or repeated, so treating every read during the window as exposed is both safer and easier to explain than a claim about what the model probably did with it.

Built on this thinking

Blast radius is a query when refusals are recorded too

Barzel Central Gateway records every attempt including refusals, with identity, parameters, rule and route — so the six blast-radius questions are queries rather than reconstruction. BarzelVault adds hash-chained integrity, so the record holds up when it becomes evidence.

Recovery and re-enablement

The pressure to restore service arrives quickly, and it arrives before anyone has decided what “fixed” means. Decide the criteria before the incident — that is the point of a runbook.

Four criteria. All four, not a majority.

The bound exists that would have made this uninteresting
Not “we removed the malicious document”. The question is what parameter bound, entitlement change or approval gate would have made the action refused regardless of what the agent believed. If the answer is none, the agent should not be re-enabled yet.
The source is removed and its class is addressed
Deleting one poisoned ticket does not stop the next one. What changed about how that content class is handled — scanning, entitlement, a bound on what it can trigger?
Compensating actions are complete
Every unintended write reversed or offset, every mid-flight workflow resolved to a consistent state, every counterparty who saw something contacted.
The refusal now happens in a test
Reproduce the attack against the changed configuration and observe the refusal in the log. Without this, the fix is a belief.

Re-enable narrower than before, then widen deliberately. An agent restored to its previous entitlements after an incident is an agent whose incident taught the estate nothing, and the second occurrence will be harder to explain than the first.

What to prepare beforehand

Five things, each of which takes an hour now and saves the incident.

PreparationWhyTest it by
One-command context snapshotNobody writes an export script mid-incidentRunning it against a live agent, quarterly
One-action credential revocationIf revocation needs a deploy, it is not revocationTiming it — target under a minute
Query templates for the six blast-radius questionsComposing SIEM queries under pressure produces wrong answersRunning them against last week’s traffic
A named on-call with authority to revokeEscalating to find who can act is the commonest delayA quarterly drill against a real agent
Documented re-enablement criteriaOtherwise service pressure decidesReviewing them when nothing is on fire

The drill is the one to insist on. Exercise a real revocation against a real agent in production once a quarter and measure time-to-stop. Teams that have never done this reliably discover their number is tens of minutes, most of it spent establishing who has access — and they discover it during the incident.

Honest limits of this runbook

None of this tells you whether an incident occurred. Detection remains the hard part, and the leading indicators — refusal-rate spikes, unusual parameter distributions, call-rate anomalies — require instrumentation most estates do not have. A runbook assumes somebody noticed.

Reconstruction is also bounded by what you recorded. If refusals were not logged, you cannot know what the attacker attempted. If parameters were not captured, you cannot establish magnitude. The quality of an incident response is fixed months earlier, by decisions about what to instrument, and no amount of process compensates for a record that was never written.

And the honest position on eradication is that there is none. You cannot patch persuadability. Every response ends in bounding consequence more tightly than before, which means an estate that has had one agent incident should expect to have another — with a smaller blast radius each time if the recovery criteria are applied, and the same one if they are not.

Frequently asked questions

How does AI agent incident response differ from conventional IR?

Three ways: there is nothing to patch, because the agent was persuaded rather than exploited; containment means revoking authority rather than isolating a host; and the evidence lives in the context window rather than on disk, so it is destroyed the moment the agent is terminated.

What should you do in the first ten minutes of an agent incident?

Snapshot the agent’s context before terminating anything; revoke the credential and entitlements rather than the process; freeze pending approvals and in-flight workflows; pull the action list from the evidence record; then check for lateral spread to other agents that read the same source.

Why should you capture context before containing an agent?

Because the agent’s reasoning exists only in its context window — the conversation, retrieved documents and tool responses — and terminating the session destroys it. Every subsequent question about what convinced the agent is answerable only from that snapshot.

Should you ask a compromised agent what it did?

No. It will produce a fluent, plausible, partly-invented account and you will spend an hour disproving it. Pull the action list from the audit record instead, which is the only source that reflects what actually happened.

What are the types of AI agent incident?

Six: injected instruction, runaway loop, over-entitlement discovered, credential compromise, poisoned tool or server, and chain escalation through delegation. Over-entitlement is the most common and the least like a conventional security incident — nothing was attacked.

How do you reconstruct the blast radius of an agent incident?

Six queries: what it read and whether any was personal or regulated data, what it wrote and where, what left the organisation, what was refused, which other agents touched the same source content, and what workflows are mid-flight. Assume every read during the window is disclosed.

How do you know when a contained agent can be re-enabled?

Four criteria, all of them: a bound now exists that would have made the action refused regardless of the agent’s belief; the source is removed and its whole class addressed; compensating actions are complete; and the attack reproduced against the new configuration produces an observable refusal.

What should you prepare before an agent incident?

Five things: a one-command context snapshot, one-action credential revocation, query templates for the blast-radius questions, a named on-call with authority to revoke, and documented re-enablement criteria. Drill the revocation quarterly and measure time-to-stop.

Why can’t you patch a compromised AI agent?

Because it was not exploited. No vulnerability was used and no unauthorised code ran — it read something and did what it was asked. A rebuilt agent is exactly as susceptible, so every response ends in bounding consequence rather than removing a flaw.

What limits the quality of an agent incident response?

What was instrumented months earlier. If refusals were not logged you cannot know what was attempted; if parameters were not captured you cannot establish magnitude. No amount of process compensates for a record that was never written.

Glossary

Agent compromise
A state in which an agent is taking actions directed by an attacker, typically through injected content rather than through code execution.
Containment
Removing an agent’s ability to act — revoking credentials and entitlements — as distinct from stopping the process running it.
Context capture
Preserving the agent’s conversation and retrieved content before termination, since it is the only record of why the agent acted.
Confabulation risk
The tendency of an agent asked to explain its own actions to produce a plausible account unsupported by the record.
Re-enablement criteria
The stated conditions under which a contained agent is permitted to act again, decided before the pressure to restore service arrives.

Sources and further reading

The incident lifecycle follows NIST SP 800-61, which is authoritative on structure. The agent-specific differences, the ten-minute sequence and the confabulation warning are our own operating judgement.

  1. 01 · NISTNIST SP 800-61 — Computer Security Incident Handling Guide ↗The incident lifecycle this article maps onto agent-specific containment.
  2. 02 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
  3. 03 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  4. 04 · MITREMITRE ATT&CK ↗Adversary tactics and techniques referenced when scoping agent-estate response.
  5. 05 · MITREMITRE ATLAS ↗Adversary technique knowledge base for AI systems, useful for naming what a risk score is scoring.
  6. 06 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
  7. 07 · European UnionGDPR — Regulation (EU) 2016/679 ↗Lawful basis, data minimisation and processing records that agent estates inherit.
  8. 08 · Simon WillisonPrompt injection — ongoing series ↗The most consistently updated practitioner record of the attack class.

Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). AI Agent Incident Response: Containing a Compromised Agent. Real Biz Digital. https://realbizdigital.net/insights/ai-agent-incident-response/

Try the mechanics on a live server

To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

Barzel Central Gateway is this layer, sold as a running product

Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so an evaluation costs an afternoon rather than a purchase order.

PlanPriceIncludedRight for
CommunityFree1,000 tool calls/mo · full policy engine, registry, auditEvaluating the estate, or a single team proving the path works
Starter$10/mo10,000 calls/moOne or two production agents against a handful of servers
Team$79/mo100,000 calls/moA platform team governing an estate of 5–20 servers
Business$149/mo250,000 calls/moEstate-wide governance with SIEM evidence and multi-team routing

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.