5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

AI governance

Preventing unauthorised AI agent actions

An agent action is unauthorised when it crosses a boundary nobody approved — and that happens in three distinct ways, which need three distinct controls. The agent can exceed assumed authority while faithfully following its instructions. Untrusted content can reach its context and redirect it. Or it can act entirely correctly on wrong data, at machine speed. Most programmes build one control and quietly assume it covers all three. It does not, and the gap usually surfaces as an action that cannot be reversed.

How Barzel applies here Start free with BarzelVault

What does "unauthorised" actually cover?

The word flattens three different mechanisms. They share an outcome and share nothing else — different triggers, different failure signatures, different places where a control can intervene.

ClassWhat happensWhere the control has to sit
Scope creepThe agent exceeds assumed authority while following its instructions correctlyA permission boundary outside the agent: action type, value, counterparty, data class
Injected instructionUntrusted content reaching the model's context redirects the agentTrust boundaries, output filtering, enforcement at the tool call
Malformed inputThe agent acts correctly on wrong data, repeatedly and fastRate and behavioural limits — not value thresholds

Scope creep is the class regulators have already named. FINRA's 2026 Annual Regulatory Oversight Report lists scope creep beyond intended authority among its AI agent concerns, alongside autonomy without human validation, auditability of multi-step reasoning, sensitive-data disclosure, domain-knowledge gaps and misaligned reward structures. Its recommendations — human-in-the-loop protocols, tracking mechanisms and behavioural guardrails — apply to all three classes, but only if built separately. Scoping permissions addresses the first class well and the third barely at all.

How an injected instruction actually reaches an agent

CVE-2026-75130 is worth studying because it documents the whole chain rather than a hypothetical. The affected component is Upstash's Context7 MCP documentation server, versions through 2.1.2; the weakness class is CWE-1427. Malicious instructions enter through the Custom AI Instructions feature. A developer then asks for library documentation — an ordinary request for an ordinary reason — and the unsanitised content reaches the coding agent's context. The documented impacts were credential exfiltration from environment files to an attacker-controlled endpoint and destructive file deletion.

Note what is absent: no compromised credential, no memory-safety bug, no privilege escalation. The agent held the permissions it was meant to hold and used them. What changed was the content of its context, arriving through the channel the tool exists to provide.

Why the CVE record is a poor instrument

The disclosure timeline is instructive for anyone planning to rely on vulnerability feeds as a control.

EventDateAttributed to
Fix deployed23 February 2026Documented by Noma Security
Public disclosure, as "ContextCrush"5 March 2026Noma Security
CVE record published18 August 2026VulnCheck

Roughly five months separate public disclosure from the CVE record, and the record carries no patch reference despite a fix deployed before disclosure. Severity is unsettled: VulnCheck scored it 9.0 under CVSS 3.1 and 6.4 under CVSS 4.0, the NVD has not assessed it, and the NVD's CPE entry misspells the vendor as Uptash.

So a team whose defence against injected instructions is "we track advisories for the servers we use" would have had no signal for five months, no patch pointer when one arrived, and a vendor search that missed it. Advisories describe what happened elsewhere; they are not a control over what your agent does next. That argues for runtime enforcement — see the MCP security analysis for the protocol-level picture.

Can detection fix an irreversible action?

This is the architectural point that decides whether a control programme works, and the one most often skipped.

Control typeWhat it doesEffect on an irreversible action
PreventiveRefuses the action before it executesChanges the outcome
DetectiveReports that the action happenedProduces a record of a loss
CorrectiveReverses the action afterwardsOnly works where reversal exists

Detective and corrective controls are cheaper, easier to retrofit and better at demonstrating diligence, which is why they accumulate. But apply them to the impacts in CVE-2026-75130. A credential sent to an attacker-controlled endpoint cannot be un-sent, only rotated after the fact, once you know. A deleted file may be restorable from backup, or may not.

So the design question is not "will we know?" but "if we know, can we act?" Where the answer is no, the action belongs behind an approval step or a hard refusal. Detection still matters — an audit trail is what makes an incident investigable, and the NCSC's Thinking carefully before adopting agentic AI guidance of 15 May 2026 recommends monitoring and AI-specific incident response. It just does not substitute for prevention on actions that cannot be walked back.

Why prompt-level defences are not a control boundary

The instinct after reading about an injection is to strengthen the system prompt: tell the agent to ignore instructions in retrieved content, to refuse to read environment files, never to send secrets to an external endpoint. That helps at the margin and is not a boundary.

The test is simple. If a change to the system prompt can change what the agent is permitted to do, the permission lives in the model's context — and anything reaching that context can move it. That is precisely the channel the injection uses. A defence and an attack occupying the same layer are not a boundary; they are an argument, and the agent adjudicates it.

The NSA's Cybersecurity Information Sheet on MCP (Ver. 1.0, May 2026) states the principle plainly: each output must be treated as untrusted input to the next phase of the pipeline. Its recommendations follow — filter for content length, disallowed keywords, rate limits and indirect prompt injection; sandbox tool execution using OS-level frameworks such as AppContainers, SELinux, AppArmor and seccomp; apply least privilege to agent processes; clearly define trust boundaries between agents, plugins, models and end users; segregate tools by data classification.

Every one of those sits outside the model, which is not incidental. Enforcement has to live in a layer the agent passes through and cannot reconfigure, so a redirected agent meets the same refusal as an honest one. The protocol layer is moving the same way: MCP 2026-07-28 added RFC 9207 issuer validation — authorization servers return iss and clients must validate it before redeeming codes, closing an authorization-server mix-up hole — and bound client credentials to the issuer that minted them.

Why do rate limits catch what value limits miss?

The third class survives good permission design, because nothing about it looks wrong. An agent processing a malformed feed issues a credit note: plausible amount, real counterparty, granted action type, well under any per-action ceiling. It is a correct action on incorrect data. Then it does it again, and again, at a rate no human could match. The value threshold is never crossed, so it never fires, and a per-action approval rule sees nothing exceptional to approve.

Rate and behavioural anomaly detection surfaces this dozens of actions before a value threshold does, because it measures the wrong thing on purpose: not whether any action is unusual, but whether the pattern is. Thirty credit notes in four minutes is anomalous even when all thirty are individually ordinary. This is why the NSA guidance lists rate limits beside content filtering, and why a per-window cap belongs alongside every per-action cap.

The NCSC reaches the same class from the other direction: model failure modes before connecting an agent to live data, use ephemeral credentials, keep explicit human accountability chains. Modelling failure modes first is how a team discovers that all its limits are per-action ones — before a feed proves it.

Frequently asked questions

What counts as an unauthorised agent action?

Any action outside a boundary someone approved. It arises through scope creep, injected instructions, or correct action on malformed data — three mechanisms needing three controls.

Can a better system prompt prevent unauthorised actions?

No. If a prompt change can change what the agent may do, the constraint sits in the model's context, and anything reaching that context can move it.

What does CVE-2026-75130 demonstrate?

The full injection chain in a real deployment: untrusted content entering through a documentation feature, reaching a coding agent's context, driving credential exfiltration and file deletion.

Is detection enough if the audit trail is good?

Not for irreversible actions. Detective and corrective controls produce a record; only a preventive control changes the outcome.

Why add rate limits when value limits exist?

Because in a malformed-input failure every individual action sits under the value limit. Only a rate or behavioural limit sees the pattern.

Related

BarzelVault sits in front of the agent as a preventive layer: deterministic risk scoring and policy-as-code decide each proposed action before execution, atomic spend and action limits cap what gets through, an emergency kill switch and revocation stop a run in progress, and dry-run simulation lets you test a policy without side effects.

In practice

Permission before the action. Evidence after it.

The duties on this page attach to the moment an automated system acts: who permitted it, on which data, under which policy version, and what a person saw before approving. Barzel enforces that decision before execution and writes the record an auditor, a regulator or a data subject can be shown.

430 days leftEU AI Act high-risk obligations (Annex III) apply from 2 December 2027

BarzelVault

The AI action firewall: decide what an agent may do before it does it.

  • Approval thresholds and policy checks enforced before execution; human approvals that expire and escalate.
  • Cryptographically signed audit receipts: trigger, inputs, policy version, approver, outcome.
  • Credential isolation, spend and action limits, and an emergency kill switch.

Free tier: 10,000 calls a monthPaid plans from $199 a monthLive on MCPize

Start free Ask by emailProduct pageDocumentation

Barzel Central Gateway

The AI governance control plane: one inventory and one policy layer across every MCP server and agent.

  • Registers and synchronises every tool; enforces identity, policy, region, cost and health per tool.
  • Identity mapping through OIDC, Entra ID, Okta, SAML and SPIFFE, with credential brokerage.
  • Trace and SIEM export (W3C trace context, OTLP) for the security team and the regulator.

Free tier: 1,000 calls a monthPaid plans from $10 a monthLive on MCPize

Start free Ask by emailProduct pageDocumentation

Enterprise: written quote by email within two business days. No sales call.

Sources

  1. VulnCheck, CVE-2026-75130, published 18 August 2026.
  2. Noma Security, ContextCrush, disclosed 5 March 2026; fix deployed 23 February 2026.
  3. NSA, CSI: Model Context Protocol (MCP) — Security Design Considerations for AI-Driven Automation, Ver. 1.0, May 2026.
  4. FINRA, 2026 Annual Regulatory Oversight Report, generative AI section.
  5. NCSC, Thinking carefully before adopting agentic AI, 15 May 2026.
  6. Model Context Protocol, 2026-07-28 specification (RFC 9207 issuer validation; client credential binding).

This article is for information and does not constitute legal or security advice.