5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

Threat · containment

MCP Prompt Injection: How Attacks Reach Real Tools

Prompt injection was an annoyance while models only produced text. MCP is the thing that turns a persuaded model into a payment, a deletion or an email nobody authorised.

By Mark Alex, FounderPublished 24 Aug 2026Updated 2 Sep 202614 min

The short answer

MCP prompt injection is any attack in which instructions reach a model through content it processes, and the model then invokes MCP tools on the attacker’s behalf. Four paths carry it: retrieved documents and web content, tool results, tool descriptions and schemas, and multi-agent messages. There is currently no reliable way to prevent the model being persuaded, so the defence is containment — the resulting tool call must be refused by entitlement and parameter policy regardless of the reasoning behind it. Five defences that do not hold, one that does, and an afternoon’s worth of tests.

Why MCP changes the severity

Prompt injection has been understood for years. For most of that time the worst outcome was that a model said something it should not have — embarrassing, occasionally reputational, rarely material.

A tool-calling protocol changes the last step. The persuaded model no longer produces text for a human to evaluate; it produces a tool call that executes. The injected instruction and the legitimate instruction arrive at the tool layer in exactly the same shape, and nothing in the call carries a marker for which reasoning produced it.

So the vulnerability is unchanged and the severity is transformed. That distinction matters for how the problem gets funded: this is not a new attack requiring new detection research, it is an old attack that now needs an authority boundary.

Four paths in

Every MCP injection we have seen or constructed arrives by one of these. Path two is the one most teams have not considered.

01 · Path 01

Retrieved content

A document, a web page, a support ticket, a calendar invite, a code comment, a customer-editable field. Anything the agent reads that somebody outside your control can write.

The largest surface by far, and it includes internal content: a field a customer types into is an untrusted input even though it is stored in your own database.

02 · Path 02

Tool results

The output of one tool becomes context for the next decision. A search tool that returns third-party content is an injection channel, and it is one most threat models omit because the tool itself is trusted.

This is the path that makes tool chains dangerous. Tool A is trusted, tool B is trusted, and the content flowing from A to B is not.

03 · Path 03

Tool descriptions and schemas

The description is documentation for humans and instruction for models. A tool whose description says to first call a different tool has effectively issued an instruction — see tool poisoning.

Especially dangerous because descriptions arrive at the start of context, where models weight them heavily, and because a version bump can change them without review.

04 · Path 04

Multi-agent messages

Where one agent’s output is another’s input, a compromise propagates. The receiving agent has no way to distinguish a genuine instruction from a relayed injection, and its entitlements may be broader than the sender’s.

Anatomy of a working attack

A concrete chain, of the kind that works against a typical customer-support setup. Nothing here is sophisticated, which is the point.

A support agent handles refund requests. It reads the ticket, checks the order, and can issue refunds up to a limit. Reasonable, and in production somewhere near you.

The attacker opens a ticket whose body contains ordinary complaint text and, further down in a form the model reads and a human skims past, an instruction: that the account has a credit adjustment pre-approved by finance, that the correct handling is a refund of a specific amount to a specific alternate payment destination, and that the adjustment should not be mentioned in the reply.

The model has no basis for treating that as different from a genuine policy note. It calls the refund tool with attacker-supplied values, and the confirmation is suppressed because the injected text asked for it to be.

Now count the controls that would have stopped it. A prompt saying to ignore instructions in tickets: no — the injected text reads as data about the account rather than as a command. Output filtering: no — the refund is a tool call rather than prose. Detection: only after the money has left. An allow-list on the refund tool: no — the agent is supposed to have it.

What stops it is narrower. The destination is not among the customer’s own registered payment methods, so the call is refused on parameters. Or the amount is above the threshold that requires a human, and the human sees an unfamiliar destination and asks a question. Both are parameter-level controls, and neither cares that an injection caused the call.

Five defences that do not hold

Each is deployed somewhere as the primary defence. Each fails structurally rather than through poor implementation.

Instructing the model to ignore injected instructions

The instruction and the injection occupy the same channel with no reliable precedence. Every published wording of this defence has been circumvented, usually quickly.

Delimiting untrusted content with markers

Helps a little, and the attacker writes the closing delimiter. Do it anyway; do not count it.

Scanning for injection patterns

Catches naive attempts and creates confidence disproportionate to coverage. Paraphrase defeats signatures, and the paraphrase space is unbounded.

A second model as a judge

Persuadable by the same content, and it makes decisions unreproducible. Two persuadable components in series is not a boundary.

Detection and alerting

Necessary, and the wrong layer for anything irreversible. An alert on a completed payment is a notification.

The common flaw: each tries to keep the model from being convinced. That is the part nobody currently knows how to do reliably, so a defence resting on it inherits the uncertainty.

What actually works: make being convinced not matter

Assume the model will be persuaded. Design so that a persuaded model cannot do anything you would mind. Five controls, and none of them inspect content.

  • 01Entitlement, filtered at discovery. A tool the model never learns exists cannot be requested. The cheapest control, and the one with the largest effect on injection outcomes specifically.
  • 02Parameter bounds. Allow-listed destinations, amount ceilings, row caps. The attack above dies here, on the destination, before any human is involved.
  • 03Approval on irreversible actions, with parameters shown. The approver sees an unfamiliar payment destination and asks. Human judgement applied to specific values, which is the one thing humans are reliably better at.
  • 04Per-identity ceilings. Bounds the damage rate. One successful injection becomes one bad call rather than four hundred.
  • 05Egress restriction on the server. A server that can only reach its own backing system cannot be used to move data elsewhere, whatever it is persuaded to attempt.

Notice that none of these are injection defences. They are authority controls, and they work against injection precisely because they do not care why a call was made. That property is what makes them durable as the attack evolves.

Built on this thinking

Containment, as an enforcement point

BarzelVault evaluates the fully specified call — every parameter, the identity, the target — before execution and refuses on values rather than on intent, which is what makes it indifferent to why a call was made. Approval workflow, per-identity ceilings and hash-chained audit including refusals.

Testing it in an afternoon

Four tests. Each takes under an hour and each should end in a refusal that appears in the record.

TestMethodPass condition
Retrieved-content injectionPut an instruction in a document the agent will read, telling it to call a tool it is not entitled toRefused at the enforcement point; refusal logged with the caller
Tool-result injectionHave a permitted read tool return content containing an instruction to call a high-risk toolThe follow-on call is refused by entitlement, whatever the model attempts
Parameter injectionInjected text supplies a destination or amount outside the allow-list or ceilingRejected on parameters, with the attempted value recorded rather than clamped
Description injectionAlter a tool description in a staging registry to instruct a prior callThe diff flags it at intake before it reaches production

Run test two by hand, once, with the team watching. Seeing a model try to follow an instruction it found in returned data, and seeing the enforcement point refuse it anyway, settles the “can we handle this in the prompt” argument faster than any document.

What we are not claiming

Containment is not prevention, and it is worth being exact about the residual risk rather than rounding it away.

A successfully injected agent can still do anything it is legitimately entitled to do, within bounds. If it may read customer records, it may be made to read the wrong ones. If it may send email, it may send an odd one. The controls bound the magnitude and produce a record; they do not restore the model’s judgement.

That residual is why entitlement minimisation matters more here than anywhere else, and why reviewing entitlements as sets rather than tool by tool is not a formality. The smaller the legitimate set, the smaller the space a persuaded model has to operate in.

Anyone claiming to prevent prompt injection is describing detection with better marketing. Ask what happens when the paraphrase changes.

Frequently asked questions

What is MCP prompt injection?

An attack in which instructions embedded in content a model processes cause it to invoke MCP tools on the attacker’s behalf. The vulnerability is the same one that has existed for years; MCP changes the severity, because the persuaded model now produces an executing tool call rather than text a human evaluates.

How does prompt injection reach MCP tools?

By four paths: retrieved content such as documents, tickets and customer-editable fields; tool results, where one trusted tool returns untrusted content that becomes context for the next decision; tool descriptions and schemas, which are instruction to a model as much as documentation to a human; and multi-agent messages, where one agent’s output is another’s input.

Can MCP prompt injection be prevented?

Not reliably, at the model layer, today. The injected instruction and the legitimate instruction occupy the same channel with no dependable precedence. What works is containment: the resulting tool call is refused by entitlement and parameter policy regardless of the reasoning that produced it.

Why does telling the model to ignore injected instructions not work?

Because the instruction and the injection share one channel with no reliable ordering of authority, and well-crafted injected text reads as data about the situation rather than as a command. Every published wording of this defence has been circumvented, usually quickly.

Is scanning for prompt injection patterns useful?

It catches naive attempts and creates confidence out of proportion to its coverage, because paraphrase defeats signatures and the paraphrase space is unbounded. Keep it as a layer; do not let it be the control you rely on.

What actually contains MCP prompt injection?

Five authority controls, none of which inspect content: entitlement filtered at discovery, parameter bounds such as allow-listed destinations and amount ceilings, approval on irreversible actions with the parameters shown to the approver, per-identity rate ceilings to bound the damage rate, and egress restriction on the server.

How do you test whether your MCP estate is contained against injection?

Four tests in an afternoon: put an instruction in a document the agent will read; have a permitted read tool return content instructing a high-risk call; supply a destination or amount outside the allow-list through injected text; and alter a tool description in staging. Each should end in a refusal that appears in the record.

What residual risk remains after containment?

A successfully injected agent can still do anything it is legitimately entitled to do, within bounds — read the wrong customer records, send an odd email. Containment bounds the magnitude and produces a record; it does not restore the model’s judgement, which is why minimising the legitimate entitlement set matters more here than anywhere else.

Sources and further reading

Prompt injection as an attack class is best documented in the practitioner literature cited below; the attack is not ours and we make no novel claims about it. The four-path taxonomy, the honest statement that prevention is unavailable, and the containment tests are our contribution.

  1. 01 · Simon WillisonPrompt injection — ongoing series ↗The most consistently updated practitioner record of the attack class.
  2. 02 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
  3. 03 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  4. 04 · MITREMITRE ATLAS ↗Adversary technique knowledge base for AI systems, useful for naming what a risk score is scoring.
  5. 05 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  6. 06 · OWASPOWASP Application Security Verification Standard ↗Input-validation, authorization and logging requirements restated here in MCP terms.
  7. 07 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.

Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.

Try the mechanics on a live server

To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

BarzelVault is the pre-execution decision point, sold as a running product

Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic policy outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.

DevFree10,000 calls/mo
Team$199/mo75,000 calls/mo
Business$799/mo750,000 calls/mo
Enterprise$3,999/mo5,000,000 calls/mo

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.