5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

Multi-agent · trust

Multi-Agent Security: Trust Between Agents

Two agents that trust each other are one agent with two identities and a worse audit trail. The interesting question is what a chain of agents is permitted to do that no single one of them is.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202616 min2,952 words

The short answer

Multi-agent security governs what happens when agents invoke other agents. Its central rule is that authority must narrow monotonically along a chain: a delegating agent may never grant a downstream agent more scope than it holds itself. Without this, a chain composes into an authority no single agent was granted, and one compromised agent becomes a compromised chain — because a receiving agent cannot distinguish a genuine instruction from a relayed injection. Four trust models, the amplification problem, cross-agent loop detection, and what a chain-level record must contain.

Key takeaways

  1. 01Authority must narrow at every hop. A chain whose scope grows is a privilege-escalation path with an architecture diagram.
  2. 02A receiving agent cannot tell a genuine instruction from a relayed injection. Treat inter-agent messages as untrusted content.
  3. 03The chain’s effective authority is the union of what it can reach, not the minimum. Review chains, not agents.
  4. 04Loops across agents are invisible to per-agent rate limits — A calls B calls A is two well-behaved agents and one runaway system.
  5. 05Every hop needs the originating human principal preserved, or the last agent’s service account owns every action.
  6. 06One chain id across every call, or a nine-hop failure is nine unrelated incidents.

What actually changes

A single agent calling tools has one identity, one entitlement set and one audit trail. Reasoning about it is tractable: you can enumerate what it can reach and bound each reach.

Add a second agent that the first can invoke, and three things change at once.

The effective authority is now a union. Agent A can reach whatever A can reach, plus whatever it can persuade B to do. If B has tools A does not, A now effectively has them — through a hop that no entitlement review examined, because the reviews looked at A and B separately.

Trust is transitive and undeserved. B receives an instruction from A and has no basis to distinguish “A decided this” from “A read a document that told it to say this”. Every injection that reaches A now reaches B, with A’s credibility attached.

Accountability dilutes at each hop. Unless the originating human is carried forward explicitly, the action lands on the last agent’s service account, and the audit record says a robot did it.

Key facts

  • ▸A chain’s effective authority is the union of its members’ reach, not the minimum.
  • ▸A receiving agent cannot distinguish a genuine instruction from a relayed injection.
  • ▸Without explicit propagation, accountability terminates at the last agent’s service account.

The narrowing rule

One rule carries most of the weight: a delegating agent may never grant more than it holds. Each hop’s scope is equal to or narrower than the hop before it, and the chain’s authority is bounded by its first member rather than by its most privileged one.

This sounds obvious and is routinely violated, because the natural implementation is for each agent to authenticate with its own service credentials. Agent A calls agent B; B authenticates as itself; B’s entitlements apply. If B is a general-purpose agent with broad reach — and general-purpose agents usually are — then A has just escalated by making a phone call.

The implementation that holds is delegation rather than re-authentication. A presents a token that names the originating human, A’s own identity, and a scope no wider than A’s own. B receives a token it cannot widen, acts within it, and the record shows the full chain.

Token exchange is the standard mechanism: each hop exchanges its inbound token for a downstream one, narrowed and audience-bound. Where that is unavailable, the enforcement point must compute the intersection of the chain’s scopes and apply it — which is more fragile but better than nothing.

Worked example · A chain that escalates, and the same chain fixed

A triage agent handles inbound support requests. It can read tickets and customer records but cannot issue refunds. A general operations agent can do almost anything.

broken — hop 1triage-agent (scope: tickets:read, customers:read) receives a ticket containing an injected instruction
broken — hop 2triage-agent delegates to ops-agent: “issue a refund for order 5512”
broken — hop 3ops-agent authenticates as itself (scope: near-unlimited), issues the refund
resultA read-only agent caused a refund. No control was violated: each agent acted within its own entitlements
fixed — hop 1triage-agent receives the same ticket
fixed — hop 2Delegation carries a token: subject=customer-service-user, actor=triage-agent, scope = intersection of triage-agent’s scope and the request = tickets:read, customers:read
fixed — hop 3ops-agent receives the narrowed token. refunds:issue is not in scope. The call is refused and logged against the originating human and the full chain

The fix costs one token exchange per hop and it converts an escalation path into a logged refusal. Note that the injection still succeeded at persuading the model — what changed is that being persuaded stopped being sufficient.

Four trust models

How much a receiving agent should believe about an instruction from another agent. Choose deliberately; the default is the first row and it is rarely right.

ModelReceiving agent treats the message asRight whenRisk
Implicit trustA legitimate instruction from a trusted peerNever, for consequential actionsOne compromised agent compromises the chain
Authenticated but untrustedA verified sender, untrusted contentDefault for most estatesRequires policy at every hop, which is the correct cost
Delegated authorityAn instruction carrying the originating human’s identity and a narrowed scopeAny chain touching consequential actionsNeeds token exchange support
MediatedAll inter-agent traffic passes an enforcement point that evaluates the chainHigh-consequence estatesAdds a hop; the enforcement point sees everything

The distinction between rows one and two is the whole design decision. Authenticated is not the same as trusted. Verifying that a message genuinely came from agent A tells you nothing about whether A was reasoning or relaying. Treat the sender as authenticated and the content as untrusted, exactly as you would content from a web page — because functionally it may be the same content.

Row four is worth the extra hop wherever a chain can reach money, personal data or external parties. It is also the only model in which chain-level policy — cumulative magnitude, composition limits, loop detection — is expressible at all.

Amplification and cross-agent loops

A single request that fans out across agents multiplies in ways that are hard to see from any one agent’s perspective.

Three failure shapes, none of which per-agent controls detect.

Shape 01

Fan-out amplification

One request becomes three delegations, each becoming three more. Twenty-seven agent invocations and several hundred tool calls from one ticket. Each agent behaved reasonably; the system did not.

Bound the fan-out explicitly: maximum delegation depth, maximum delegations per request, and a total call budget for the chain rather than per agent.

Shape 02

Cross-agent loops

A delegates to B, B delegates back to A with a slight variation. Neither exceeds its own rate limit. The system runs until something external breaks or the budget is exhausted.

Per-agent rate limits are structurally blind to this. Detection needs the chain id: count calls per chain, and refuse when a chain revisits an agent more than a small number of times.

Shape 03

Cost concentration

One agent’s delegation pattern consumes the shared budget, starving unrelated workflows. The expensive agent is not the one that caused it.

Attribute cost to the chain and to the originating request, not only to the agent that made the call — otherwise the invoice blames the wrong component.

The unifying control is a chain budget: a total call count, spend and depth allocated at the entry point and decremented across every hop. It is the multi-agent equivalent of a per-identity ceiling, and without it the estate has no bound on what one request can consume.

Built on this thinking

Chains need one enforcement point that sees every hop

Barzel Central Gateway mediates agent traffic, preserves the originating principal across hops and applies per-identity ceilings that a chain budget builds on. BarzelVault evaluates the action at each hop against the narrowed scope, so a relayed instruction is refused on values regardless of which agent sent it.

What a chain record must contain

Per-call records remain necessary. Six chain-level fields make a multi-agent failure reconstructable rather than forensic.

FieldWhy
Chain idThe join across every call in one execution. Without it, a nine-hop failure is nine unrelated events
Originating human principalCarried to every hop. Otherwise the action belongs to the last agent’s service account
Delegation depth and parent agentReconstructs the tree, and makes fan-out visible
Scope at this hopProves narrowing held — or shows exactly where it did not
Delegation reasonWhat the calling agent said it wanted, which is where a relayed injection becomes visible on review
Chain budget consumedCalls and spend so far, so the ceiling is auditable and the runaway is attributable

The delegation reason field is the one worth arguing for. When a chain does something unexpected, the question is which hop introduced the intent, and the reason text at each hop is the only place that is legible. It is also, uncomfortably, where you will find the injected instruction quoted verbatim.

Five mistakes worth naming

Agents re-authenticating as themselves

The escalation path. Each hop applies its own entitlements, so a chain reaches whatever its most privileged member can.

Trusting authenticated senders

Verifying who sent a message says nothing about whether they were reasoning or relaying.

Per-agent rate limits only

Blind to cross-agent loops, which are two well-behaved agents forming one runaway system.

No chain id

Turns a multi-hop failure into an archaeology exercise against timestamps.

Reviewing agents individually

The authority that matters is the chain’s union, and no per-agent review sees it.

What this does not settle

Multi-agent standards are early. Token exchange and mediated enforcement work today, but agent-to-agent interoperability protocols are still forming, and an estate built now will likely need to adapt. Design so the enforcement point is the thing that changes rather than every agent — that is the argument for mediation independent of its security properties.

Narrowing also has a real cost: a chain can only do what its weakest link permits, which means legitimate delegations get refused and someone has to widen the first agent’s scope. That pressure is the mechanism by which narrowing quietly stops being enforced. Watch for entry-point scopes growing over time; it is the signal that the rule is being circumvented rather than applied.

And nothing here resolves whether multi-agent architecture is the right choice. Many systems described as multi-agent would be simpler, cheaper and safer as one agent with more tools, or as an ordinary workflow with a model in one step. The security cost of a chain is real, and it should be paid only where the architecture earns it.

Frequently asked questions

What is multi-agent security?

The discipline of governing what happens when AI agents invoke other agents. Its central rule is that authority must narrow monotonically along a chain: a delegating agent may never grant a downstream agent more scope than it holds itself.

Why is agent-to-agent delegation a privilege escalation risk?

Because the natural implementation has each agent authenticate with its own credentials. A read-only agent delegating to a general-purpose agent effectively gains that agent’s reach, and no entitlement review catches it because the reviews examined each agent separately.

Can a receiving agent trust an instruction from another agent?

No. It cannot distinguish a genuine instruction from a relayed injection, because the sending agent may simply be repeating what a document told it. Treat inter-agent messages as authenticated but untrusted — verify the sender, distrust the content.

How should authority propagate across agent hops?

By delegation rather than re-authentication. Each hop presents a token naming the originating human, the calling agent, and a scope no wider than the caller’s own — typically via OAuth token exchange, with each hop narrowing and audience-binding the token it passes on.

What are the trust models for multi-agent systems?

Four: implicit trust, which is never right for consequential actions; authenticated but untrusted, the sensible default; delegated authority carrying the originating identity and a narrowed scope; and mediated, where all inter-agent traffic passes an enforcement point that can evaluate the whole chain.

Why do per-agent rate limits miss multi-agent loops?

Because A calling B calling A is two agents each staying within their own limits while the system as a whole runs away. Detection requires a chain id and a chain-level budget, counting calls per execution rather than per identity.

What is a chain budget?

A total call count, spend and delegation depth allocated at a chain’s entry point and decremented across every hop. It is the multi-agent equivalent of a per-identity ceiling, and without it there is no bound on what one request can consume.

What should a multi-agent audit record contain?

Six chain-level fields beyond the per-call record: chain id, originating human principal, delegation depth and parent agent, scope at this hop, the delegation reason the calling agent gave, and chain budget consumed so far.

Which chain field is most useful during an incident?

The delegation reason at each hop. When a chain does something unexpected, the question is which hop introduced the intent, and the reason text is the only place that is legible — it is also where a relayed injection appears verbatim.

Is multi-agent architecture always worth the security cost?

No. Many systems described as multi-agent would be simpler, cheaper and safer as one agent with more tools, or as an ordinary workflow with a model in a single step. The chain’s security cost is real and should be paid only where the architecture earns it.

Glossary

Multi-agent system
An arrangement in which agents invoke other agents, delegating subtasks rather than calling tools directly.
Authority propagation
How an agent’s permissions and the originating human’s identity are carried — or lost — across a delegation to another agent.
Monotonic narrowing
The requirement that each hop in a delegation chain holds scope equal to or narrower than the hop before it.
Amplification
The multiplication of tool calls and cost as one request fans out across several agents, each of which may fan out further.
Chain id
A single identifier attached to every call in one multi-agent execution, without which a failure cannot be reconstructed.

Sources and further reading

The delegation mechanics draw on OAuth token exchange and zero-trust principles, cited below; the emerging agent-interop work is referenced for context. The narrowing rule, the trust models and the amplification analysis are our own.

  1. 01 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  2. 02 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
  3. 03 · A2A projectAgent2Agent (A2A) Protocol ↗Emerging interoperability layer for agent-to-agent messaging and trust.
  4. 04 · NISTNIST SP 800-207 — Zero Trust Architecture ↗The policy decision point / policy enforcement point split this architecture borrows directly.
  5. 05 · IETFRFC 8693 — OAuth 2.0 Token Exchange ↗The mechanism for preserving caller authority across a gateway hop.
  6. 06 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  7. 07 · MITREMITRE ATLAS ↗Adversary technique knowledge base for AI systems, useful for naming what a risk score is scoring.
  8. 08 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.

Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). Multi-Agent Security: Trust Between Agents. Real Biz Digital. https://realbizdigital.net/insights/multi-agent-security/

Try the mechanics on a live server

To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

Barzel Central Gateway is this layer, sold as a running product

Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so an evaluation costs an afternoon rather than a purchase order.

PlanPriceIncludedRight for
CommunityFree1,000 tool calls/mo · full policy engine, registry, auditEvaluating the estate, or a single team proving the path works
Starter$10/mo10,000 calls/moOne or two production agents against a handful of servers
Team$79/mo100,000 calls/moA platform team governing an estate of 5–20 servers
Business$149/mo250,000 calls/moEstate-wide governance with SIEM evidence and multi-team routing

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.