Pillar guide · MCP Governance
MCP Governance: Policy and Operating Model
Security asks whether a tool call can be exploited. Governance asks who decided it was allowed, on what basis, whether that decision is still current, and who can prove it. This guide covers the four domains, the risk tiers, the operating model and the first thirty days.
The short answer
MCP governance is the set of decisions, controls and records that determine which AI agents may invoke which Model Context Protocol tools, under what conditions, with whose approval, and with what evidence retained. It has four domains — estate, identity, policy and evidence. Security keeps a single tool call from being exploited; governance keeps the organisation able to answer what its agents are permitted to do, who permitted it, and what they actually did.
Governance is not security, and the difference is operational
The two words get used interchangeably, and the conflation costs real money. MCP security is largely an engineering property of a system: can this tool call be exploited, is the credential scoped, is the input validated. You can improve it by writing better code. MCP governance is an organisational property: can this company state which of its agents may call which tools, show who decided that, show that the decision is still current, and produce the record afterwards. You cannot code your way to it, because most of what is missing is a decision nobody has made.
The practical test is a question. Ask an engineer which tools your finance agent can reach; you will usually get an answer. Ask who approved the third of them, when, and against what criteria; the answer is almost always silence. The first question is security’s. The second is governance’s, and the second is the one an auditor, a customer security review or an incident post-mortem actually asks.
Can this call be abused? Is the token short-lived, the parameter bounded, the output distrusted? Answered by the team that built the server, verified by testing.
Who decided this agent may make this call, on what basis, is that still true, and where is the record? Answered by an operating model, verified by producing evidence on demand.
Why it starts mattering at server three
One MCP server is a project. Two are a pattern. Three are an estate, and an estate is the point at which nobody’s memory is the inventory any more. The threshold is not arbitrary: with two servers a single engineer can hold the whole permission surface in their head and answer questions about it correctly. At three or four, held knowledge becomes stale knowledge, and stale knowledge is indistinguishable from no knowledge when something goes wrong.
The second threshold is narrower and arrives sooner: the first server that can write. A read-only estate has a bounded worst case — disclosure. A single mutating tool changes the worst case to action: money moved, records deleted, messages sent under your name. Governance should exist before the first write tool ships, not after.
The shape of MCP sprawl
Sprawl is the accumulation of servers, tools and credentials faster than anyone inventories or retires them. It is almost always additive. Teams add a server to unblock a piece of work, the work ships, the server stays, and the capability it granted is never revisited — not because anyone decided to keep it, but because removing things requires someone to own the question. Sprawl is not a security failure. It is a governance vacuum with a security consequence.
The four domains you govern
Everything worth calling MCP governance falls into four domains. They are ordered by dependency: you cannot write useful policy without identity, and you cannot establish identity for an estate you have not inventoried.
Know what exists, and who owns it
A current list of every MCP server, every tool each one advertises, every agent connected to it, and a named human owner per server. The list has to be generated rather than maintained by hand — a hand-maintained inventory is accurate on the day it is written. Include third-party servers explicitly: a community server installed to save an afternoon is a supply-chain decision with your credentials attached.
One identity per agent, revocable in one place
Entitlement is meaningless if you cannot say who is asking. A shared API key across four agents means you can neither grant differently nor revoke individually, and every log line names the key rather than the actor. Distinct, short-lived, verifiable identity per agent is the hinge the other three domains turn on. This is the domain that most often blocks the rest, because it usually requires the estate to stop reading credentials from local configuration files.
Rules that bind identity to permitted action
Which identity may call which tool, within what parameter bounds, at what rate, and when a human has to agree first. Note that tool-level permission alone is insufficient: a payment tool with no amount ceiling is a materially different tool at 10 and at 10 million. Policy has to reach parameter values, or it is only pretending to be a control.
One record per attempt, including refusals
Caller identity, server, tool, parameters, policy decision, approver where applicable, result and timestamp — written where enforcement happens rather than reassembled later from three partial logs. Denials matter more than successes: a rising denial rate against one tool is either a policy that is wrong or an agent that is being steered, and both are things you want to find in a dashboard rather than in a headline.
Risk tiers policy can key on
Policy needs something to key on other than the tool’s name. Three tiers, assigned by consequence rather than by which system the tool talks to, are enough to run an estate. Tier the tool, not the server: one server frequently holds tools from two different tiers, and tiering at server granularity forces you to govern the safe ones as if they were dangerous.
| Tier | Test | Default posture |
|---|---|---|
| T1 · Reversible | Reads, or writes you could undo in a minute without telling anyone. | Allow by entitlement. Rate-limit per identity. Log. |
| T2 · Consequential | Costs money, is visible outside the company, or is awkward but survivable to reverse. | Allow within bounded parameters. Approval above a threshold. Alert the owner. |
| T3 · Irreversible | Destroys data, moves funds, changes access, or cannot be taken back at all. | Deny by default. Named identities only, human approval per call, dual review on policy change. |
Two rules keep tiering honest. Tier by the widest thing the tool’s credential permits, not by the narrowest thing it is meant to do — an order-lookup tool holding a credential that can also delete orders is T3 whatever its description says. And re-tier on change: a tool that gains a parameter can change tier without anyone editing its name.
Policy as code, because the review is the point
Entitlements declared in a console are invisible: there is no diff, no reviewer, no reason recorded, and no way to tell staging from production except by looking at both. Entitlements declared in version-controlled files get four properties for free that governance otherwise has to manufacture — a diff showing exactly what changed, a named reviewer who agreed, a commit message saying why, and an identical declaration deployable to every environment.
Declare the grant, not the exception
Deny by default and enumerate what is allowed. A policy file listing prohibitions is a policy that grows every time an attacker is inventive.
Bound the values
Amount ceilings, allowed recipients, row limits, date ranges. Parameter bounds are where policy stops being a list of names.
Expire the grant
Give every T2 and T3 entitlement an end date. Renewal forces the review that voluntary cleanup never gets around to.
Require dual review at T3
The policy that grants irreversible capability should need two humans, the same way the irreversible action itself does.
Test policy like code
Assertions that a named identity is refused a named tool. Otherwise a refactor silently widens what an agent can reach.
Keep the reason in the commit
Six months later the only question anyone asks about an entitlement is why it exists. Git answers it if you write it down.
Five stages of maturity
Not a scorecard to aspire to in full. Most organisations should be at stage 2 within a quarter of connecting their third server, and stage 3 before any T3 tool is exposed. Stage 4 earns its cost only once agents act without a human watching each run.
-
Stage 0 · Unmanaged
Configuration lives on laptops
Servers are wired up per developer machine with static keys. Nobody holds a list, and the honest answer to “what can our agents do?” is that it depends whose machine you ask about.
-
Stage 1 · Inventoried
A list exists and has an owner
Every server and tool is enumerated, tiered and attributed to a named owner. Enforcement is still per server and still inconsistent, but the estate is now a known quantity rather than a rumour.
-
Stage 2 · Mediated
All traffic passes one enforcement point
Agents reach servers through a gateway rather than directly. Identity is per agent, issued centrally and revocable in one place. This is the stage that makes every later control cheap instead of impossible.
-
Stage 3 · Policied
Entitlements are reviewed like code
Grants live in version control with bounds, expiry and named reviewers. Approval thresholds are set by risk tier rather than by instinct, and policy changes are testable.
-
Stage 4 · Evidenced
Denials drive the policy
Every attempt is queryable in one place, retained on a stated schedule, and reviewed. Denial trends feed policy revision, and producing a week of agent activity is one query rather than a project.
Who owns what
MCP governance fails most often for an organisational reason, not a technical one: the capability is granted by engineering, the consequence lands on a business owner, and neither believes the decision was theirs. Four roles, with the boundary stated, resolve most of it.
Platform engineering — owns the enforcement point
Runs the gateway, issues agent identity, keeps the inventory generated rather than typed, and makes the policy mechanism work. Does not decide what any given agent should be allowed to do.
Security — owns tiering and policy review
Assigns risk tiers, sets the default posture per tier, reviews T3 policy changes, and watches denial trends. Reviews grants; does not author them.
System owner — owns exposure and entitlement
The business owner of the underlying system decides which of its tools are exposed to agents at all and which identities get them. This is the role most estates leave vacant, and its vacancy is why capability accumulates.
Audit or compliance — owns whether the evidence answers
States the questions the record must be able to answer and the retention period, then tests it by asking. The test is a rehearsal, not an inspection: it is much cheaper to discover a gap in the log during a drill than during an incident.
Built on this thinking
Two products cover mediation and policy
Barzel Central Gateway is the single mediation point that gets you to stage 2 — one identity model, one route, one audit stream in front of every server. BarzelVault is the action-level permission, approval and audit layer that makes stages 3 and 4 practical.
The first thirty days
A governance programme that begins with a framework document stalls. One that begins with an inventory produces an argument in week one, which is the point. Four weeks, one deliverable each.
- W1Enumerate. Every server, every tool, every connected agent, every credential and its scope. Generate it by querying, not by asking people. Expect the count to surprise someone senior.
- W2Tier and name. Assign T1/T2/T3 per tool by widest permitted action, and put a named owner against every server. Anything that cannot be attributed to an owner is a candidate for removal, not for a meeting.
- W3Mediate the T3 path. Do not attempt the whole estate. Route every irreversible tool through one enforcement point with per-agent identity, and leave the read-only servers where they are for now.
- W4Rehearse the question. Pick a day last month and produce every tool execution from it, with parameters and decisions, in one query. Whatever you cannot answer is your actual backlog — and it is a shorter, better-argued backlog than any framework document would have given you.
Nothing here requires a committee, a budget cycle or a vendor decision. It requires someone to own the four deliverables for a month.
Frequently asked questions
What is MCP governance?
MCP governance is the set of decisions, controls and records that determine which AI agents may invoke which Model Context Protocol tools, under what conditions, with whose approval, and with what evidence retained. It covers four domains: the estate of servers and tools, the identity of each calling agent, the policy binding identity to permitted action, and the evidence that the policy was applied.
What is the difference between MCP security and MCP governance?
Security asks whether a given tool call can be exploited. Governance asks who decided the call was allowed, on what basis, whether that decision is still current, and who can prove it afterwards. Security is largely an engineering property of one system; governance is an organisational property spanning every system an agent can reach.
When does an organisation need MCP governance?
In practice at the third server, or at the first server that can write. Two read-only servers can be reasoned about in someone’s head. Three or more, or any server that mutates state, exceeds what informal memory tracks reliably.
What is MCP sprawl?
MCP sprawl is the accumulation of servers, tools and credentials faster than anyone inventories or retires them. It is usually additive: teams add servers to unblock work, rarely remove them, and never total up the capability surface they have assembled.
What is policy as code in an MCP context?
Declaring agent entitlements — which identity may call which tool, within what parameter bounds, at what rate, with what approval — in version-controlled files rather than in a console. The review happens in a pull request, the history is the audit trail, and the same declaration deploys to every environment.
Who owns MCP governance in an organisation?
Four roles with distinct duties: platform engineering owns the enforcement point, security owns risk tiering and policy review, the business owner of each system owns which tools are exposed and who may use them, and internal audit or compliance owns whether the evidence answers the questions they will be asked.
Can MCP governance be enforced inside each MCP server?
It can, but the same controls implemented independently in each server drift apart, and the weakest implementation sets your real posture. A single mediation point gives one identity model, one policy set, one audit stream and one place to revoke.
How does MCP governance map to the NIST AI RMF or ISO/IEC 42001?
Both frameworks require you to know what your AI systems can do, to manage the risk of those capabilities, and to keep records demonstrating the management. An MCP estate inventory, risk-tiered entitlements and a complete execution log are the artefacts those requirements ask for, expressed at the level where agents actually act.
Go deeper on MCP governance
Twenty-four articles that each take one section of this guide and work it through in detail.
Sources and further reading
Primary specifications and standards this article relies on. The risk tiers, maturity stages and thirty-day sequence are our own model, drawn from building and operating five MCP servers — where a claim is judgement rather than something a standard states, the text says so.
- 01 · MCP project Model Context Protocol — specification ↗ The normative spec, including capability negotiation and the authorization sections governance has to build on.
- 02 · MCP project Model Context Protocol — official documentation ↗ Primary source for protocol structure, transports and tool definitions.
- 03 · NIST AI Risk Management Framework (AI RMF 1.0) ↗ The Govern, Map, Measure and Manage functions that the four domains here are an MCP-level expression of.
- 04 · ISO/IEC ISO/IEC 42001:2023 — AI management systems ↗ The certifiable management-system standard whose documentation requirements the evidence domain satisfies.
- 05 · OWASP GenAI Security Project OWASP GenAI LLM Top 10 (2026) ↗ Current consensus risk list; excessive agency is the entry this guide’s tiering exists to bound.
- 06 · NIST SP 800-207: Zero Trust Architecture ↗ Where the policy-enforcement-point and policy-decision-point separation used in stage 2 comes from.
- 07 · European Union Regulation (EU) 2024/1689 — Artificial Intelligence Act ↗ Consulted for its record-keeping and human-oversight obligations. This article is not legal advice.
Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.
Try the enumeration step against a real server
Week one asks you to enumerate rather than ask. To see what that produces — the full tool list a server advertises, unfiltered — point a client at a public one. Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run — no signup, no key, 54 tools. A tools/list call takes about a minute and shows you the shape an inventory row needs. Setup is in the reference.
Reference documentation
Want the specification rather than the argument?
Transport, auth, capability enumeration, policy evaluation and refusal semantics for every Barzel server.
Dev free · 10,000 calls/mo · paid plans from $199/mo on the MCPize marketplace
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.