5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

Pillar guide · MCP Governance

MCP Governance: Policy and Operating Model

Security asks whether a tool call can be exploited. Governance asks who decided it was allowed, on what basis, whether that decision is still current, and who can prove it. This guide covers the four domains, the risk tiers, the operating model and the first thirty days.

By Mark Alex, Founder Published 24 Aug 2026 Updated 2 Sep 2026 18 min read

The short answer

MCP governance is the set of decisions, controls and records that determine which AI agents may invoke which Model Context Protocol tools, under what conditions, with whose approval, and with what evidence retained. It has four domains — estate, identity, policy and evidence. Security keeps a single tool call from being exploited; governance keeps the organisation able to answer what its agents are permitted to do, who permitted it, and what they actually did.

Governance is not security, and the difference is operational

The two words get used interchangeably, and the conflation costs real money. MCP security is largely an engineering property of a system: can this tool call be exploited, is the credential scoped, is the input validated. You can improve it by writing better code. MCP governance is an organisational property: can this company state which of its agents may call which tools, show who decided that, show that the decision is still current, and produce the record afterwards. You cannot code your way to it, because most of what is missing is a decision nobody has made.

The practical test is a question. Ask an engineer which tools your finance agent can reach; you will usually get an answer. Ask who approved the third of them, when, and against what criteria; the answer is almost always silence. The first question is security’s. The second is governance’s, and the second is the one an auditor, a customer security review or an incident post-mortem actually asks.

Security question

Can this call be abused? Is the token short-lived, the parameter bounded, the output distrusted? Answered by the team that built the server, verified by testing.

Governance question

Who decided this agent may make this call, on what basis, is that still true, and where is the record? Answered by an operating model, verified by producing evidence on demand.

Why it starts mattering at server three

One MCP server is a project. Two are a pattern. Three are an estate, and an estate is the point at which nobody’s memory is the inventory any more. The threshold is not arbitrary: with two servers a single engineer can hold the whole permission surface in their head and answer questions about it correctly. At three or four, held knowledge becomes stale knowledge, and stale knowledge is indistinguishable from no knowledge when something goes wrong.

The second threshold is narrower and arrives sooner: the first server that can write. A read-only estate has a bounded worst case — disclosure. A single mutating tool changes the worst case to action: money moved, records deleted, messages sent under your name. Governance should exist before the first write tool ships, not after.

The shape of MCP sprawl

Sprawl is the accumulation of servers, tools and credentials faster than anyone inventories or retires them. It is almost always additive. Teams add a server to unblock a piece of work, the work ships, the server stays, and the capability it granted is never revisited — not because anyone decided to keep it, but because removing things requires someone to own the question. Sprawl is not a security failure. It is a governance vacuum with a security consequence.

The four domains you govern

Everything worth calling MCP governance falls into four domains. They are ordered by dependency: you cannot write useful policy without identity, and you cannot establish identity for an estate you have not inventoried.

01 · Estate

Know what exists, and who owns it

A current list of every MCP server, every tool each one advertises, every agent connected to it, and a named human owner per server. The list has to be generated rather than maintained by hand — a hand-maintained inventory is accurate on the day it is written. Include third-party servers explicitly: a community server installed to save an afternoon is a supply-chain decision with your credentials attached.

02 · Identity

One identity per agent, revocable in one place

Entitlement is meaningless if you cannot say who is asking. A shared API key across four agents means you can neither grant differently nor revoke individually, and every log line names the key rather than the actor. Distinct, short-lived, verifiable identity per agent is the hinge the other three domains turn on. This is the domain that most often blocks the rest, because it usually requires the estate to stop reading credentials from local configuration files.

03 · Policy

Rules that bind identity to permitted action

Which identity may call which tool, within what parameter bounds, at what rate, and when a human has to agree first. Note that tool-level permission alone is insufficient: a payment tool with no amount ceiling is a materially different tool at 10 and at 10 million. Policy has to reach parameter values, or it is only pretending to be a control.

04 · Evidence

One record per attempt, including refusals

Caller identity, server, tool, parameters, policy decision, approver where applicable, result and timestamp — written where enforcement happens rather than reassembled later from three partial logs. Denials matter more than successes: a rising denial rate against one tool is either a policy that is wrong or an agent that is being steered, and both are things you want to find in a dashboard rather than in a headline.

Risk tiers policy can key on

Policy needs something to key on other than the tool’s name. Three tiers, assigned by consequence rather than by which system the tool talks to, are enough to run an estate. Tier the tool, not the server: one server frequently holds tools from two different tiers, and tiering at server granularity forces you to govern the safe ones as if they were dangerous.

Tier Test Default posture
T1 · Reversible Reads, or writes you could undo in a minute without telling anyone. Allow by entitlement. Rate-limit per identity. Log.
T2 · Consequential Costs money, is visible outside the company, or is awkward but survivable to reverse. Allow within bounded parameters. Approval above a threshold. Alert the owner.
T3 · Irreversible Destroys data, moves funds, changes access, or cannot be taken back at all. Deny by default. Named identities only, human approval per call, dual review on policy change.

Two rules keep tiering honest. Tier by the widest thing the tool’s credential permits, not by the narrowest thing it is meant to do — an order-lookup tool holding a credential that can also delete orders is T3 whatever its description says. And re-tier on change: a tool that gains a parameter can change tier without anyone editing its name.

Policy as code, because the review is the point

Entitlements declared in a console are invisible: there is no diff, no reviewer, no reason recorded, and no way to tell staging from production except by looking at both. Entitlements declared in version-controlled files get four properties for free that governance otherwise has to manufacture — a diff showing exactly what changed, a named reviewer who agreed, a commit message saying why, and an identical declaration deployable to every environment.

Declare the grant, not the exception

Deny by default and enumerate what is allowed. A policy file listing prohibitions is a policy that grows every time an attacker is inventive.

Bound the values

Amount ceilings, allowed recipients, row limits, date ranges. Parameter bounds are where policy stops being a list of names.

Expire the grant

Give every T2 and T3 entitlement an end date. Renewal forces the review that voluntary cleanup never gets around to.

Require dual review at T3

The policy that grants irreversible capability should need two humans, the same way the irreversible action itself does.

Test policy like code

Assertions that a named identity is refused a named tool. Otherwise a refactor silently widens what an agent can reach.

Keep the reason in the commit

Six months later the only question anyone asks about an entitlement is why it exists. Git answers it if you write it down.

Five stages of maturity

Not a scorecard to aspire to in full. Most organisations should be at stage 2 within a quarter of connecting their third server, and stage 3 before any T3 tool is exposed. Stage 4 earns its cost only once agents act without a human watching each run.

  1. Stage 0 · Unmanaged

    Configuration lives on laptops

    Servers are wired up per developer machine with static keys. Nobody holds a list, and the honest answer to “what can our agents do?” is that it depends whose machine you ask about.

  2. Stage 1 · Inventoried

    A list exists and has an owner

    Every server and tool is enumerated, tiered and attributed to a named owner. Enforcement is still per server and still inconsistent, but the estate is now a known quantity rather than a rumour.

  3. Stage 2 · Mediated

    All traffic passes one enforcement point

    Agents reach servers through a gateway rather than directly. Identity is per agent, issued centrally and revocable in one place. This is the stage that makes every later control cheap instead of impossible.

  4. Stage 3 · Policied

    Entitlements are reviewed like code

    Grants live in version control with bounds, expiry and named reviewers. Approval thresholds are set by risk tier rather than by instinct, and policy changes are testable.

  5. Stage 4 · Evidenced

    Denials drive the policy

    Every attempt is queryable in one place, retained on a stated schedule, and reviewed. Denial trends feed policy revision, and producing a week of agent activity is one query rather than a project.

Who owns what

MCP governance fails most often for an organisational reason, not a technical one: the capability is granted by engineering, the consequence lands on a business owner, and neither believes the decision was theirs. Four roles, with the boundary stated, resolve most of it.

Platform engineering — owns the enforcement point

Runs the gateway, issues agent identity, keeps the inventory generated rather than typed, and makes the policy mechanism work. Does not decide what any given agent should be allowed to do.

Security — owns tiering and policy review

Assigns risk tiers, sets the default posture per tier, reviews T3 policy changes, and watches denial trends. Reviews grants; does not author them.

System owner — owns exposure and entitlement

The business owner of the underlying system decides which of its tools are exposed to agents at all and which identities get them. This is the role most estates leave vacant, and its vacancy is why capability accumulates.

Audit or compliance — owns whether the evidence answers

States the questions the record must be able to answer and the retention period, then tests it by asking. The test is a rehearsal, not an inspection: it is much cheaper to discover a gap in the log during a drill than during an incident.

Built on this thinking

Two products cover mediation and policy

Barzel Central Gateway is the single mediation point that gets you to stage 2 — one identity model, one route, one audit stream in front of every server. BarzelVault is the action-level permission, approval and audit layer that makes stages 3 and 4 practical.

The first thirty days

A governance programme that begins with a framework document stalls. One that begins with an inventory produces an argument in week one, which is the point. Four weeks, one deliverable each.

  1. W1Enumerate. Every server, every tool, every connected agent, every credential and its scope. Generate it by querying, not by asking people. Expect the count to surprise someone senior.
  2. W2Tier and name. Assign T1/T2/T3 per tool by widest permitted action, and put a named owner against every server. Anything that cannot be attributed to an owner is a candidate for removal, not for a meeting.
  3. W3Mediate the T3 path. Do not attempt the whole estate. Route every irreversible tool through one enforcement point with per-agent identity, and leave the read-only servers where they are for now.
  4. W4Rehearse the question. Pick a day last month and produce every tool execution from it, with parameters and decisions, in one query. Whatever you cannot answer is your actual backlog — and it is a shorter, better-argued backlog than any framework document would have given you.

Nothing here requires a committee, a budget cycle or a vendor decision. It requires someone to own the four deliverables for a month.

Frequently asked questions

What is MCP governance?

MCP governance is the set of decisions, controls and records that determine which AI agents may invoke which Model Context Protocol tools, under what conditions, with whose approval, and with what evidence retained. It covers four domains: the estate of servers and tools, the identity of each calling agent, the policy binding identity to permitted action, and the evidence that the policy was applied.

What is the difference between MCP security and MCP governance?

Security asks whether a given tool call can be exploited. Governance asks who decided the call was allowed, on what basis, whether that decision is still current, and who can prove it afterwards. Security is largely an engineering property of one system; governance is an organisational property spanning every system an agent can reach.

When does an organisation need MCP governance?

In practice at the third server, or at the first server that can write. Two read-only servers can be reasoned about in someone’s head. Three or more, or any server that mutates state, exceeds what informal memory tracks reliably.

What is MCP sprawl?

MCP sprawl is the accumulation of servers, tools and credentials faster than anyone inventories or retires them. It is usually additive: teams add servers to unblock work, rarely remove them, and never total up the capability surface they have assembled.

What is policy as code in an MCP context?

Declaring agent entitlements — which identity may call which tool, within what parameter bounds, at what rate, with what approval — in version-controlled files rather than in a console. The review happens in a pull request, the history is the audit trail, and the same declaration deploys to every environment.

Who owns MCP governance in an organisation?

Four roles with distinct duties: platform engineering owns the enforcement point, security owns risk tiering and policy review, the business owner of each system owns which tools are exposed and who may use them, and internal audit or compliance owns whether the evidence answers the questions they will be asked.

Can MCP governance be enforced inside each MCP server?

It can, but the same controls implemented independently in each server drift apart, and the weakest implementation sets your real posture. A single mediation point gives one identity model, one policy set, one audit stream and one place to revoke.

How does MCP governance map to the NIST AI RMF or ISO/IEC 42001?

Both frameworks require you to know what your AI systems can do, to manage the risk of those capabilities, and to keep records demonstrating the management. An MCP estate inventory, risk-tiered entitlements and a complete execution log are the artefacts those requirements ask for, expressed at the level where agents actually act.

Go deeper on MCP governance

Twenty-four articles that each take one section of this guide and work it through in detail.

Implementation · 14 min How to Secure an MCP Server → Ten ordered steps from credential scope to verification, plus what to do about a server you did not write. Estate · 9 min Building a Capability Registry → The inventory domain in full: what a registry records, how it stays current, and why generated beats maintained. Policy · 10 min Designing Approval Thresholds → Where to put the human, chosen by consequence rather than instinct — and how to keep approval from becoming rubber-stamping. Evidence · 11 min How to Audit AI Agent Actions → What one record per execution has to contain, why denials matter more than successes, and the questions the log must answer. Identity · 9 min AI Agent Governance vs Traditional IAM → Why role-based access control assumes a human decision-maker, and what breaks when the caller reasons about its own next step. Architecture · 12 min What Is an MCP Control Plane? → The five responsibilities of the deciding layer, the hard boundary it shares with servers, and four signals you have outgrown per-server configuration. Estate · 12 min MCP Server Inventory → Six discovery sources for finding every server you actually run, eleven fields per entry, and what to do with a server nobody owns. Governance · 12 min Enterprise MCP Registry → Four approval states, twelve fields, a five-step intake, and the one design decision that determines whether engineers use it or route around it. Policy · 13 min MCP Policy Engine → Ten inputs a decision needs, four outcomes richer than allow and deny, and a shadow mode so the first rule you ship is not also the first outage. Risk · 12 min MCP Tool Risk Scoring → Two questions, five classes, and why parameter values — not tool names — determine what a call actually risks. Scale · 12 min MCP Tool Discovery → Why a flat tool list stops scaling past roughly forty tools, and the four patterns that fix selection accuracy and token cost together. Operations · 13 min MCP Tool Routing → Six signals worth routing on, routebooks as reviewable artefacts, and the rule that a fallback may never widen authority. Instrumentation · 12 min MCP Observability → Twelve fields on every tool-call span, six metrics worth alerting on, and why refusal rate moves before error rate does. Buyer’s guide · 13 min Best MCP Gateways in 2026 → Four categories, twelve criteria, five disqualifiers and a five-day evaluation — including where our own product does not fit. Architecture · 17 min Enterprise MCP Gateway → Six components, three deployment models, failure behaviour per risk class, a rollout sequence and the total cost of ownership nobody quotes. Evidence · 17 min MCP Audit Logs → Sixteen fields per attempt, hash chaining, two-tier retention and the five questions an auditor actually asks. Approvals · 16 min Human Approval for MCP Workflows → Seven elements of an approval request, queue design, timeout semantics and the arithmetic of approver capacity. Orchestration · 16 min MCP Workflow Orchestration → Compensation instead of transactions, partial-failure design, and the eight fields that make a half-applied workflow diagnosable. Testing · 16 min MCP Workflow Simulation → Three fidelity levels, a fixture corpus built from real refusals, and a promotion criterion you can put in a pull request. Routing · 15 min What Is an MCP Routebook? → Eight fields per capability, eligibility separated from preference, and the no-eligible-route clause everybody omits. Change · 16 min MCP Change Impact Analysis → Six change types, five blast-radius queries, and why a reworded tool description passes every compatibility test. Capacity · 16 min MCP Capacity Planning → Why agent load is bursty rather than diurnal, a volume model you can build in a spreadsheet, and the four limits that bind before compute. Operations · 16 min MCP Estate Management → Six recurring routines, eight metrics on one page, and the five practices that break somewhere past thirty servers. Contracts · 16 min MCP Tool Contracts → Why two servers offering the same capability are rarely interchangeable, and the normalization layer that makes them so.

Sources and further reading

Primary specifications and standards this article relies on. The risk tiers, maturity stages and thirty-day sequence are our own model, drawn from building and operating five MCP servers — where a claim is judgement rather than something a standard states, the text says so.

  1. 01 · MCP project Model Context Protocol — specification ↗ The normative spec, including capability negotiation and the authorization sections governance has to build on.
  2. 02 · MCP project Model Context Protocol — official documentation ↗ Primary source for protocol structure, transports and tool definitions.
  3. 03 · NIST AI Risk Management Framework (AI RMF 1.0) ↗ The Govern, Map, Measure and Manage functions that the four domains here are an MCP-level expression of.
  4. 04 · ISO/IEC ISO/IEC 42001:2023 — AI management systems ↗ The certifiable management-system standard whose documentation requirements the evidence domain satisfies.
  5. 05 · OWASP GenAI Security Project OWASP GenAI LLM Top 10 (2026) ↗ Current consensus risk list; excessive agency is the entry this guide’s tiering exists to bound.
  6. 06 · NIST SP 800-207: Zero Trust Architecture ↗ Where the policy-enforcement-point and policy-decision-point separation used in stage 2 comes from.
  7. 07 · European Union Regulation (EU) 2024/1689 — Artificial Intelligence Act ↗ Consulted for its record-keeping and human-oversight obligations. This article is not legal advice.

Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.

Try the enumeration step against a real server

Week one asks you to enumerate rather than ask. To see what that produces — the full tool list a server advertises, unfiltered — point a client at a public one. Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run — no signup, no key, 54 tools. A tools/list call takes about a minute and shows you the shape an inventory row needs. Setup is in the reference.

Reference documentation

Want the specification rather than the argument?

Transport, auth, capability enumeration, policy evaluation and refusal semantics for every Barzel server.

Dev free · 10,000 calls/mo · paid plans from $199/mo on the MCPize marketplace

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.