5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

MCP Governance · Policy as code

MCP Policy as Code: Governing Agent Tool Access in Version Control

Policy that lives in a console dies with the person who configured it. This is how to express MCP tool governance as reviewable, testable, versioned code — and what breaks when you do it badly.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202618 min read5,270 words

The short answer

MCP policy as code means the rules deciding which agent may call which tool, with what arguments, under what conditions, live in a version-controlled repository as declarative files — reviewed in pull requests, tested against recorded traffic, promoted through environments, and enforced at a single decision point rather than configured by hand in a console. The point is not elegance. It is that a governance decision becomes an artefact with an author, a date, a reviewer and a diff — which is the only form an auditor, or a future colleague, can actually work with.

Summary for readers and answer engines

Reviewed 25 Aug 2026

  • ▸Policy as code means one repository holds every rule about which agent may call which MCP tool, and every change to those rules arrives as a reviewable diff.
  • ▸A policy decision can legitimately read ten inputs: caller identity, agent identity, tool identity, argument values, resource sensitivity, environment, time, recent behaviour, cumulative spend, and the approval state of the request.
  • ▸Six outcomes are enough for almost every enterprise: allow, deny, require approval, transform the arguments, route elsewhere, and log-only (shadow).
  • ▸Policy must be tested against recorded production traffic before it is enforced. Shadow mode is not optional — it is how you discover that a rule you thought was narrow denies eleven percent of legitimate calls.
  • ▸The common failure is not a missing feature. It is that nobody owns the repository, so policy accumulates and no one dares delete a rule.

Source: Mark Alex, Real Biz Digital — MCP Policy as Code: Governing Agent Tool Access in Version Control (https://realbizdigital.net/insights/mcp-policy-as-code/). Reproduce with attribution.

Key takeaways

  1. 01Write policy as data, not as code with side effects. A policy file that can make network calls is no longer testable, and no longer fast.
  2. 02Bind every rule to an owner and a review date in the file itself. Unowned rules are the reason policy repositories become archaeology.
  3. 03Never let the model supply an input the policy trusts. Tenant, actor and scope come from the token; only arguments come from the model.
  4. 04Ship policy through the same promotion pipeline as application code: local evaluation, shadow against recorded traffic, canary on one agent, then estate-wide.
  5. 05Measure the deny rate and the approval-queue depth. A policy set that denies nothing is decoration; one that denies constantly is a workflow bug wearing a security costume.
  6. 06Keep the decision fast and local. If policy evaluation needs a round trip to a third-party service, the failure mode of that service becomes the failure mode of your agents.
Part of the clusterMCP Governance →

Quick answers

One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.

What is MCP policy as code?
Expressing the governance rules for Model Context Protocol tool access as declarative, version-controlled files that a policy engine evaluates at call time, rather than as settings clicked into a console.
How is it different from an MCP policy engine?
The engine is the runtime that evaluates a decision. Policy as code is the discipline of authoring, reviewing, testing and promoting the rules that engine evaluates. You need both; only one of them is a file format.
What language should MCP policy be written in?
Prefer a declarative format with a schema — YAML or JSON with strict validation, or a purpose-built policy language such as Rego. Avoid general-purpose code: it invites side effects and defeats static review.
Can I test MCP policy before enforcing it?
Yes, and you must. Replay recorded tool calls through the new policy set in shadow mode, compare outcomes against the current set, and only promote when the diff is understood line by line.
Who should own the MCP policy repository?
One named platform or security engineering team owns the repository and the review gate; individual tool and data owners approve rules that touch their systems. Shared ownership with no single owner is how repositories rot.
Does policy as code slow agents down?
A local, in-process decision on a static rule set typically costs single-digit milliseconds. Latency problems come from remote lookups inside the decision path, not from the policy itself.
What happens when the policy engine is unavailable?
You decide in advance, per risk class, and write it down: fail closed for money and destructive actions, fail open with elevated logging for read-only ones. An undeclared answer becomes an outage or a breach, chosen at random.

Why console-configured MCP governance always decays

Every governance programme starts in a user interface, because a user interface is the fastest way to get one rule working. The problem arrives at rule forty.

A console is a superb place to make a first decision and a terrible place to keep a hundred. Nothing in it records why a rule exists. There is no diff when someone widens a scope at 6pm. There is no test that proves the change you are about to make will not deny the finance agent’s month-end run. And when the engineer who configured it leaves, the reasoning leaves with them — leaving a successor who is, quite rationally, too frightened to delete anything.

This is not an argument against consoles. It is an argument about where the source of truth lives. The console should render the policy; the repository should define it. When those two are inverted, three specific things fail, and they fail in the same order every time.

  • 01Review disappears. A policy change with real blast radius gets made by one person with no second pair of eyes, because the interface offers no natural place for a reviewer to stand.
  • 02History becomes unreconstructable. Six months later, an auditor asks who authorised the agent to issue refunds above £500. The console can show the current state. It cannot show the decision.
  • 03Rules accumulate and never leave. Deletion requires confidence that nothing depends on the rule. Without tests, nobody has that confidence, so the set only ever grows — and a large set nobody understands is functionally the same as no policy at all.

Policy as code fixes all three with boring, well-understood machinery: files, pull requests, tests, and a pipeline. There is nothing novel here. That is the appeal.

The ten inputs an MCP policy decision may legitimately use

A policy decision is a pure function of its inputs. Getting the input list right is most of the design work, because an input you cannot trust is worse than an input you do not have. The distinction that matters throughout: identity and context come from the transport and the token; only arguments come from the model.

Key facts

  • ▸Four of the ten inputs — recent behaviour, cumulative volume, approval state and resource sensitivity — require the decision point to hold state. A stateless policy engine cannot express rate ceilings or approvals, no matter how expressive its language.
  • ▸Argument values are the only input the model controls, and they are where every injection-driven abuse arrives. Validate them against the tool’s JSON Schema first, then against policy.
  • ▸If your policy reads a tenant identifier out of the tool arguments, you do not have tenant isolation. You have a suggestion.
InputSourceTrustworthy?What it lets you express
Caller identityValidated OAuth/OIDC access tokenYes — cryptographicallyAlice may refund; the contractor account may not
Agent identityClient credential or workload identityYesThe support agent may read tickets; the marketing agent may not
Tool identityServer registration plus tool name and versionYespayments.refund is governed; payments.list is not
Argument valuesThe model’s tool callNo — treat as hostile inputAmount > 500, path outside allowlist, URL not in the permitted host set
Resource sensitivityData classification on the target systemYes, if maintainedAny tool touching a system classified restricted requires approval
EnvironmentDeployment context of the calling agentYesProduction writes are governed differently from sandbox writes
Time and calendarServer clock plus business calendarYesNo ledger writes during close; no bulk sends outside working hours
Recent behaviourDecision history for this agentYes, if you keep itFourteen refunds in ten minutes is a pattern, not fourteen decisions
Cumulative spend or volumeCounter maintained by the decision pointYesDaily ceiling per agent, per tool class
Approval stateApproval workflow recordYesThis call carries a valid, unexpired, single-use approval token

Barzel Central Gateway evaluates ten inputs of exactly this shape and returns one of six outcomes. The list above is not a product description — it is what we found we needed after the first version had four inputs and could not express “this specific approval, once”.

Six outcomes, and why allow/deny is not enough

Binary policy engines produce binary politics. If the only answers are yes and no, every genuinely ambiguous case becomes an argument between a security team that must say no and a product team that needs a yes. Four additional outcomes dissolve most of those arguments.

Allow
The call proceeds unchanged. Recorded with the policy version that permitted it, so the decision is reconstructable later.
Deny
The call is refused before execution, with a machine-readable reason the agent can act on — not a bare 403 the model will simply retry. RFC 9457 problem details are a good shape for this.
Require approval
Execution suspends; a human with the authority to accept the consequence approves or rejects. The approval is bound to this exact call, single-use and time-limited.
Transform
The call proceeds with modified arguments: a redacted field, a clamped amount, a narrowed date range, an injected tenant filter. Transform is the outcome that turns a hard no into a safe yes.
Route
The call is redirected to a different server or tool that satisfies the same capability under acceptable conditions — the in-region instance, the read replica, the sandboxed equivalent.
Log only (shadow)
The call proceeds; the decision that would have been taken is recorded. This is how every new rule enters the estate.
Two outcomes
  • ›Every ambiguity escalates to a human argument
  • ›New rules go straight to enforcement, so nobody dares write them
  • ›Teams route around the gateway to get work done
  • ›Policy set stays small because each rule is expensive to add
Six outcomes
  • ›Ambiguity has a mechanism: approve, transform or route
  • ›New rules enter in shadow and prove themselves on real traffic
  • ›Routing around the gateway loses you capability, not just oversight
  • ›Policy set grows safely because each rule is cheap to test

Note that transform and route are the two outcomes teams most often skip, and the two that most reduce friction. A rule that clamps a refund to the approved maximum is strictly better than a rule that denies the refund and generates a support ticket.

Repository layout that survives contact with fifty rules

The structure below is deliberately dull. It optimises for one thing: when something is denied in production, an engineer who has never seen the repository should find the responsible rule in under two minutes.

One rule, one file
A file per rule, named for what it does. Bundling twelve rules into payments.yaml guarantees that a change to one requires reviewing all twelve.
Owner and review date in the file
Every rule carries owner, created, review_by and rationale. The rationale field is the single highest-value line in the repository and the one teams always omit.
Classes, not per-tool copies
Express “destructive infrastructure actions require two approvers” once, against a class, then assign tools to the class. Otherwise every new server means fifteen new rules.
Fixtures alongside rules
Recorded traffic lives in the repository, scrubbed. Without fixtures, tests are hypothetical and shadow mode is your only test environment.
Environments as overlays
The same rules, with different enforcement levels per environment. Never fork the rule set per environment — the forks diverge within a month.

policy/ — suggested layout

policy/
  meta.yaml                  # schema version, owners, review cadence
  classes/
    risk-tiers.yaml          # tool risk classes and what each implies
    data-sensitivity.yaml    # classification -> required controls
  rules/
    payments/
      refund-ceiling.yaml
      supplier-payment-approval.yaml
    data/
      customer-export.yaml
      pii-redaction.yaml
    infra/
      destructive-actions.yaml
  fixtures/
    recorded/2026-08/*.jsonl # replayed tool calls, argument values scrubbed
    expected/*.yaml          # asserted outcome per fixture
  environments/
    shadow.yaml  canary.yaml  production.yaml

A useful discipline: make rationale a required field that the schema validator enforces. It costs an author one sentence and saves a successor an afternoon of guessing.

The four-stage promotion pipeline

Policy changes are production changes with an unusual property: a bad one fails closed, which means the outage looks like the security control working correctly. That is exactly why the pipeline matters more here than in ordinary application code.

Stage 01

Local evaluation and schema validation

Every rule validates against the policy schema, and the rule set evaluates against unit fixtures on the author’s machine and in CI. Static checks catch the majority of real mistakes: a mistyped tool name that silently matches nothing, a scope that is wider than intended, a missing owner.

Fail the build on an unmatched tool reference. A rule that matches no tool is not conservative — it is a rule you believe is protecting you and is not.

Stage 02

Shadow against recorded traffic

Replay a month of recorded tool calls through the new set and diff outcomes against the current set. Report every changed decision, grouped by rule. This is where you discover that your “narrow” export rule would have blocked the nightly reconciliation job.

Set a promotion criterion in advance, in writing: no unexplained outcome changes, and every intended change accounted for by a named rule.

Stage 03

Canary on one agent

Enforce for a single agent, or a single team, for a fixed window — a week is usually enough to cross one business cycle. Watch deny rate, approval-queue depth and time-to-approval, not just error counts.

The metric that predicts trouble is not denials. It is approvals that expire unactioned, because that means you have moved work onto humans who did not agree to receive it.

Stage 04

Estate-wide, with a documented rollback

Promote, keeping the previous version deployable in one step. Record the policy set version on every decision so that any later question about a call resolves to a specific commit.

Rollback plans that require a human to reconstruct the previous state are not rollback plans.

Three of these four stages need recorded traffic. If you take one operational decision from this article, start recording tool calls and their arguments — scrubbed appropriately — before you write your second policy rule.

A worked example: the refund ceiling that was almost an outage

This is the shape of the problem in practice, with the numbers changed. A support agent handles refunds. The rule looks obviously correct. Shadow mode disagrees.

Worked example · Refund ceiling, from draft to production
Draft rulepayments.refund: deny if amount > 500
IntentSmall refunds automatic; larger ones need a human
Shadow result11.4% of calls denied — expected roughly 2%

The gap was currency. A minority of refunds arrived in cents from one upstream integration and in pounds from another, and the tool schema said only number. Every cents-denominated refund above £5 tripped a rule intended to catch refunds above £500.

Second draftRequire currency and minor_units explicitly; normalise before comparison; deny on absent currency
Shadow result2.1% denied, all genuinely above the ceiling
RefinementReplace deny with require approval for 500–5,000, deny above 5,000
CanaryOne team, one week: 31 approvals, median time-to-approval 7 minutes, zero expired

Production for six weeks now. The interesting artefact is not the rule — it is the commit message explaining the currency ambiguity, which stopped the same class of mistake being made twice on a different tool.

Cost of finding this in productionOne denied refund for a customer already waiting, times roughly 400 a month

The lesson generalises past currency. Any policy that compares a number supplied by a model against a threshold must first pin down the unit. Tool contracts that leave units implicit make policy unenforceable, which is why schema normalisation and policy authoring are the same project.

Six ways policy-as-code programmes fail

Mistake

Policy that makes network calls

A rule that fetches a user’s group membership from an HTTP endpoint at decision time is untestable, non-deterministic and puts a third party in the latency path of every tool call.

Instead: Push the data to the decision point out-of-band and evaluate against a local snapshot. Accept staleness measured in minutes; declare it in the rule.

Mistake

No owner on the repository

Everyone can raise a pull request; nobody is accountable for the set as a whole. The result is a rule set that only grows and is never pruned.

Instead: One named team owns the repository and the review gate. Tool owners approve rules touching their systems. Write both into the meta file.

Mistake

Enforcing on day one

A rule goes straight to production because it is “obviously correct”. See the refund example: obviously correct rules deny 11% of legitimate traffic surprisingly often.

Instead: Shadow mode is mandatory for every new rule, without exception, including the trivial ones. The trivial ones are the ones nobody tests.

Mistake

Modelling per tool instead of per class

Fifteen new rules every time a server is onboarded. Within two quarters nobody can answer “what governs destructive actions?” without reading the whole repository.

Instead: Define risk and sensitivity classes; assign tools to classes; write rules against classes. Onboarding becomes one classification decision.

Mistake

Trusting model-supplied context

The policy reads tenant_id or user_id from the tool arguments, so a persuaded model can select its own authorisation context.

Instead: Derive identity and tenancy from the validated token, inject them server-side, and treat any matching argument as adversarial input to be overwritten.

Mistake

No declared behaviour when the engine is down

Nobody decided, so the answer is whatever the code happened to do — usually fail open for everything, discovered during an incident.

Instead: Declare fail-closed and fail-open sets by risk class, in the repository, and test both paths in CI.

What to measure once policy is running

Policy health is measurable, and the useful metrics are not the ones dashboards default to. Volume of decisions tells you nothing. These eight tell you whether the programme is working.

Policy health metrics and their failure signal
MetricHealthy range (our estates)What a bad value means
Deny rate, all tools0.5–3%Near zero: policy is decoration. Above 8%: rules are fighting real workflows
Approval requests per approver per dayunder 12Above 20 and approvals become rubber stamps — the control is now theatre
Median time-to-approvalunder 15 minHours mean agents stall and teams start routing around the gateway
Expired approvalsunder 2%High expiry means you moved work to humans who never agreed to receive it
Transform rate2–10%Zero means you are denying things you could have made safe
Rules with no matching traffic in 90 daysunder 10%Higher means the repository is accumulating fiction
Rules past review_by date0Anything above zero is an audit finding waiting to be found
p99 decision latencyunder 25 msAbove that, something remote is in the decision path

These are our operating ranges across the estates we have instrumented, not an industry benchmark, and they are stated so you can argue with them rather than inherit them silently. Treat them as starting thresholds and recalibrate against your own first quarter of data.

Next step

Write your first rule against recorded traffic, not against a guess

The gateway records ten inputs per decision and can evaluate a candidate rule set in shadow mode against what already happened. That is the cheapest way to find out whether the rule you are about to write means what you think it means.

What policy as code does not solve

Policy as code is a discipline for managing rules. It does not tell you which rules are correct for your business, and it will faithfully version a bad rule set forever.

Three limits are worth stating plainly, because vendors rarely do.

  • 01It cannot govern a path that does not pass through the decision point. An agent holding a direct credential to an upstream system bypasses every rule you wrote. Policy as code is downstream of credential hygiene.
  • 02It cannot judge intent. A refund of £499 issued four hundred times is within every per-call rule and is still fraud. Cumulative and behavioural rules help; they do not close the gap. Detection and reconciliation remain necessary.
  • 03It does not make a policy set correct, only reviewable. Reviewability is worth a great deal — and it is a different claim from safety.

Frequently asked questions

What is MCP policy as code?

MCP policy as code means the rules governing which agent may call which Model Context Protocol tool, with which arguments, under which conditions, live in a version-controlled repository as declarative files. They are reviewed in pull requests, tested against recorded traffic, promoted through environments, and enforced at a single decision point rather than configured by hand in a console.

How is policy as code different from an MCP policy engine?

The policy engine is the runtime component that evaluates a decision when a tool call arrives. Policy as code is the practice of authoring, reviewing, testing and promoting the rules that engine evaluates. A policy engine configured entirely through a web console is still a policy engine; it just has no reviewable history.

What language should MCP policy be written in?

Prefer a declarative format validated against a schema: YAML or JSON with strict validation, or a purpose-built policy language such as Rego. Avoid general-purpose programming languages for rule bodies, because they permit side effects, make static review unreliable, and put arbitrary execution in the latency path of every tool call.

Which inputs can an MCP policy decision safely use?

Ten: caller identity, agent identity, tool identity, argument values, resource sensitivity, environment, time and calendar, recent behaviour, cumulative spend or volume, and approval state. All except argument values can be made trustworthy. Argument values come from the model and must be treated as hostile input validated against the tool schema first.

How do I test MCP policy before enforcing it?

Replay recorded production tool calls through the candidate policy set in shadow mode, then diff the outcomes against the current set and account for every change by naming the rule responsible. Promote only when the diff is fully explained. Set the promotion criterion in writing before you look at the results.

What outcomes should an MCP policy engine support?

Six are sufficient for almost every enterprise: allow, deny, require human approval, transform the arguments, route to an alternative server or tool, and log only. Transform and route are the two most often omitted and the two that most reduce friction, because they convert an unsafe request into a safe one instead of refusing it.

Who should own the MCP policy repository?

One named platform or security engineering team owns the repository and the review gate. Individual tool and data owners approve rules that touch their systems. Diffuse ownership produces a repository that only grows, because deletion requires a confidence nobody has authority to assert.

Does policy evaluation add latency to agent tool calls?

A local, in-process evaluation against a static rule set typically costs single-digit milliseconds and is not perceptible next to model inference time. Latency problems come from remote lookups inside the decision path: group membership queries, external risk scores, third-party APIs. Push that data to the decision point out-of-band.

What should happen if the policy engine is unavailable?

Decide in advance per risk class and record it in the repository. Fail closed for money movement, destructive actions and anything touching restricted data; fail open with elevated logging and a hard time limit for read-only calls. Both paths should be exercised in CI, because an undeclared answer becomes either an outage or a breach chosen at random.

How many policy rules should an estate have?

Fewer than teams expect. Rules written against risk and sensitivity classes cover most of an estate in twenty to forty files, whereas per-tool rules reach several hundred and become unmaintainable. If your rule count grows linearly with server count, you are modelling at the wrong level.

Can policy as code satisfy an auditor?

It provides much of what an auditor asks for: who authorised a control, when, with what review, and what changed since. It does not by itself demonstrate that the controls are sufficient, and it cannot evidence enforcement without decision logs. Pair the repository with tamper-evident decision records.

How does this relate to MCP tool contracts and schema normalisation?

Directly, and it is the dependency teams discover late. Policy compares values, and values without declared units or stable semantics cannot be compared safely. Normalising tool contracts across servers is a prerequisite for policy that means what it says, as the refund-ceiling example on this page shows.

Glossary

MCP policy as code
The practice of expressing Model Context Protocol governance rules as declarative, version-controlled files that are reviewed, tested and promoted like application code.
Policy decision point (PDP)
The component that evaluates a request against policy and returns an outcome, without itself performing the action.
Policy enforcement point (PEP)
The component that receives the decision and applies it — permitting, refusing or altering the call.
Shadow mode
Evaluating a policy set against live or recorded traffic and recording the outcome that would have been applied, without applying it.
Transform outcome
A decision that permits a call after modifying its arguments, such as redacting a field, clamping an amount or injecting a tenant filter.
Risk class
A named grouping of tools sharing consequence characteristics, so rules can be written once against the class rather than per tool.
Promotion pipeline
The staged path a policy change travels: local validation, shadow, canary, estate-wide, each with a defined promotion criterion.
Fail-closed set
The explicitly enumerated tools or classes whose calls must be refused when the policy engine cannot render a decision.
Rationale field
A required per-rule note recording why the rule exists, which is what makes later deletion possible.
Decision record
The persisted evidence of one policy evaluation: inputs, outcome, rule matched and policy set version.

Standards and entities referenced

Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.

  1. 01 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  2. 02 · CNCFOpen Policy Agent — documentation ↗Reference implementation of decoupled policy decisions and policy as code.
  3. 03 · NISTNIST SP 800-207 — Zero Trust Architecture ↗The policy decision point / policy enforcement point split this architecture borrows directly.
  4. 04 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
  5. 05 · IETFRFC 9457 — Problem Details for HTTP APIs ↗A machine-readable shape for the denial reasons a policy layer returns.
  6. 06 · JSON SchemaJSON Schema Specification ↗How tool parameter contracts are expressed, and what a validator can enforce.
  7. 07 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
  8. 08 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  9. 09 · ISOISO/IEC 42001 — AI management systems ↗The management-system standard auditors increasingly map AI governance evidence against.
  10. 10 · SemVerSemantic Versioning 2.0.0 ↗The versioning contract tool schemas should honour but frequently do not.

Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). MCP Policy as Code: Governing Agent Tool Access in Version Control. Real Biz Digital. https://realbizdigital.net/insights/mcp-policy-as-code/

Try the mechanics on a live server

To watch a real tools/list response, and see how much surface one server exposes, before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

Barzel Central Gateway is this layer, sold as a running product

Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so an evaluation costs an afternoon rather than a purchase order.

PlanPriceIncludedRight for
CommunityFree1,000 tool calls/mo · full policy engine, registry, routing, auditEvaluating the estate, or one team proving the path works
Starter$10/mo10,000 calls/mo · everything in CommunityOne or two production agents against a handful of servers
Team$79/mo100,000 calls/mo · routebooks, simulation, change impactA platform team governing an estate of 5–20 servers
Business$149/mo250,000 calls/mo · estate-wide evidence exportMulti-team governance with SIEM obligations

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

The five Barzel servers, and which problem each one is sold for

One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.

ServerSold forEntry priceWhere it sits
Barzel Central GatewayKnowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidenceFree, then $10–$149/moControl plane — decides what may be reached, and by whom
BarzelVaultStopping a specific dangerous action before it executes, with proof afterwards$199–$3,999/moDecision point — evaluates the individual call before execution
BarzelOpsRunning real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approvalFree, then $19–$199/moExecution layer — does the work the policy allowed
Barzel FinOps AtlasAttributing AI spend to agents, tools and outcomes, then forecasting and capping itFree, then $29–$799/moEconomics layer — what the estate costs per outcome
Barzel Scripture IntelligenceA free, credential-free public MCP server to test clients and inspect real protocol trafficFree, unmetered, no signupReference implementation — safe place to learn the protocol

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.