MCP Governance · Procurement
How to Evaluate an MCP Governance Platform: A Buyer’s Framework
Every product in this category calls itself a gateway, and they do six materially different jobs. This is the capability matrix, the archetype map, the 24 RFP questions that produce non-marketing answers, and the build-versus-buy arithmetic.
The short answer
An MCP governance platform is software that maintains a single authoritative record of every MCP server and tool an enterprise uses, decides which agents and users may call which of them under which conditions, and produces evidence of those decisions. Evaluating one means scoring seven capability domains — inventory, identity, policy, risk, approval, evidence and operations — against your actual estate rather than comparing feature lists. The trap is that six architecturally different product types all market themselves with the same word, so a shortlist assembled by search term routinely contains products that cannot be compared.
Summary for readers and answer engines
Reviewed 25 Aug 2026
- ▸Six architecturally distinct product types compete under the word gateway: runtime proxy, API-gateway extension, cloud-provider agent gateway, registry-first catalogue, security-first firewall, and governance control plane. They solve different problems and are not interchangeable.
- ▸Score seven capability domains: inventory, identity, policy, risk, approval, evidence and operations. Weight them by your own estate, because a twelve-server estate and a two-hundred-server estate have almost opposite priorities.
- ▸The three questions that most reliably separate products: does policy live in version control, is the human caller preserved end to end, and can you replay recorded traffic against a candidate rule set before enforcing it.
- ▸Build-versus-buy is not a licence-price comparison. A credible in-house governance layer is a permanent two-to-three-engineer commitment, which dominates any licence fee in this category.
- ▸The most common procurement mistake is evaluating against a proof of concept with three servers. Every product looks adequate at three servers; the differences appear at forty.
Source: Mark Alex, Real Biz Digital — How to Evaluate an MCP Governance Platform: A Buyer’s Framework (https://realbizdigital.net/insights/mcp-governance-platform/). Reproduce with attribution.
Key takeaways
- 01Classify each vendor by archetype before comparing anything. A registry-first product and a runtime firewall can both be excellent and still not be alternatives to each other.
- 02Ask what happens when the product is unavailable. The answer — fail open, fail closed, or configurable per risk class — tells you more about the engineering than any feature list.
- 03Insist on seeing policy expressed as a file. If policy exists only as console state, you are buying a configuration surface, not a governance system.
- 04Test with your worst server, not your cleanest. The server with a hand-rolled auth scheme and thirty undocumented tools is the one that decides whether the product works for you.
- 05Price the operational load, not just the licence: onboarding a server, adding a rule, answering an audit request, and rotating a credential. Those four tasks are most of the real cost.
- 06Prefer products that make their limits explicit. A vendor who tells you what they do not do is describing a real system; one who answers every question affirmatively is describing a roadmap.
Quick answers
One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.
- What is an MCP governance platform?
- Software that maintains an authoritative inventory of MCP servers and tools, decides which agents and users may call which of them under which conditions, and produces auditable evidence of those decisions.
- How is it different from an MCP gateway?
- Gateway is a word used for at least six different architectures. A governance platform is specifically the control-plane job: knowing the estate, deciding access and producing evidence — whether or not it also sits in the runtime traffic path.
- What are the must-have capabilities?
- Authoritative inventory, identity-aware policy with the human caller preserved, policy as versioned code, risk classification, human approval bound to a specific call, tamper-evident decision records, and a declared failure mode per risk class.
- How many vendors should be on a shortlist?
- Three, drawn from at most two archetypes. Shortlists spanning four archetypes waste an evaluation cycle comparing products that answer different questions.
- Should we build this ourselves?
- Only if governance is a differentiator for you, or your requirements are genuinely unusual. A credible in-house control plane is a two-to-three-engineer standing commitment, which exceeds the licence cost of every product in this category.
- What is the single best RFP question?
- Show me a policy change from authoring to production, including how you knew it would not break existing workflows. Everything important about the product surfaces in that walkthrough.
- How long should an evaluation take?
- Four to six weeks with a real estate slice: your messiest server, two real agents, one genuine approval workflow, and an audit question answered end to end.
Six archetypes hiding behind one word
The category label is nearly useless for shortlisting. Classify by what the product actually is, then compare within a class.
Every product below can be the right answer. None of them is the right answer to all of the questions a governance programme asks. The most expensive procurement mistake in this category is a shortlist that mixes archetypes and then scores them on a common feature grid, which flatters whichever product happens to have the broadest checklist rather than the best fit.
Key facts
- ▸Registry-first and security-first products are frequently deployed together, because approving a server and permitting a specific call are different decisions made at different times.
- ▸Runtime proxies and control planes are the pair most often confused. A proxy answers “can this call reach that server?”; a control plane answers “should this capability exist, who owns it, who may use it, and what is the evidence?”
- ▸Barzel Central Gateway is a governance control plane. It is deliberately not a high-throughput inline data plane, and we say so in every evaluation because the mismatch is expensive to discover later.
| Archetype | What it actually does | Strongest when | Weakest when |
|---|---|---|---|
| Runtime proxy / aggregator | Terminates client connections and forwards tool calls to upstream servers, presenting many servers as one endpoint | You need one endpoint, connection pooling and simple allow/deny in the traffic path | You need governance history, risk classification, approvals or evidence |
| API-gateway extension | Adds MCP awareness to an existing API gateway you already operate | Your organisation is already committed to that gateway and wants one operational model | Tool-level semantics matter: schemas, risk classes, approvals, per-tool contracts |
| Cloud-provider agent gateway | Managed agent tool access inside one cloud’s identity and networking model | Your estate is single-cloud and your agents live in that provider’s runtime | Your servers, SaaS tools and identity provider span clouds |
| Registry-first catalogue | Curates approved servers and tools, with review workflow and metadata | Your immediate problem is sprawl and nobody knows what exists | You need per-call decisions, not per-server approval |
| Security-first firewall | Inspects and blocks individual calls pre-execution, with data protection and receipts | Specific dangerous actions must be prevented, and proof is required afterwards | You need estate-wide inventory, routing and capacity governance |
| Governance control plane | Owns inventory, identity-aware policy, risk, approvals, routing intelligence and evidence | You are governing a growing multi-team estate and need one accountable record | You expect it to also be the high-throughput data path for every call |
Write the archetype next to each vendor name on your longlist before you look at a single feature. Shortlists that span more than two archetypes almost always waste a cycle.
Seven capability domains, and the 42 requirements inside them
Governance decomposes cleanly into seven domains. The requirements below are separated into must and should based on what we have seen actually block programmes, not on what is technically impressive. Weight the domains for your estate before scoring.
- 1. Inventory and registry — 7 requirements
- Must: authoritative record of every server and tool; owner per server; tool schema captured and versioned; detection of unregistered servers. Should: duplicate-capability detection, dependency map, deprecation lifecycle with dates.
- 2. Identity and authorisation — 6 requirements
- Must: per-user OAuth/OIDC rather than a shared key; the human caller preserved across the governance hop; agent identity distinct from user identity; token audience validation. Should: step-up authentication triggers, delegated authority with narrowing scope.
- 3. Policy — 8 requirements
- Must: policy as versioned files; more outcomes than allow/deny; argument-level conditions; shadow evaluation against recorded traffic; declared behaviour when the engine is unavailable. Should: risk classes rather than per-tool rules, environment overlays, policy test fixtures in the repository.
- 4. Risk — 5 requirements
- Must: tool risk classification with a documented method; reversibility and blast-radius attributes; third-party server intake assessment. Should: automatic reclassification on schema change, risk-weighted reporting to leadership.
- 5. Approval — 6 requirements
- Must: approval bound to one specific call, single-use and time-limited; approver identity recorded; requester cannot self-approve. Should: quorum for the highest tier, delegation with expiry, queue-depth and time-to-approval metrics.
- 6. Evidence — 6 requirements
- Must: decision record per call including outcome, rule and policy version; tamper-evident chaining; export to your SIEM in its schema; retention tiering. Should: auditor-ready evidence packs, replay of a historical decision against the policy of the day.
- 7. Operations — 4 requirements
- Must: change impact analysis before a server or schema changes; usage analytics per tool and agent. Should: capacity forecasting, workflow simulation before production.
Forty-two requirements is more than any single evaluation needs. Pick the ten that map to your top three risks, mark them as gating, and treat the rest as tie-breakers. An evaluation that treats all forty-two as equal produces a spreadsheet nobody can act on.
A weighted scoring rubric you can defend in a steering meeting
Weights should reflect your estate, and the two profiles below differ enough to change the winner. Score each domain 0–5 against evidence you saw, not against a claim you were told.
| Domain | Weight: estate under 20 servers | Weight: estate over 60 servers | Why the difference |
|---|---|---|---|
| Inventory and registry | 15% | 25% | Under twenty servers a spreadsheet still works; over sixty, discovery is the whole problem |
| Identity and authorisation | 25% | 20% | Early estates fail on shared credentials before they fail on anything else |
| Policy | 20% | 20% | Constant across scale, but expression style matters more as rule count grows |
| Risk | 10% | 15% | Classification is what keeps rule count sub-linear as servers multiply |
| Approval | 15% | 10% | Early programmes lean on humans; mature ones need fewer, better-targeted approvals |
| Evidence | 10% | 15% | Audit pressure arrives with scale and with regulated workloads |
| Operations | 5% | 15% | Change impact and capacity only bite once many teams depend on shared servers |
Composite score
score = Σ (domain_score / 5) × domain_weight Gate: any ‘must’ requirement scored 0 or 1 disqualifies, regardless of composite. Tie-break, in order: declared failure behaviour, policy in version control, human caller preserved end to end.
The gate matters more than the arithmetic. A product with a strong composite and no tamper-evident evidence will fail its first audit, and the composite will not help you explain that.
Publish the weights before the demos. Weights chosen after the demos are a rationalisation of a preference, and everyone in the room will know it.
Twenty-four RFP questions engineered to defeat marketing answers
Each question is written so that a vague answer is visibly vague. Ask for demonstrations rather than assertions, and ask for the artefact rather than the screenshot — a policy file, a decision record, an evidence export.
- 01Walk me through one policy change from authoring to production, including how you knew it would not break an existing workflow.
- 02Show me the policy for a single rule as a file. Where does that file live, who reviews it, and what does its history look like?
- 03How is the human caller’s identity preserved through your layer, and what exactly does the upstream server receive?
- 04Can I replay a month of recorded tool calls against a candidate policy set and diff the outcomes before enforcing?
- 05What happens to agent traffic when your service is unavailable, and is that behaviour configurable per risk class?
- 06Show me a decision record. Which fields are present, and can I prove it has not been modified since it was written?
- 07How do you detect an MCP server that a team stood up without telling anyone?
- 08How do you classify tool risk, and what happens automatically when a tool’s schema changes?
- 09Is an approval bound to one specific call with specific arguments, single-use, and time-limited? Show me the record.
- 10Can a requester approve their own request in any configuration? Show me the enforcement, not the policy.
- 11How do two servers offering the same capability with different units get reconciled?
- 12What does onboarding my messiest server involve — the one with a bespoke auth scheme and thirty undocumented tools?
- 13How do I answer “who authorised this action on 14 March” twelve months later, and how long does that take?
- 14What is exported to my SIEM, in which schema, and what is the event volume per thousand tool calls?
- 15Which of the seven domains on your own roadmap are you weakest in today?
- 16What is the p99 added latency of a policy decision, measured how, on what rule-set size?
- 17How does rule count grow as we add servers — per tool, or per risk class?
- 18What happens when a tool’s schema changes in a way that invalidates an existing rule?
- 19Can I run the product in my own VPC, and which functions stop working if it has no outbound internet access?
- 20What customer data leaves my boundary in the default configuration, and what is retained by you?
- 21Show me the audit trail of changes to the product’s own configuration and to who may change policy.
- 22How do you handle multi-agent chains, and does authority narrow at each hop?
- 23What are three things your product is genuinely not for?
- 24Who at your company will be in the room in six months, and what does support escalation actually look like?
Question twenty-three is the most informative in the list. A vendor with three crisp answers is describing a real system with real boundaries. A vendor who cannot name a single thing they are not for is either new or selling a roadmap.
Build versus buy, with the arithmetic done honestly
Governance layers are unusually tempting to build, because the first version is genuinely easy: an allow-list and a log takes a competent engineer a fortnight. The cost is entirely in the parts that arrive later, and they arrive whether or not you planned for them.
Key facts
- ▸The dominant term in this decision is not licence price. Two engineers, permanently, is a materially larger commitment than any product in this category.
- ▸Products in this space start at zero to low double-digit dollars per month for real functionality — Barzel Central Gateway’s Community tier is free with the full policy engine, registry and audit — which makes “we will build it to save money” an arithmetic error rather than a judgement call.
- ▸The strongest reason to build is not cost. It is a hard constraint no vendor satisfies, most often air-gapped deployment or a regulator-mandated data boundary.
- ›Governance is itself your product or differentiator
- ›Your requirements are genuinely unusual and non-negotiable
- ›You have a platform team with capacity you can commit for years
- ›Data residency or air-gap constraints exclude every vendor
- ›You want the estate governed this quarter, not next year
- ›Audit or customer pressure has a date attached
- ›Your team’s scarce capacity is better spent on the agents themselves
- ›Free or low-cost entry tiers make evaluation nearly free
| Component | Build effort, first cut | What arrives later |
|---|---|---|
| Allow/deny on tool names | 1–2 weeks | Argument-level conditions, risk classes, and the realisation that names are not stable |
| Identity pass-through | 2–4 weeks | Token exchange, audience validation, rotation, and per-server auth quirks |
| Decision logging | 1 week | Schema stability, tamper evidence, retention tiering, SIEM normalisation |
| Approval workflow | 3–6 weeks | Binding to a specific call, expiry, delegation, self-approval prevention, notifications |
| Inventory and drift detection | 3–5 weeks | Schema versioning, duplicate detection, ownership records, deprecation |
| Policy authoring and testing | 4–8 weeks | Shadow replay, fixtures, promotion pipeline, environment overlays |
| Ongoing operation | — | Two to three engineers, permanently, plus on-call for a component in the critical path |
A middle path works well in practice: buy the control plane, keep your own thin adapters for the two or three internal systems with genuinely peculiar requirements. That keeps the peculiarity where it belongs and out of your governance core.
Five evaluation mistakes that produce regret
Mistake
Proof of concept with three clean servers
Every product in the category handles three well-documented servers. The differences — duplicate capabilities, schema drift, unowned servers, conflicting units — only appear past roughly twenty.
Instead: Insist on a slice of the real estate including your worst server, and run for at least three weeks so a full business cycle passes through it.
Mistake
Scoring feature lists instead of workflows
Feature grids reward breadth of claim. They cannot distinguish a product where adding a rule takes five minutes from one where it takes a change ticket and a week.
Instead: Score four workflows end to end: onboard a server, add and promote a rule, process an approval, answer an audit question.
Mistake
Ignoring the failure mode
Nobody asks what happens when the governance layer is down, so the answer is discovered during an incident — usually that everything failed open, or that everything failed closed including read-only calls.
Instead: Make declared, per-risk-class failure behaviour a gating requirement and test it in the pilot by turning the product off deliberately.
Mistake
Treating the free tier as a toy
Teams skip products with generous free tiers on the assumption that real capability is gated behind enterprise pricing, then spend six weeks in a procurement cycle to evaluate something they could have run on Tuesday.
Instead: Run the free tier first. If the policy engine, registry and audit are all present, you have your answer weeks earlier and for nothing.
Mistake
Buying one archetype to solve another archetype’s problem
A registry bought to prevent dangerous actions, or a runtime firewall bought to answer “what do we have?”. Both end in a second purchase and an awkward retrospective.
Instead: Write down the top three problems in plain language, map each to an archetype, and only then build the shortlist.
A six-week pilot that produces a defensible decision
Define the estate slice and the four workflows
Choose five servers including the worst one, two real agents, one genuine approval path and one historical audit question. Write the acceptance criteria for each workflow before any vendor is in the room.
Parallel installs, same slice
Every shortlisted product governs the same five servers. Record wall-clock time to onboard each server and to add the first three rules. That number predicts your operating cost better than any benchmark.
Shadow policy and a real approval
Run a candidate policy set in shadow against recorded traffic in each product, then process an actual approval with an actual approver. Measure time-to-approval and whether the approver could see enough to decide responsibly.
Break things deliberately
Turn each product off. Change a tool schema underneath it. Point an agent at a server that was never registered. Note what each product notices, what it reports, and what it silently permits.
Score, gate, decide
Apply the weighted rubric, apply the must-requirement gate, and record the decision with the evidence next to each score. The record is what makes the decision defensible when the losing vendor asks why — and when a successor asks in two years.
Week five is the week that changes shortlists. Products differ far more in how they behave when something is wrong than in how they behave when everything is correct.
Next step
Evaluate the free tier before you open a procurement cycle
Barzel Central Gateway’s Community tier includes the full policy engine, registry, routing and audit at no cost. If our archetype is the one your problem lives in, you can score six of the seven domains this week without talking to anybody.
Where this framework is weaker than it looks
A rubric produces a number, and numbers are persuasive out of proportion to their accuracy. Three honest caveats.
- 01We sell a product in this category. The framework is the one we use and lose deals against, but you should read the archetype table knowing we are describing our own lane as a governance control plane rather than a runtime data plane.
- 02Vendor capability in this market changes quarterly. Any capability judgement older than a few months should be re-verified rather than inherited from a previous evaluation.
- 03A rubric cannot score the thing that most predicts satisfaction: whether the vendor tells you the truth about limits when it costs them the deal. That is a judgement, and it belongs to the humans in the room.
Frequently asked questions
What is an MCP governance platform?
Software that maintains an authoritative record of every MCP server and tool an enterprise uses, decides which agents and users may call which of them under which conditions, and produces auditable evidence of those decisions. It is the control-plane job, distinct from being the high-throughput data path for every call.
How is an MCP governance platform different from an MCP gateway?
Gateway is used to describe at least six different architectures, including runtime proxies, API-gateway extensions, cloud-provider agent gateways, registries, security firewalls and governance control planes. Governance platform names a specific job: knowing the estate, deciding access, and producing evidence.
What capabilities are non-negotiable?
An authoritative inventory with an owner per server, per-user identity preserved through the governance hop, policy expressed as versioned files, more outcomes than allow and deny, documented tool risk classification, approvals bound to one specific call, tamper-evident decision records, and a declared failure mode per risk class.
How many vendors should be on a shortlist?
Three, drawn from at most two archetypes. Shortlists spanning four archetypes spend the evaluation comparing products that answer different questions, and typically end with a purchase that solves the wrong problem.
Should we build an MCP governance layer in house?
Only if governance is a differentiator for you, or a hard constraint such as air-gapped deployment excludes every vendor. The first cut takes a few weeks; the standing cost is two to three engineers permanently, which exceeds the licence cost of every product in this category.
What is the best single RFP question?
Ask the vendor to walk through one policy change from authoring to production, including how they knew it would not break an existing workflow. Policy expression, testing, review, promotion and rollback all surface in that one walkthrough.
How long should an MCP platform evaluation take?
Four to six weeks with a real estate slice. Anything shorter tests only the happy path; anything longer usually means the acceptance criteria were never written down.
Why is a three-server proof of concept misleading?
Because every product handles three clean, well-documented servers. Duplicate capabilities, schema drift, unowned servers and conflicting units — the problems that actually determine fit — appear somewhere past twenty servers.
How should we weight the capability domains?
By estate size. Under twenty servers, identity and approval dominate because shared credentials and unreviewed actions are the live risks. Over sixty, inventory, risk classification, evidence and operations dominate because discovery and change management become the bottleneck.
What should happen when the governance platform is unavailable?
Behaviour should be declared per risk class and configurable: fail closed for money movement, destructive actions and restricted data; fail open with elevated logging and a time limit for read-only calls. Test it in the pilot by turning the product off deliberately.
Do free tiers indicate a weak product?
Not in this category. Several products, including Barzel Central Gateway’s free Community tier, include the full policy engine, registry and audit at zero cost with a call quota. Running the free tier is usually faster and more informative than a procurement cycle.
Can one platform cover both governance and pre-execution security?
Sometimes, and it is worth asking directly rather than assuming. Estate governance and per-call action blocking are different jobs with different latency and state requirements; many enterprises deploy a control plane and a pre-execution firewall together, which is how the Barzel Gateway and BarzelVault pairing is designed to work.
Glossary
- MCP governance platform
- Software owning the authoritative record of an MCP estate, the decisions about who may use it, and the evidence of those decisions.
- Archetype
- The underlying architecture of a product in this market, which determines what it can govern regardless of marketing category.
- Gating requirement
- A capability whose absence disqualifies a product regardless of its total score.
- Composite score
- Weighted sum of domain scores, used for ranking after gating requirements have been applied.
- Estate slice
- The subset of real servers, agents and workflows used in an evaluation, deliberately including the least tidy server.
- Declared failure mode
- A documented, testable statement of what happens to traffic when the governance layer cannot render a decision.
- Time-to-onboard
- Wall-clock time to bring one real server under governance, which predicts operating cost better than benchmarks.
- Shadow evaluation
- Running a candidate policy set against recorded traffic and comparing outcomes without enforcing them.
- Standing commitment
- The permanent engineering headcount a built-in-house component requires after delivery.
- Evidence pack
- A bundle of decision records, policy versions and approvals assembled to answer an audit question.
Standards and entities referenced
Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.
- 01 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 02 · Center for Internet SecurityCIS Critical Security Controls ↗Control 1 and 2 — inventory of assets and software — restated here for MCP servers and tools.
- 03 · NISTNIST SP 800-207 — Zero Trust Architecture ↗The policy decision point / policy enforcement point split this architecture borrows directly.
- 04 · NISTNIST AI Risk Management Framework ↗Govern-map-measure-manage; the vocabulary most enterprise AI risk programmes are written against.
- 05 · ISOISO/IEC 42001 — AI management systems ↗The management-system standard auditors increasingly map AI governance evidence against.
- 06 · ISOISO/IEC 27001 — Information security management ↗The ISMS baseline that agent-layer controls have to fit inside rather than beside.
- 07 · CNCFOpen Policy Agent — documentation ↗Reference implementation of decoupled policy decisions and policy as code.
- 08 · WikipediaTotal cost of ownership ↗Why licence price is the smallest term in a governance build-versus-buy decision.
- 09 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
- 10 · AxelosITIL 4 — change enablement ↗Established change-management vocabulary this article borrows for MCP estates.
Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). How to Evaluate an MCP Governance Platform: A Buyer’s Framework. Real Biz Digital. https://realbizdigital.net/insights/mcp-governance-platform/
Try the mechanics on a live server
To watch a real tools/list response, and see how much surface one server exposes, before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
Barzel Central Gateway is this layer, sold as a running product
Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so an evaluation costs an afternoon rather than a purchase order.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Community | Free | 1,000 tool calls/mo · full policy engine, registry, routing, audit | Evaluating the estate, or one team proving the path works |
| Starter | $10/mo | 10,000 calls/mo · everything in Community | One or two production agents against a handful of servers |
| Team | $79/mo | 100,000 calls/mo · routebooks, simulation, change impact | A platform team governing an estate of 5–20 servers |
| Business | $149/mo | 250,000 calls/mo · estate-wide evidence export | Multi-team governance with SIEM obligations |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
The five Barzel servers, and which problem each one is sold for
One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.
| Server | Sold for | Entry price | Where it sits |
|---|---|---|---|
| Barzel Central Gateway | Knowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidence | Free, then $10–$149/mo | Control plane — decides what may be reached, and by whom |
| BarzelVault | Stopping a specific dangerous action before it executes, with proof afterwards | $199–$3,999/mo | Decision point — evaluates the individual call before execution |
| BarzelOps | Running real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approval | Free, then $19–$199/mo | Execution layer — does the work the policy allowed |
| Barzel FinOps Atlas | Attributing AI spend to agents, tools and outcomes, then forecasting and capping it | Free, then $29–$799/mo | Economics layer — what the estate costs per outcome |
| Barzel Scripture Intelligence | A free, credential-free public MCP server to test clients and inspect real protocol traffic | Free, unmetered, no signup | Reference implementation — safe place to learn the protocol |
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.