MCP Governance · Deployment
Self-Hosted vs Managed MCP Gateway: How to Decide
Self-hosting is usually chosen for a reason that turns out not to require it. Here are the four reasons that genuinely do, the operational cost of each deployment model, and the six questions that settle the decision.
The short answer
Self-hosting an MCP gateway means you operate the control plane inside your own boundary and accept its availability, upgrade and on-call burden in exchange for keeping decision data and credentials within your perimeter. It is genuinely required in four situations — regulated data residency, air-gapped networks, contractual prohibition on third-party processing, and latency-critical co-location — and is otherwise an expensive way to buy a feeling. The honest comparison is not features. It is who carries the pager for a component that sits in the path of every agent action.
Summary for readers and answer engines
Reviewed 25 Aug 2026
- ▸Four models exist: fully managed, managed in your VPC, self-hosted in your cloud, and on-premise or air-gapped. Operational load roughly doubles at each step.
- ▸Four requirements genuinely mandate self-hosting: legally binding data residency, air-gapped networks, contractual prohibition on third-party processing, and sub-millisecond co-location with the upstream.
- ▸The most common stated reason — “our data cannot leave” — frequently dissolves on inspection, because what leaves is decision metadata rather than payloads, and metadata scope is configurable.
- ▸Self-hosting transfers seven duties to you: availability, upgrades, credential storage, backup and restore of decision history, capacity, on-call, and security patching of a component in the critical path.
- ▸Whatever you choose, the decision that matters more is the declared failure behaviour per risk class. A self-hosted gateway that fails open is worse than a managed one that fails closed.
Source: Mark Alex, Real Biz Digital — Self-Hosted vs Managed MCP Gateway: How to Decide (https://realbizdigital.net/insights/self-hosted-mcp-gateway/). Reproduce with attribution.
Key takeaways
- 01Establish precisely what data would cross the boundary before arguing about where the boundary should be. Frequently the answer changes the decision.
- 02Price the on-call rota. A component in the path of every agent action needs 24/7 coverage, and that is the largest line in the self-hosting case.
- 03Ask which functions degrade without outbound internet access. Some products lose threat intelligence, some lose licensing, some lose nothing.
- 04Insist on a documented backup and restore path for decision history. Evidence you cannot restore is evidence you do not have.
- 05Treat upgrade cadence as a security property. A self-hosted deployment three versions behind is carrying known issues you have chosen to keep.
- 06Air-gapped is a different product requirement from self-hosted, and conflating them wastes evaluation cycles.
Quick answers
One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.
- What is a self-hosted MCP gateway?
- An MCP governance control plane you deploy and operate inside your own infrastructure, so that decision data, policy and credentials remain within your perimeter and their availability becomes your responsibility.
- When is self-hosting genuinely necessary?
- Four cases: legally binding data residency, air-gapped networks, contracts prohibiting third-party processing of the relevant metadata, and latency requirements that demand co-location with the upstream system.
- What does self-hosting actually cost?
- Beyond compute, seven operational duties: availability, upgrades, credential storage, backup and restore of decision history, capacity planning, 24/7 on-call, and security patching of a critical-path component.
- What is a VPC deployment?
- The vendor’s software runs inside your cloud account and network, often with a managed control loop, so data stays in your boundary while the vendor retains responsibility for upgrades.
- Does self-hosting improve security?
- It changes the threat model rather than improving it. You remove third-party processing risk and add the risk that you patch late, misconfigure, or lack the operational maturity the component needs.
- What breaks in an air-gapped deployment?
- Typically licence validation, threat intelligence feeds, model-provider metadata and automatic updates. Ask for the explicit list rather than assuming the product simply works offline.
- Which decision matters more than deployment model?
- The declared failure behaviour per risk class. Where the software runs matters far less than what happens to agent traffic when it cannot render a decision.
Four deployment models, compared on what actually differs
Key facts
- ▸Operational load roughly doubles at each step down this table, and it is a standing cost rather than a project cost.
- ▸The VPC model is under-considered and often the right answer: it satisfies most data-boundary requirements while leaving software operation with the people who wrote the software.
- ▸Air-gapped is a distinct requirement, not an extreme version of self-hosted. It changes licensing, updates and threat-intelligence dependencies, and many products simply do not support it.
| Model | Data boundary | Operational load | Upgrade cadence | Right when |
|---|---|---|---|---|
| Fully managed (SaaS) | Decision metadata processed by vendor; payload handling configurable | Lowest — vendor carries availability and on-call | Continuous, vendor-driven | Most estates, most of the time |
| Managed in your VPC | Data stays in your account; vendor operates the software | Low to moderate — you own network and identity, vendor owns the software | Vendor-driven, on your maintenance window | Data residency requirements without operational appetite |
| Self-hosted in your cloud | Everything inside your perimeter | High — you own availability, patching, backup, on-call | Your cadence, your risk | Contractual prohibition on third-party processing |
| On-premise or air-gapped | No external dependency at all | Highest — plus offline licensing and update logistics | Manual, typically quarterly | Genuinely isolated networks |
Establish which model you actually need before shortlisting, because it eliminates vendors faster than any capability requirement.
The four reasons that genuinely require self-hosting
We hear roughly ten reasons for self-hosting. Four of them hold up under examination.
- ›“Our data cannot leave” — often only payloads are in scope, and payloads need not leave
- ›“Security requires it” — unless you patch faster than the vendor, this inverts
- ›“We already run everything ourselves” — a preference, not a requirement
- ›“Cheaper at our scale” — almost never, once on-call is priced
- ›Named regulation or contract clause covering decision metadata
- ›No egress path exists, at all
- ›Sub-processor addition genuinely unavailable
- ›Measured latency budget a hop cannot meet
Legally binding data residency
A regulator or a customer contract requires that specified data remain within a jurisdiction, and the decision metadata falls within that specification. This is a fact you can point at, not a preference.
Test it: name the regulation or clause, and identify precisely which fields it covers. Frequently the answer is payloads rather than decision metadata, which a configuration change resolves.
Air-gapped or severely restricted networks
Defence, critical national infrastructure, certain manufacturing environments. No outbound internet exists, so managed is not an option at any price.
Test it: confirm the network truly has no egress path, including through a proxy. Many “air-gapped” environments have a documented, monitored egress route.
Contractual prohibition on third-party processing
A customer contract or a data processing agreement forbids introducing a new sub-processor for this class of data, and renegotiation is not available in your timeframe.
Test it: check whether the vendor can be added as a sub-processor, and how long that takes. Sometimes the blocker is a six-week paperwork exercise rather than a prohibition.
Latency-critical co-location
The decision must be rendered within a budget that a network hop cannot satisfy — genuinely rare for agent workloads, where model inference dominates by two orders of magnitude.
Test it: measure. If model inference is 800ms and the decision hop is 20ms, latency is not your reason and something else is doing the arguing.
The purpose of this framing is not to argue against self-hosting. It is to make sure the reason survives the first difficult on-call week, because a deployment model chosen on preference gets abandoned the first time it costs a weekend.
Seven duties self-hosting transfers to you
These are the parts that do not appear in the comparison table and do appear in the following year.
- 01Availability. A governance control plane sits in the path of every governed action. Its availability target is your agents’ availability target, which usually means multi-zone and a tested failover.
- 02Upgrades. Policy engine changes, schema migrations and protocol updates. MCP is moving quickly; a deployment three versions behind is carrying known issues by choice.
- 03Credential storage. The gateway holds or brokers credentials to upstream systems. That makes it one of the highest-value targets in your estate, with the key management obligations that implies.
- 04Backup and restore of decision history. Evidence you cannot restore is evidence you do not have. This obligation grows with retention period and is routinely discovered during the first audit.
- 05Capacity. Agent load is bursty rather than diurnal. Sizing for the average produces refusals during the bursts that matter most.
- 06On-call. Twenty-four-hour coverage for a critical-path component. In most organisations this is the single largest cost in the self-hosting case and the one most often omitted.
- 07Security patching. Not merely applying patches, but tracking advisories for the component and its dependencies, and having a path to emergency deployment.
The cost line that decides it
on-call rota (1 in 4, 24/7) ≈ 4 engineers partially committed ≈ 0.6–1.0 FTE equivalent, permanently plus upgrade and patch effort ≈ 0.2–0.4 FTE total standing cost ≈ 0.8–1.4 FTE against a licence of tens of dollars per month
This is why self-hosting on cost grounds is almost always an arithmetic error. It is a legitimate choice for boundary reasons and a poor one for budget reasons.
Six questions that settle the decision
- 01Exactly which fields would cross the boundary in the managed model? Ask for the field list, not a category. Decision metadata and payload content are very different answers, and only one of them usually matters.
- 02Which functions degrade without outbound internet access? Licence validation, threat feeds, update checks, telemetry. Get the explicit list, in writing.
- 03What is the documented backup and restore procedure for decision history, and has it been tested? Restore, specifically. Backups that have never been restored are a hope.
- 04What is the upgrade path, and what breaks between versions? Schema migrations and policy-format changes are where self-hosted deployments stall and then stay stalled.
- 05What is the declared failure behaviour, and is it configurable per risk class? This matters more than deployment model. A self-hosted gateway that fails open on payment tools is worse than a managed one that fails closed.
- 06Who carries the pager, by name, and what is the rota? If this question has no answer, the self-hosting decision has not actually been made yet.
Question five is the one we would ask first if we could only ask one. Deployment model is a boundary question; failure behaviour is a safety question, and the second has more effect on outcomes than the first.
The hybrid pattern most estates land on
In practice the sharp distinction dissolves, and the arrangement below is what mature estates tend to converge on.
Decision point close to the action
The component that renders per-call decisions runs where the agents are, so latency and availability are local concerns and a network partition does not stop work.
Registry and policy authored centrally
Policy repository, registry and risk classification live in one place regardless of where enforcement runs. Distributing the source of truth is how estates diverge.
Evidence exported, not stored twice
Decision records are written locally and exported to your SIEM and long-term archive. The archive is yours; the format is standard; no vendor holds your only copy.
Credentials brokered, never mirrored
Upstream credentials stay in your secret manager. The gateway requests short-lived tokens rather than holding long-lived secrets, which shrinks the blast radius wherever it runs.
Configurable metadata scope
What is sent to any external service is a configuration decision with a documented default, so a residency requirement becomes a setting rather than a deployment model.
One declared failure policy
Fail-closed and fail-open sets defined once, in the policy repository, and honoured identically by every deployment. Consistency here matters more than topology.
This pattern is why we build Barzel Central Gateway to keep policy in a repository, evidence exportable in standard formats and credential handling brokered rather than stored: it makes the deployment question smaller, which is the useful outcome for everyone.
Next step
Make the deployment question smaller
Policy in a repository, evidence exported in standard formats, credentials brokered rather than stored: the more of those you have, the less the deployment model decides. Start on the free tier and see which of your boundary requirements actually bind.
Our position, stated plainly
We sell a governance product, and that shapes this page. Two disclosures.
- 01We benefit commercially when self-hosting is unnecessary, so read the “reasons that dissolve” section with that in mind. The four valid reasons are stated as we would state them to a customer we expected to lose.
- 02This page describes deployment trade-offs generally rather than enumerating our own supported topologies, which change; the product pages and our team are authoritative on what we currently support.
Frequently asked questions
What is a self-hosted MCP gateway?
An MCP governance control plane deployed and operated inside your own infrastructure, so decision data, policy and brokered credentials remain within your perimeter — and so its availability, upgrades, backups, capacity and on-call become your responsibility.
When is self-hosting an MCP gateway genuinely necessary?
In four situations: legally binding data residency covering the relevant decision metadata, air-gapped or severely restricted networks with no egress path, contractual prohibition on adding a third-party processor, and a measured latency budget that a network hop cannot satisfy.
What is the difference between self-hosted and VPC deployment?
In a VPC deployment the vendor’s software runs inside your cloud account and network, so data stays in your boundary, but the vendor retains responsibility for operating and upgrading the software. Self-hosting transfers that operational responsibility entirely to you.
Does self-hosting an MCP gateway improve security?
It changes the threat model rather than strictly improving it. You remove third-party processing exposure and take on the risks of patching late, misconfiguring a critical-path component, and holding high-value credentials in an environment that may have less operational maturity than the vendor’s.
What does self-hosting cost operationally?
Seven duties: availability engineering, upgrades and schema migrations, credential storage, tested backup and restore of decision history, capacity planning for bursty load, twenty-four-hour on-call, and security patching. Realistically 0.8 to 1.4 full-time-equivalent engineers, permanently.
What breaks in an air-gapped MCP gateway deployment?
Typically licence validation, threat intelligence feeds, automatic updates and any vendor-side telemetry. Ask for the explicit degradation list in writing, because air-gapped support is a distinct product capability rather than an extreme form of self-hosting.
Is self-hosting cheaper than a managed MCP gateway?
Almost never. Governance control-plane licences in this category are priced in tens of dollars per month at platform-team volumes, while the on-call rota alone for a critical-path component costs the equivalent of most of an engineer permanently.
What should I ask a vendor about data boundaries?
Ask for the field-level list of what would cross the boundary in the managed model, not a category description. Decision metadata and payload content are very different answers, and residency requirements frequently cover only the latter — which makes it a configuration decision.
Which matters more: deployment model or failure behaviour?
Failure behaviour. A self-hosted gateway that fails open on payment tools when it cannot render a decision is more dangerous than a managed gateway that fails closed. Deployment model is a boundary question; failure behaviour is a safety question.
What is the hybrid pattern for MCP governance deployment?
Render decisions close to the agents, author policy and hold the registry centrally, export evidence to your own SIEM and archive in standard formats, broker rather than mirror credentials, and declare one fail-closed and fail-open policy honoured identically by every deployment.
How should backup and restore be handled for decision history?
Require a documented and actually tested restore procedure, not merely a backup schedule. Retention obligations of several years mean the restore path will be exercised under audit pressure, and evidence you cannot restore is evidence you do not have.
How often should a self-hosted MCP gateway be upgraded?
Treat upgrade cadence as a security property and stay within one or two releases of current. MCP is evolving quickly, and a deployment several versions behind is carrying known protocol and policy-engine issues as a deliberate choice.
Glossary
- Self-hosted deployment
- Operating the governance control plane inside your own infrastructure, accepting its full operational burden.
- VPC deployment
- Vendor software running inside your cloud account and network, typically with vendor-managed operation.
- Air-gapped deployment
- Operation in a network with no external connectivity, requiring offline licensing and update paths.
- Data boundary
- The perimeter within which specified data must remain, defined by regulation or contract.
- Decision metadata
- The fields describing a policy decision — caller, agent, tool, outcome, rule — as distinct from the payload content of the call.
- Sub-processor
- A third party processing data on your behalf, whose addition is often contractually controlled.
- Declared failure behaviour
- The documented, testable statement of what happens to traffic when the decision point cannot respond.
- Credential brokering
- Requesting short-lived upstream tokens on demand rather than storing long-lived secrets in the gateway.
- Standing cost
- Permanent recurring engineering commitment, as opposed to one-off implementation effort.
- Upgrade cadence
- How frequently a deployment adopts new releases, treated here as a security property.
Standards and entities referenced
Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.
- 01 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 02 · NISTNIST SP 800-207 — Zero Trust Architecture ↗The policy decision point / policy enforcement point split this architecture borrows directly.
- 03 · ISOISO/IEC 27001 — Information security management ↗The ISMS baseline that agent-layer controls have to fit inside rather than beside.
- 04 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
- 05 · European UnionGDPR — Regulation (EU) 2016/679 ↗Lawful basis, data minimisation and processing records that agent estates inherit.
- 06 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
- 07 · Kubernetes projectKubernetes — cluster architecture ↗The canonical control-plane / data-plane separation, and the closest well-understood analogue.
- 08 · WikipediaTotal cost of ownership ↗Why licence price is the smallest term in a governance build-versus-buy decision.
Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). Self-Hosted vs Managed MCP Gateway: How to Decide. Real Biz Digital. https://realbizdigital.net/insights/self-hosted-mcp-gateway/
Try the mechanics on a live server
To watch a real tools/list response, and see how much surface one server exposes, before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
Barzel Central Gateway is this layer, sold as a running product
Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so an evaluation costs an afternoon rather than a purchase order.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Community | Free | 1,000 tool calls/mo · full policy engine, registry, routing, audit | Evaluating the estate, or one team proving the path works |
| Starter | $10/mo | 10,000 calls/mo · everything in Community | One or two production agents against a handful of servers |
| Team | $79/mo | 100,000 calls/mo · routebooks, simulation, change impact | A platform team governing an estate of 5–20 servers |
| Business | $149/mo | 250,000 calls/mo · estate-wide evidence export | Multi-team governance with SIEM obligations |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
The five Barzel servers, and which problem each one is sold for
One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.
| Server | Sold for | Entry price | Where it sits |
|---|---|---|---|
| Barzel Central Gateway | Knowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidence | Free, then $10–$149/mo | Control plane — decides what may be reached, and by whom |
| BarzelVault | Stopping a specific dangerous action before it executes, with proof afterwards | $199–$3,999/mo | Decision point — evaluates the individual call before execution |
| BarzelOps | Running real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approval | Free, then $19–$199/mo | Execution layer — does the work the policy allowed |
| Barzel FinOps Atlas | Attributing AI spend to agents, tools and outcomes, then forecasting and capping it | Free, then $29–$799/mo | Economics layer — what the estate costs per outcome |
| Barzel Scripture Intelligence | A free, credential-free public MCP server to test clients and inspect real protocol traffic | Free, unmetered, no signup | Reference implementation — safe place to learn the protocol |
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.