Architecture · operations
Enterprise MCP Gateway: Architecture, Governance and Operations
A gateway that governs one team’s agents is a component. A gateway every agent in the company must pass through is infrastructure — and infrastructure is judged on the day it fails, not on the day it demos.
The short answer
An enterprise MCP gateway is a gateway operated as load-bearing infrastructure: inline and unavoidable, highly available with a declared failure mode per risk class, deployable where your data-residency obligations require, and instrumented so that a platform team can run it without the vendor. Enterprise-grade is not more policy features — it is the six operational properties that only start mattering once every agent depends on it. A reference architecture, three deployment models, failure design by risk class, and the TCO lines nobody quotes.
Key takeaways
- 01Enterprise-grade means operable, not feature-complete. The differentiator is what happens at 3am, not what happens in the demo.
- 02Inline or nothing. A gateway agents can route around governs only the agents that chose to comply.
- 03Decide fail-closed versus fail-open per risk class before an outage decides for you, uniformly, in production.
- 04The gateway becomes a single point of failure the moment it is unavoidable. That is the trade, and it has to be designed for rather than discovered.
- 05Deployment model is usually decided by data residency, not by preference — tool parameters carry regulated data.
- 06TCO is not the licence. Add the enforcement-point on-call, policy authorship, intake staffing and adapter maintenance.
What enterprise-grade actually means here
Vendors reach for “enterprise” to mean more features. In this category it means something narrower and harder: the gateway has become a dependency for every agent in the organisation, and it now inherits every obligation that comes with being in a critical path.
The features that win a bake-off — policy expressiveness, approval workflow, risk scoring — are table stakes and mostly comparable across serious products. What separates a gateway that survives two years of production from one that gets ripped out is a different list entirely, and none of it demos well.
Six operational requirements. Each is boring, each is decisive.
- 1. Unavoidable placement
- Agents cannot reach servers directly. Enforced by network reachability, not by policy documents or good intentions.
- 2. Declared failure behaviour
- For each risk class: proceed or stop when the gateway is unavailable. Written down, tested, and different per class.
- 3. Availability commensurate with fail-closed scope
- If irreversible actions fail closed, the gateway’s availability target becomes those workflows’ availability target.
- 4. Deployable inside your data boundary
- Tool parameters carry regulated data. Where the gateway runs is frequently a legal question rather than an architectural preference.
- 5. Operable without the vendor
- A platform team can deploy, upgrade, roll back, debug and capacity-plan it. Every escalation that requires the vendor is an outage extension.
- 6. Policy portable and reviewable
- Rules exportable as text, diffable, version-controlled. Policy locked in a console is policy you cannot review or take with you.
Reference architecture
Six components. The separation matters because it determines what fails independently and what can be scaled, replaced or audited on its own.
Ingress and identity termination
Terminates agent connections, validates the token on every request rather than at connection, and resolves the caller to an identity with roles and entitlements. Stateless, horizontally scalable, and the only component that needs to be reachable by agents.
Per-request validation is the detail that gets skipped. A long-lived session with connection-time-only validation means a revoked agent keeps working for hours.
Registry and routebook store
The authoritative record of approved servers, tools, owners, risk classes, entitlements and candidate routes. Read on nearly every request, written rarely — so cache aggressively and invalidate explicitly.
This is the component whose availability quietly gates everything. A gateway that cannot read its registry cannot make a decision, which is why it needs its own replica strategy rather than sharing a general-purpose database.
Policy evaluator
Deterministic, versioned, single-digit millisecond. Takes identity, tool, parameters, environment, target and recent behaviour; returns one of the enforcement outcomes. Must be the same code path used by simulation.
Approval broker
Holds calls above threshold, presents them with full context, records the decision. Needs durable state rather than held connections — a four-hour approval must survive a gateway restart without auto-approving or vanishing.
The component most often bolted on afterwards, and the one whose restart semantics cause duplicated side effects when it is.
Egress and route execution
Selects the route from the routebook, holds the credential for each backing server, forwards the permitted call, and normalizes the response. The only component holding mutating credentials, which makes it the one to isolate hardest.
Evidence writer
Emits one record per attempt, refusals included, hash-chained, to the store the security team already queries. Write path must not block the call path — but a dropped record is a governance failure, so buffer durably rather than dropping.
The trade here is explicit: never block execution on evidence write, never silently drop evidence. Durable local buffer with backpressure alerting is the design that satisfies both.
Three deployment models
Usually decided by one question: may tool parameters leave your network? Because they frequently carry customer records, payment details and regulated data, this is a compliance answer rather than an architecture preference.
| Model | Parameters leave your network? | Operating burden | Right when |
|---|---|---|---|
| Managed / SaaS | Yes | Lowest — vendor operates | No regulated data in parameters, or the vendor is in-scope and contracted for it |
| VPC / private deployment | No — control plane may be external, data plane is yours | Moderate — you run it, vendor supplies | Regulated data in parameters, but you want vendor-supplied upgrades |
| Self-hosted | No | Highest — you own everything | Air-gapped, sovereign, or the gateway is strategic enough to own outright |
Two questions to ask any vendor, in these words. “Do tool parameter values leave my network in any deployment model, including for logging, analytics or support?” — the qualifiers are where the surprising answers live. And “if our contract ends, what runs and what stops?” A gateway whose policy evaluation depends on a vendor endpoint is a gateway whose failure mode includes commercial disputes.
It is also worth being honest that self-hosting a governance layer means owning an availability target you may not want. See the failure design below before assuming it is the safe choice.
Availability and failure design
Making the gateway unavoidable makes it a single point of failure. That is not an argument against it — the alternative is ungoverned agents — but it is a design obligation that has to be discharged explicitly.
The central decision is fail-closed versus fail-open, and the mistake is deciding it once, uniformly. Uniform fail-closed blocks harmless reads during every deploy and gets switched off globally after the second incident. Uniform fail-open means an outage silently removes every control you have.
| Risk class | On gateway unavailability | Rationale |
|---|---|---|
| R0 read, non-sensitive | Proceed; log the ungoverned window | Blocking costs more than the residual risk. Reconcile the window afterwards |
| R1 read, sensitive | Stop, after a short grace for in-flight calls | Disclosure is not reversible either |
| W1 bounded write | Stop, short grace | Reversible, but reversal costs a real afternoon |
| W2 high-reach write | Stop, no grace | The specific case the gateway exists for |
| X irreversible | Stop, no grace, alert on-call | Never proceeds unevaluated. A pending approval surviving a restart must not auto-approve |
Three availability consequences follow from that table. If X-class actions fail closed, the gateway’s availability target is the availability target of every workflow containing one — so size the SLO from the business process, not from a generic number. Deploy with the ingress layer stateless and horizontally scaled so upgrades are rolling rather than disruptive. And rehearse the failure: kill the enforcement point mid-call for each risk class and confirm the behaviour matches what you wrote down. An untested failure mode is a hypothesis.
Rollout sequence
Estate-wide cutover is the failure pattern. It stalls on whoever owns the least-loved server and delivers nothing until it delivers everything. Onboard by agent instead: each step is independently useful and reversible.
- 01One agent, one server, non-production. Prove the path, measure added latency yourself, find what breaks in the client. Products that make this hard on day one make everything hard.
- 02Same agent, production, monitor only. Inline but permissive, emitting the evidence record. You are establishing the traffic baseline and finding the calls nobody knew about.
- 03Move one rule out of a server — and delete the local copy. The deletion is the step people skip. A rule in two places is two rules that will disagree.
- 04Enforce for that agent, per risk class. Start with X-class refusals only. Read the refusals daily for a week; they will teach you more about your estate than the inventory did.
- 05Remove direct network reachability to the servers that agent uses. Until this step the gateway is advisory. Verify by attempting a direct connection and confirming it fails.
- 06Repeat per agent, then retire the direct paths per server. Coverage grows continuously and every step can be rolled back independently.
Two milestones are worth tracking publicly: percentage of agents inline and percentage of servers with no direct path. The gap between them is your remaining bypass surface, and it is the number that determines whether the programme has actually governed anything.
Built on this thinking
Evaluate it on your own traffic, on the free tier
Barzel Central Gateway covers the six components: per-request identity, registry and routebooks, deterministic policy with simulation, approval brokering, route execution and SIEM-ready evidence. The Community tier is free at 1,000 calls a month — enough to run the rollout’s first two steps before any purchase decision.
Total cost of ownership
Licence cost is the visible line and rarely the largest. Four others are real, recurring, and absent from every quote we have seen — including our own, which is why they are listed here.
| Cost line | Typical shape | Notes |
|---|---|---|
| Licence or usage fee | Per call, per seat, or flat | Per-call pricing on a system whose commonest incident is a retry loop deserves a question about what happens during one |
| Enforcement-point on-call | A rota, indefinitely | This is the real cost of unavoidable placement, and it starts the day you go inline |
| Policy authorship and tuning | Continuous, a few hours weekly | Rules need owners, review dates and refusal-log reading. Unowned rules get widened |
| Intake staffing | Roughly two days per new server | Slower than two working days and teams route around you, which costs more than the staffing |
| Adapter and contract maintenance | Per third-party server, per upgrade | Normalization adapters break when vendors change. Budget for it or the estate silently de-normalizes |
The honest summary: the licence is often the smallest line, and a free or cheap gateway does not make the other four go away. Evaluate on operability, because operability is what you will be paying for.
Where we do not fit, and where a gateway is the wrong answer
For one MCP server and one agent, a gateway is overhead. Harden the server, scope its credential, bound its parameters and revisit when a second server appears. The threshold is the first moment a governance decision has to be made twice.
For a fully air-gapped requirement, ask before assuming: Barzel Central Gateway is distributed through the MCPize marketplace, and a genuinely sovereign deployment is a conversation rather than a download. If you need certification documents today, read our security page first — we are engineered against SOC 2 and ISO 27001 control objectives and hold neither certificate, and it is better to learn that here than in week three of procurement.
And a gateway does not make an agent correct. It bounds what a wrong decision can do; it does not improve the decision. Estates that treat the gateway as the whole safety story end up with well-governed agents doing the wrong thing efficiently, which is a process quality problem no enforcement point will surface.
Frequently asked questions
What makes an MCP gateway enterprise-grade?
Six operational properties rather than more features: unavoidable inline placement, declared failure behaviour per risk class, availability commensurate with what fails closed, deployability inside your data boundary, operability by a platform team without vendor escalation, and portable reviewable policy.
What components does an enterprise MCP gateway need?
Six: ingress with per-request identity validation, a registry and routebook store, a deterministic policy evaluator, an approval broker with durable state, an egress and route-execution layer holding the credentials, and an evidence writer emitting one hash-chained record per attempt.
Should an MCP gateway be managed, VPC or self-hosted?
Usually decided by whether tool parameters may leave your network, which is a compliance question since parameters frequently carry regulated data. Managed suits estates without regulated parameters; VPC keeps the data plane yours with vendor-supplied upgrades; self-hosted suits sovereign or air-gapped requirements at the highest operating burden.
Should an MCP gateway fail closed or fail open?
Per risk class, never uniformly. Non-sensitive reads proceed with the ungoverned window logged. Sensitive reads and bounded writes stop after a short grace for in-flight calls. High-reach and irreversible actions stop with no grace, and a pending approval surviving a restart must not auto-approve.
Does an MCP gateway become a single point of failure?
Yes, the moment it is unavoidable — that is the trade for governing anything. Discharge it explicitly: declare failure behaviour per risk class, size the availability target from the business processes whose X-class actions fail closed, keep ingress stateless for rolling upgrades, and rehearse failure by killing the enforcement point mid-call.
How should an enterprise roll out an MCP gateway?
By agent, not by server. One agent against one server in non-production; the same agent in production in monitor-only mode; move one rule out of a server and delete the local copy; enforce X-class refusals for that agent; remove direct network reachability; then repeat. Every step is independently useful and reversible.
When does a gateway stop being advisory?
When agents can no longer reach servers directly. Until direct network reachability is removed, the gateway governs only the agents that chose to route through it. Verify by attempting a direct connection and confirming it fails.
What is the total cost of owning an MCP gateway?
Five lines, of which the licence is often smallest: licence or usage fee, an indefinite on-call rota for the enforcement point, continuous policy authorship and tuning, intake staffing at roughly two days per new server, and adapter maintenance as third-party contracts change.
What should you ask a vendor about data handling?
Two questions in these words: do tool parameter values leave my network in any deployment model, including for logging, analytics or support — the qualifiers are where surprising answers live. And if our contract ends, what runs and what stops.
Is a gateway enough to secure an AI agent estate?
No. It bounds what a wrong decision can do; it does not make decisions better. Estates treating the gateway as the whole safety story end up with well-governed agents doing the wrong thing efficiently, which is a process-quality problem no enforcement point surfaces.
Glossary
- Enterprise MCP gateway
- A Model Context Protocol gateway operated as shared, load-bearing infrastructure: inline for every agent, highly available, deployable within a required data boundary, and operable by a platform team.
- Inline placement
- Positioning the gateway so that agents cannot reach MCP servers without passing through it, typically by removing direct network reachability to the servers.
- Failure mode by risk class
- A declared policy stating, for each tool risk class, whether calls proceed or stop when the enforcement point is unavailable.
- Data boundary
- The network and jurisdictional perimeter within which tool parameters, which frequently carry regulated data, are permitted to be processed.
- Total cost of ownership
- The full cost of operating a gateway, including on-call, policy authorship, intake staffing and adapter maintenance, not only the licence.
Sources and further reading
Architecture and availability practice follows established infrastructure engineering, cited below. The six operational requirements, the failure-mode-by-risk-class table and the TCO breakdown are our own, from building and operating Barzel Central Gateway.
- 01 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 02 · Kubernetes projectKubernetes — cluster architecture ↗The canonical control-plane / data-plane separation, and the closest well-understood analogue.
- 03 · NISTNIST SP 800-207 — Zero Trust Architecture ↗The policy decision point / policy enforcement point split this architecture borrows directly.
- 04 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
- 05 · IETFRFC 8693 — OAuth 2.0 Token Exchange ↗The mechanism for preserving caller authority across a gateway hop.
- 06 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.
- 07 · EU AI Act (unofficial consolidated text)EU AI Act — full text ↗Obligations around logging, human oversight and traceability for higher-risk systems.
- 08 · ISOISO/IEC 42001 — AI management systems ↗The management-system standard auditors increasingly map AI governance evidence against.
- 09 · CNCFOpen Policy Agent — documentation ↗Reference implementation of decoupled policy decisions and policy as code.
Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). Enterprise MCP Gateway: Architecture, Governance and Operations. Real Biz Digital. https://realbizdigital.net/insights/enterprise-mcp-gateway/
Try the mechanics on a live server
To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
Barzel Central Gateway is this layer, sold as a running product
Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so an evaluation costs an afternoon rather than a purchase order.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Community | Free | 1,000 tool calls/mo · full policy engine, registry, audit | Evaluating the estate, or a single team proving the path works |
| Starter | $10/mo | 10,000 calls/mo | One or two production agents against a handful of servers |
| Team | $79/mo | 100,000 calls/mo | A platform team governing an estate of 5–20 servers |
| Business | $149/mo | 250,000 calls/mo | Estate-wide governance with SIEM evidence and multi-team routing |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.