Buyer’s guide · evaluation
Best MCP Gateways in 2026: How to Choose One
We are not going to hand you a scored table of competitors. We sell one of these, and any table we produced would be marketing wearing a lab coat. What follows is the evaluation we would run if we were buying.
The short answer
MCP gateways fall into four categories: open-source proxies, extensions to existing API gateways, MCP-native governance platforms, and building it yourself. Choose on twelve criteria, of which five are outright disqualifiers — no per-identity authentication, no parameter-level policy, no record of refused calls, no way to test policy without production traffic, and any design an agent can bypass. Run the evaluation against your own traffic in a week rather than reading comparison tables. Including an honest account of where our own product does not fit.
Why there is no table here
Search for the best MCP gateway and you will find scored tables where the publisher wins. We publish one of these products, so any table we wrote would be the same artefact, and you would be right to discount it.
There is also a practical objection. This category is roughly eighteen months old. Products in it are changing monthly, several are pre-1.0, and a comparison written in August is materially wrong by November. A framework you can apply yourself has a longer shelf life than a snapshot of a moving field.
So: the four shapes these products come in, twelve criteria, five disqualifiers, and a week-long evaluation you can run against your own traffic. Where our product fails a criterion, that is stated in the last section rather than omitted.
Four categories, four different bets
Most of the decision is choosing a category. Within a category the products are more similar than the marketing suggests; across categories they are barely comparable.
| Category | What you get | What it costs you |
|---|---|---|
| Open-source MCP proxy | Transparent, self-hosted, free, auditable code. Usually solid on transport, connection multiplexing and basic allow-listing. | Policy is typically tool-level rather than parameter-level, approval workflow and tamper-evident audit are usually absent, and you own the operational burden. |
| Existing API gateway, extended | Reuses infrastructure your platform team already runs, monitors and trusts. Excellent at rate limiting, TLS and routing. | Its vocabulary is endpoints and methods, not tools and parameters. Approval workflow and parameter thresholds usually have to be bolted on — see AI gateway vs API gateway. |
| MCP-native governance platform | Built for this problem: parameter-level policy, per-identity entitlement, approval workflow, risk classification, audit designed as evidence. | A young vendor, a new dependency in your critical path, and pricing on a metric (calls) that is hard to forecast before you have traffic. |
| Build it yourself | Exact fit, no vendor, full control. Genuinely the right answer for a small estate with one unusual constraint. | The unglamorous 80% — token validation, approval UI, tamper-evident logging, rate limiting per identity, policy testing — is most of the work and none of the interest. |
A pattern worth knowing: teams that start with an open-source proxy and grow past roughly ten servers tend to end up needing parameter-level policy and an approval workflow, and those two are the expensive things to retrofit. Teams that start with an API gateway extension tend to be fine until the first approval requirement arrives.
Five disqualifiers
Check these before anything else. Any one of them present means the evaluation is over, whatever else the product does well — not because the product is bad, but because it cannot do the job you are buying it for.
- 01No per-identity authentication. If the gateway accepts a shared static key and forwards it, you cannot attribute a call, revoke one caller, or rate-limit an individual agent. Every other control degrades to guesswork.
- 02Policy on tool names only. Permitting a tool approves its entire parameter range. A transfer tool with no ceiling is a different tool at 10 and at 10 million, and a name-only policy cannot tell them apart.
- 03No record of refused calls. Refusals are the half that proves the control works and the half auditors ask about. A gateway that logs only what it permitted has no evidence of governing anything.
- 04No way to test policy without production traffic. If a rule can only be validated by enforcing it live, rules will not be changed, and unchanged rules are widened rather than tuned.
- 05Bypassable by design. If an agent can reach a server directly — because servers stay individually addressable and nothing enforces the path — the gateway governs the agents that chose to comply.
Twelve criteria worth scoring
Score each on your own traffic. The first four separate products meaningfully; the rest separate them at the margin, which still matters when two candidates are close.
Policy granularity
Tool name, parameter values, caller attributes, environment, target region, data classification present. Ask for the six inputs by name and ask to see a rule using each. Rules on tool names alone are the single most common gap.
Decision outcomes
Allow and deny is two outcomes for a problem that needs four. Dry-run and require approval both describe real decisions — policy engine outcomes.
Identity model
Per-agent identity, short-lived tokens, and preservation of the original human’s authority across the hop rather than collapsing everyone into one service account. Ask specifically about token exchange.
Evidence quality
One structured record per attempt including refusals, with rule id, approver, route and result. Tamper-evidence if you have a compliance obligation. Ask to see a real record, not a schema.
The test that separates audit-grade from log-grade: can you produce, for a named call last Tuesday, who authorised it and what the parameters were, in one query?
Approval workflow
Does the approver see the tool, the parameters and the caller, or a generic prompt? Generic prompts produce rubber stamps within a week, which is worse than no approval because it manufactures a record of oversight that did not happen.
Latency added, measured
Ask for a p95 figure and then measure it yourself on your own traffic. Policy evaluation should be single-digit milliseconds; anything above about 50ms p95 changes how interactive agents feel.
Failure behaviour
Fail-closed or fail-open, and configurable per risk class? Both are defensible; not having decided is not. Ask what happens to an in-flight approval when the gateway restarts.
Deployment model
SaaS, VPC or self-hosted. If regulated data flows through tool parameters, this may be the whole decision. Ask what leaves your network, and specifically whether parameters do.
Discovery filtering
Does it filter tools/list by calling identity? This affects security and token cost simultaneously, and surprisingly many products do not do it — see tool discovery.
Third-party server intake
Version pinning, tool-description diffing on upgrade, credential isolation per server. The supply-chain surface is where most estates are actually exposed.
Policy portability
Can rules be exported and read? Policy locked in a vendor UI is policy you cannot review, diff, or take with you.
Pricing shape
Per call, per server, per seat or flat. Per-call pricing on a system whose most common incident is a retry loop deserves a question about what happens during one.
A week-long evaluation, not a six-week bake-off
Five days, run against your own estate. This beats any comparison table because it measures the thing you actually care about: whether the product refuses what you need refused, on your traffic.
Point one agent at it
One real agent, one real server, in a non-production environment. You are testing whether the path works and what breaks in the client. Products that make this hard on day one make everything hard.
Write your hardest rule
Not a demo rule — the actual one you need. Usually something like: this agent may refund up to 500, above that a human decides, and never for accounts flagged in dispute. If it cannot be expressed, stop here.
Try to bypass it
Point the agent directly at the server. Use an expired token. Ask the model, in retrieved content, to use a tool it is not entitled to. Every one should be refused and every refusal should appear in the record.
This is the day that produces the real answer. Run it before you look at pricing.
Break it deliberately
Kill the gateway mid-call. Restart it with an approval pending. Burst well above the rate limit. You are looking for what happens, not for it to survive — documented failure beats surprising resilience.
Produce an audit answer
Pick a call from Tuesday and answer: who called it, with what parameters, under which rule, approved by whom, on which route. If that takes more than one query, the evidence model is log-grade rather than audit-grade.
Where Barzel Central Gateway fits, and where it does not
We sell in the MCP-native governance category. Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence; ten policy inputs including credential mode and target region; six enforcement outcomes; per-user OAuth/OIDC. The Community tier is free at 1,000 calls a month, which is enough to run the five-day evaluation above without a purchase order.
Where it fits well: an estate of several MCP servers where governance decisions need to be consistent, parameter-level policy is required, and somebody will eventually ask for evidence of refused actions. It pairs with BarzelVault, which holds the pre-execution decision, approval workflow and hash-chained audit for the actions that carry real consequence.
Where it does not fit, stated plainly. If you run one MCP server and one agent, this is overhead you do not need — harden the server instead. If your requirement is fully air-gapped self-hosting, ask us before assuming; our distribution is marketplace-hosted. If you need certification documents today, read our security page first: we are engineered around SOC 2 and ISO 27001 control objectives and we hold neither certificate, and we would rather you learn that here than in week three of a procurement cycle.
And if you are choosing between us and building it yourself for a two-server estate, building it yourself is often correct. The 80% that is unglamorous is smaller at two servers than at twelve.
Run the evaluation
Five days, on your own traffic, on the free tier
The Community tier of Barzel Central Gateway is free at 1,000 calls a month — enough to point one agent at it, write your hardest rule, and spend day three trying to bypass it. If it fails your evaluation, you will have a sharper set of criteria for the next candidate.
Frequently asked questions
What are the categories of MCP gateway?
Four: open-source MCP proxies, existing API gateways extended for MCP, MCP-native governance platforms, and building it yourself. Most of the decision is choosing the category — within a category the products are more similar than the marketing suggests, and across categories they are barely comparable.
What should disqualify an MCP gateway immediately?
Five findings: no per-identity authentication, policy expressed on tool names only, no record of refused calls, no way to test policy without enforcing it in production, and a design an agent can bypass by addressing servers directly. Any one of them means the product cannot do the job you are buying it for.
How do you evaluate an MCP gateway?
In five days on your own traffic. Day one, point one agent at it. Day two, write your hardest real rule. Day three, try to bypass it — direct connection, expired token, injected instruction. Day four, break it deliberately. Day five, answer an audit question about a specific call in one query.
Why does this page not include a comparison table of MCP gateway vendors?
Because we sell one of these products, so any table we produced would tell you about our marketing rather than their engineering. The category is also about eighteen months old and changing monthly, so a scored snapshot is materially wrong within a quarter. The criteria here are testable by you, against them and against us.
How much latency should an MCP gateway add?
Policy evaluation should be single-digit milliseconds. Above roughly 50ms at p95 it becomes noticeable in interactive agent workflows. Ask for the vendor’s figure, then measure it yourself on your own traffic — the two frequently differ.
Is an open-source MCP proxy enough?
For a small estate, often yes. The two things teams typically outgrow are parameter-level policy and approval workflow, and both are expensive to retrofit — so if you can already name a rule that needs a parameter threshold, factor that in before choosing.
Should we build our own MCP gateway?
For a two-server estate with one unusual constraint, frequently yes. The unglamorous 80% — token validation, approval UI, tamper-evident logging, per-identity rate limiting, policy testing — is most of the work and it scales with estate size, so the calculation shifts against building somewhere around ten servers.
What is the single best test of an MCP gateway’s audit quality?
Pick a specific call from last week and ask: who called it, with what parameters, under which rule, approved by whom, on which route. If the answer takes more than one query, the evidence model is log-grade rather than audit-grade.
Sources and further reading
This page cites standards and specifications rather than third-party product claims. We do not publish scored comparisons of other vendors: we sell a competing product, we cannot verify their internals, and a table we produced would tell you about our marketing rather than their engineering. The criteria below are testable by you, against them and against us.
- 01 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 02 · MCP projectModel Context Protocol — official documentation ↗Primary source for protocol structure, transports and the shape of a tools/list response.
- 03 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
- 04 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
- 05 · NISTNIST SP 800-207 — Zero Trust Architecture ↗The policy decision point / policy enforcement point split this architecture borrows directly.
- 06 · IETFRFC 8693 — OAuth 2.0 Token Exchange ↗The mechanism for preserving caller authority across a gateway hop.
- 07 · CNCFOpen Policy Agent — documentation ↗Reference implementation of decoupled policy decisions and policy as code.
- 08 · ISOISO/IEC 42001 — AI management systems ↗The management-system standard auditors increasingly map AI governance evidence against.
Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.
Try the mechanics on a live server
To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
Barzel Central Gateway is this layer, sold as a running product
Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so the evaluation costs an afternoon rather than a purchase order.
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.