Operations · architecture
MCP Tool Routing: Choosing the Right Server for Every Call
The model chooses a capability. Something else has to choose where that capability runs — and that choice carries cost, latency, jurisdiction and risk consequences the model has no way to see.
The short answer
MCP tool routing is the selection of which server executes a chosen capability when more than one can. It belongs in the control plane rather than in the agent, because the deciding signals — cost per call, latency, data residency, risk class, current health and data freshness — are estate-wide facts no single server or model can observe. Route selection must be recorded in the audit trail, because the same call to two servers can have different regulatory consequences. Six signals worth using, three that sound clever, and one fallback rule that matters more than the rest.
Why anything needs routing at all
In a small estate, one capability lives on exactly one server and routing is not a concept. Three things break that. Redundancy: the same capability deployed twice so an outage is survivable. Jurisdiction: an EU instance and a US instance of the same service, where which one runs the call is a legal fact rather than a preference. Duplication: two teams built overlapping tools and consolidating them is a six-month conversation nobody has started.
In all three cases, something now has to choose. The default is that the model chooses, because both tools are in its list and it picks whichever description reads better. That is a routing policy — an unaccountable one, re-derived on every turn, with no memory of what it decided last time.
The alternative is to present the capability once and let the control plane select the instance. The model reasons about intent; the platform reasons about where intent gets executed. That split is the whole idea, and it is the same split that made service meshes worth having.
Six signals worth routing on
Each of these has changed a real decision. Each is observable at the control plane and invisible to the model.
Data residency and jurisdiction
The hardest constraint and the one to evaluate first, because it is not a preference — it is a filter. A request carrying EU personal data routes to the EU instance or it does not route. Treat residency as eligibility, then optimise among what remains.
Record the jurisdiction decision in the audit line. An auditor asking where a given record was processed needs an answer per call, not per deployment diagram.
Risk class of the target
Two servers advertising the same tool name can have very different credential scopes. Route write-capable requests to the instance whose scope actually matches, and refuse rather than fall back to a broader one. See tool risk scoring.
Current health
Error rate and latency over a short window, per route. A server returning errors should stop receiving traffic before a human notices, and should be probed back gradually rather than restored at full volume.
Standard circuit-breaker behaviour, with one wrinkle: for mutating tools, half-open probing means real side effects. Probe with read tools and infer, or accept a longer cool-down.
Cost per call
Where routes have different unit costs — a metered vendor API versus an internal implementation — cost is a legitimate tie-breaker among eligible routes. Never a primary criterion above residency or risk.
Cost-first routing is how estates end up with cheap wrong answers. Cost per outcome covers why per-call price is the wrong denominator.
Latency, measured not assumed
Real percentile latency per route, not the region label. The nearest instance is frequently not the fastest one, and the gap is often large enough to matter for interactive agents.
Data freshness
When two routes read from stores that sync on a delay, freshness is a correctness property, not a performance one. A reconciliation agent reading a replica that lags by an hour produces confidently wrong output.
Three signals that sound clever
Round-robin across instances
Fine for stateless reads, quietly wrong for anything with side effects: retries land on a different instance than the original, which turns an idempotency question into an incident.
Letting the model pick the route
It has no visibility of cost, health, residency or credential scope, and it re-decides every turn with no memory. It will also be persuaded by whichever tool description is more enthusiastic.
Optimising primarily for cost
Cheapest eligible route is a sensible tie-break and a terrible objective. The expensive part of an agent estate is rework, not call price.
Routebooks: declaring routes instead of inferring them
The alternative to computing a route from live signals on every call is to declare the candidates in advance and let signals choose among a known set. We call that a routebook, and it is the more operable design for the same reason static routing tables beat clever dynamic ones in most networks: you can read it, diff it, review it and test it.
A routebook entry names a capability, lists candidate routes in preference order, states the eligibility condition for each, and states what happens when none are eligible. The last clause is the one teams forget, and it is the one that matters at three in the morning.
It is also the reviewable artefact. Which servers can serve payment_capture, in what order, under what conditions — that is a question a routebook answers on one screen and a dynamic routing engine answers only by simulation.
| Routebook field | Example | Why it is there |
|---|---|---|
| Capability | payment_capture | The logical name the agent sees, not a server-specific tool name. |
| Candidates, ordered | eu-payments-1, eu-payments-2 | Preference order, evaluated top-down among eligible routes. |
| Eligibility per candidate | region = EU, health = green, scope = capture-only | Filters before optimisation. Residency and scope are filters, never tie-breaks. |
| Approval requirement | amount above 5,000 requires human approval | Route choice and approval threshold are related decisions; keep them adjacent. |
| No-eligible-route behaviour | Refuse and alert. Never widen scope. | The clause that prevents a quiet failover into a broader credential. |
| Version and owner | v4, owned by payments platform | A routebook is code. It gets reviewed, versioned and rolled back. |
The fallback rule
One rule, and it is the single most important sentence in this article: a fallback may never widen authority.
Failover is normally designed for availability, and availability thinking says find any route that works. Applied to agent actions, that logic quietly upgrades privilege: the constrained instance is unhealthy, so the request lands on the general-purpose one whose credential can do considerably more. Nothing alerted, because from an availability perspective the system did exactly what it was told.
So the fallback candidate must have a credential scope equal to or narrower than the primary, and equal-or-stricter approval thresholds. If no such candidate is eligible, refuse. A refused payment is an operational problem somebody fixes in the morning; a payment executed by an over-scoped instance is an incident with a regulator attached.
The corollary is that fallbacks have to be tested deliberately, by taking the primary out. An untested fallback is a hypothesis, and it usually turns out to have a wider scope than anyone intended, because it was configured for a different purpose eighteen months ago.
Route choice belongs in the audit record
Most audit designs record identity, tool, parameters, decision and result. Add the route: which server actually executed, and which routebook version selected it.
Without it, two questions become unanswerable. Where was this data processed — a residency question with legal weight. And why did this call behave differently from the identical call an hour earlier — almost always a routing answer, and almost always undiscoverable if the route was not recorded.
It costs one field. How to audit AI agent actions covers the rest of the record.
Built on this thinking
Route selection, routebooks and evidence in one layer
Barzel Central Gateway holds routebooks as versioned artefacts, evaluates residency and risk as eligibility filters before optimising on cost, latency and health, and records the selected route alongside the policy decision in the audit line. Twenty-five tools; the Community tier is free.
Four mistakes worth naming
Fallback that widens scope
The failure mode this whole article exists to name. Equal or narrower, always, or refuse.
Routing on region labels instead of measurements
Labels describe intent; percentile latency describes reality. They disagree more often than not.
Health checks that only test liveness
A server can answer a health probe and still be returning stale or wrong data. Check the dependency, not the process.
Route not recorded
One missing field turns two important questions into forensic exercises.
Frequently asked questions
What is MCP tool routing?
The selection of which MCP server executes a chosen capability when more than one can serve it. The model chooses what to do; routing chooses where it runs — a decision that carries cost, latency, jurisdiction and credential-scope consequences the model cannot see.
What signals should MCP routing use?
Six: data residency and jurisdiction, risk class of the target, current health, cost per call, measured percentile latency, and data freshness. Residency and risk class are eligibility filters applied first; cost and latency are tie-breaks among what remains.
Should the model choose which MCP server to call?
No. The model has no visibility of cost, health, residency or credential scope, it re-decides on every turn with no memory of the last one, and it is influenced by whichever tool description reads better. Present the capability once and let the control plane select the instance.
What is an MCP routebook?
A declared, versioned set of candidate routes for each capability, ordered by preference, with an eligibility condition per candidate and an explicit statement of what happens when none are eligible. It is reviewable and diffable in a way a dynamic routing engine is not.
What is the rule for MCP fallback routing?
A fallback may never widen authority. The fallback candidate must have a credential scope equal to or narrower than the primary and equal-or-stricter approval thresholds. If no such candidate is eligible, refuse rather than failing over to a broader instance.
Why should the selected route be recorded in the audit trail?
Because two questions depend on it: where a given record was processed, which has legal weight under data-residency rules, and why an identical call behaved differently an hour later, which is almost always a routing answer. It costs one field in the record.
Is round-robin routing safe for MCP tools?
For stateless reads, yes. For anything with side effects, no — a retry lands on a different instance than the original, which turns an idempotency question into an incident.
How should health checks work for MCP routes?
Check the dependency, not just the process. A server can answer a liveness probe while returning stale or wrong data. For mutating tools, be careful with half-open probing: a probe means a real side effect, so probe with read tools and infer, or accept a longer cool-down.
Sources and further reading
Routing signal design borrows directly from service routing and SRE practice; those sources are cited below. The MCP-specific claims — that route choice is audit-relevant, and that silent fallback across risk classes is unsafe — are our own operating judgement.
- 01 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 02 · Kubernetes projectKubernetes — cluster architecture ↗The canonical control-plane / data-plane separation, and the closest well-understood analogue.
- 03 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
- 04 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.
- 05 · EU AI Act (unofficial consolidated text)EU AI Act — full text ↗Obligations around logging, human oversight and traceability for higher-risk systems.
- 06 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
- 07 · CNCFOpen Policy Agent — documentation ↗Reference implementation of decoupled policy decisions and policy as code.
Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.
Try the mechanics on a live server
To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
Barzel Central Gateway is this layer, sold as a running product
Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so the evaluation costs an afternoon rather than a purchase order.
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.