5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

Scale · architecture

MCP Tool Discovery: Finding the Right Tool at Scale

Every tool an agent can see costs tokens on every single turn and competes for the model’s attention. At three servers nobody notices. At thirty, the tool list is the largest fixed cost in the context window and the main cause of wrong tool selection.

By Mark Alex, FounderPublished 24 Aug 2026Updated 2 Sep 202612 min

The short answer

MCP tool discovery is how an agent learns which tools exist and picks one — by default a flat tools/list response containing every tool on every connected server. That breaks down past roughly 40 tools, because the list consumes context on every turn and selection accuracy falls as near-duplicate descriptions accumulate. The four fixes are identity-based filtering, progressive disclosure, semantic retrieval over a capability index, and strict namespacing. Filtering is the one to do first: it is also a security control.

How discovery works by default

An MCP client connects to each configured server and issues tools/list. Each server returns every tool it advertises, with a name, a natural-language description and a JSON Schema for its parameters. The client merges the responses, and the merged list goes into the model’s context. The model then chooses.

This is a good design and it is exactly right for one server. Two properties become expensive as the estate grows. First, the list is resident: descriptions and schemas occupy context on every turn of the conversation, not just the turn where a tool is called. Second, the list is undifferentiated: the model sees the same catalogue regardless of who is asking or what the task is.

Both properties are fine at ten tools and painful at a hundred. The failure is gradual, which is why it is usually diagnosed late — as a model-quality problem rather than an architecture one.

What tool overload looks like

The signature is a set of symptoms that individually look like other problems and collectively look like nothing else.

  • 01Wrong-tool selection rises, especially between similar tools. Two teams each ship a search_customer; the model picks whichever description reads better, not whichever is correct.
  • 02Prompts get longer to compensate. Somebody adds guidance about which tool to use when, which costs more context to fix a context problem.
  • 03Token cost per turn rises without task complexity rising. The tool list is a fixed cost paid on every turn, including the ones that call nothing.
  • 04Latency to first token climbs as the prompt grows, on every request.
  • 05New servers become politically difficult because everyone has learned that adding tools makes the agent worse — a symptom that quietly stops adoption.

In our experience the inflection is somewhere around 40 tools for general-purpose agents, earlier if descriptions are verbose or overlapping. It is worth measuring rather than assuming: log selection errors against tool count and the curve is usually obvious.

Four patterns that work

In rough order of effort-to-benefit. The first is close to free and is also a security control; the third is the one that scales furthest but needs infrastructure.

01 · Pattern 01

Filter tools/list by identity

Return only the tools the calling identity is entitled to. A finance agent sees the finance tools. This single change usually cuts list size more than every other pattern combined, because most estates entitle broadly and expose universally.

It is a security control wearing an efficiency costume: a tool the model never learns exists cannot be talked into existence by injected content. That makes it the rare change with a case for both the platform team and the security team — indirect prompt injection explains why discovery filtering matters more than it sounds.

02 · Pattern 02

Progressive disclosure

Present capability groups first — billing, CRM, documents, infrastructure — and return full schemas only for the group the model selects. Two round trips instead of one, in exchange for a resident list that is an order of magnitude smaller.

Group names have to be genuinely distinguishable. Groups called data and information reproduce the original selection problem one level up, and now with a round trip attached.

03 · Pattern 03

Semantic retrieval over a capability index

Embed tool descriptions, retrieve the top handful for the current intent, and present only those. This scales to thousands of tools and keeps the resident list tiny.

It also introduces a retrieval step that can be wrong, and an index that can go stale. Log what was retrieved alongside what was called — without that, a selection failure is indistinguishable from a retrieval failure and you will debug the wrong layer.

04 · Pattern 04

Strict namespacing

Prefix every tool with its owning domain and forbid duplicate bare names across the estate. Cheap, unglamorous, and it removes an entire class of ambiguity that no amount of retrieval sophistication fixes.

Namespacing is also what makes duplicates visible — a capability graph turns that visibility into a consolidation decision.

Which pattern, at which size

A rough guide rather than a rule. Estate size here means tools reachable by a single agent, not tools in the organisation.

Reachable toolsDo thisDo not bother with
Under 25Namespacing, and nothing else. A flat list is genuinely fine.Retrieval and progressive disclosure — both add failure modes you have no need for.
25–60Identity filtering plus namespacing. Most estates never need to go further.Semantic retrieval. The index maintenance is not repaid at this size.
60–200Add progressive disclosure by capability group.Bespoke ranking models.
200+Semantic retrieval over an indexed capability catalogue, with filtering underneath it.Presenting anything close to the full list, ever.

Note that filtering appears at every size above 25 and underneath every other pattern. Retrieval that searches tools the caller is not entitled to call is a retrieval system that occasionally recommends a refusal.

Discovery is an entitlement boundary

It is tempting to treat discovery as a performance concern and entitlement as a separate security concern enforced at call time. Enforce at call time regardless — discovery filtering is not a substitute. But the two are more entangled than they look.

A model that can see a tool will eventually attempt it, particularly if content it has just read suggests doing so. Every such attempt is a refusal you have to generate, log, and eventually explain to somebody reviewing refusal rates. More importantly, an advertised-but-refused tool tells an attacker exactly what exists and what it is called, which is free reconnaissance delivered by your own protocol handshake.

So: filter discovery and enforce at call time. Discovery filtering reduces the attack surface and the token bill; call-time enforcement is what actually holds. Treating either as sufficient alone is the mistake.

This is also why discovery belongs in the control plane rather than in each server — only a layer that knows the calling identity can filter for it. Control plane responsibilities covers the boundary.

Built on this thinking

Filtered discovery needs a layer that knows who is asking

Barzel Central Gateway resolves the calling identity before discovery, returns only entitled capability, and routes the selected call to the right server — so the resident tool list shrinks and the refusal surface shrinks with it. Tool discovery, route selection and policy evaluation as one pass.

Four mistakes worth naming

Fixing selection with prompt text

Adding guidance about tool choice spends context to fix a context problem, and it degrades as the list grows.

Deduplicating by deleting

Two similar tools usually reflect two real requirements. Namespace and route between them rather than deleting the one with the quieter owner.

Retrieval without observability

If you do not log retrieved-versus-called, every failure looks like a model failure and you will tune the wrong layer for a month.

Discovery filtering as the only control

A filtered list is not an enforcement point. Refuse at call time as well, and test that the refusal holds when the model is told to try anyway.

Frequently asked questions

What is MCP tool discovery?

The process by which an agent learns which tools are available and selects one. By default the client issues a tools/list request to each connected server and merges the responses into the model’s context, where the full list stays resident for the whole conversation.

How many MCP tools is too many?

In our experience selection accuracy starts degrading around 40 tools reachable by a single agent, earlier if descriptions overlap or are verbose. The number matters less than the trend — log wrong-tool selections against reachable tool count and the inflection is usually visible.

What is MCP tool overload?

The condition in which the number of tools presented to a model degrades selection accuracy and consumes a material share of the context window before any work begins. It shows up as wrong-tool selection between similar tools, rising token cost per turn, and longer prompts written to compensate.

How do you reduce MCP tool token usage?

Filter tools/list by calling identity first — it usually cuts the list more than every other technique combined. Then apply progressive disclosure by capability group, and semantic retrieval only above roughly 200 reachable tools. Strict namespacing costs nothing and removes an entire class of ambiguity.

What is progressive tool discovery?

A pattern in which the agent is first shown a small set of capability groups — billing, CRM, documents — and receives full tool schemas only for the group it selects. It trades one extra round trip for a resident list an order of magnitude smaller.

Is filtering tool discovery a security control?

Yes, though not a sufficient one. A tool the model never learns exists cannot be talked into existence by injected content, and an advertised-but-refused tool hands an attacker free reconnaissance. Filter discovery and still enforce entitlement at call time; neither substitutes for the other.

Should tool discovery be handled by the server or the gateway?

The gateway, because only a layer that knows the calling identity can filter for it. Individual servers cannot filter by entitlement they do not know about, and per-server filtering drifts apart across an estate.

Does semantic tool retrieval introduce new failure modes?

Yes. Retrieval can return the wrong candidates and the index can go stale. Log what was retrieved alongside what was called, or a selection failure is indistinguishable from a retrieval failure and you will debug the wrong layer.

Sources and further reading

The protocol mechanics below are documented in the specification. The scaling thresholds and pattern comparison are our own measurements and operating judgement, from running servers with tool counts from 25 to 54; treat the numbers as order-of-magnitude guidance, not benchmarks.

  1. 01 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  2. 02 · MCP projectModel Context Protocol — official documentation ↗Primary source for protocol structure, transports and the shape of a tools/list response.
  3. 03 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
  4. 04 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  5. 05 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.
  6. 06 · Simon WillisonPrompt injection — ongoing series ↗The most consistently updated practitioner record of the attack class.
  7. 07 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.

Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.

Try the mechanics on a live server

To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

Barzel Central Gateway is this layer, sold as a running product

Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so the evaluation costs an afternoon rather than a purchase order.

CommunityFree1,000 calls/mo
Starter$10/mo10,000 calls/mo
Team$79/mo100,000 calls/mo
Business$149/mo250,000 calls/mo

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.