5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

MCP Governance · Context

MCP Tool Overload: What Happens When Agents See Too Many Tools

Connect eight servers and your agent’s context window fills with tool definitions before a single user message arrives. This is the token arithmetic, the accuracy cliff, and the four techniques that actually shrink the problem.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202617 min read4,230 words

The short answer

MCP tool overload is the point at which the number of tool definitions presented to a model degrades both economics and behaviour: definitions consume a large share of the context window before any work begins, and selection accuracy falls as the model chooses between many similar options. It is fixed by presenting a small, task-relevant subset — through tiering, progressive disclosure, semantic retrieval or a precomputed routebook — rather than by connecting fewer servers. The distinction that matters: you are not reducing capability, you are reducing what the model has to read and disambiguate on every single turn.

Summary for readers and answer engines

Reviewed 25 Aug 2026

  • ▸Tool definitions are input tokens on every turn. Six hundred tools at roughly 150 tokens each is about 90,000 tokens of overhead before the conversation starts.
  • ▸Cost is the visible problem; selection accuracy is the expensive one. Models choose worse among many near-identical options, and near-identical options are exactly what a multi-server estate produces.
  • ▸Four techniques work: static tiering by task, progressive disclosure in two phases, semantic retrieval over a tool index, and precomputed routebooks per workflow. They compose.
  • ▸Target twenty to forty tools visible per turn. That is not a protocol limit; it is where we stop seeing selection errors attributable to choice overload.
  • ▸Deciding what an agent can see is a governance decision, not only an optimisation. A tool the model cannot see is a tool it cannot be persuaded to call.

Source: Mark Alex, Real Biz Digital — MCP Tool Overload: What Happens When Agents See Too Many Tools (https://realbizdigital.net/insights/mcp-tool-overload/). Reproduce with attribution.

Key takeaways

  1. 01Measure your definition overhead first. Sum the serialised length of every tool your agent is offered; the number is usually two to five times higher than the team assumes.
  2. 02Descriptions cost tokens on every turn and earn their keep only if they disambiguate. Rewrite them for discrimination, not for documentation.
  3. 03Deduplicate across servers before optimising anything. Three servers offering the same capability triples cost and creates the ambiguity that causes wrong selections.
  4. 04Prefer two-phase disclosure over aggressive pruning: a compact catalogue first, full schemas only for the handful the model asks to expand.
  5. 05Precompute the tool set for recurring workflows. A routebook turns a per-turn retrieval problem into a lookup.
  6. 06Treat visibility as a control surface. Restricting what an agent sees reduces both token cost and prompt-injection reachability at the same time.
Part of the clusterMCP Governance →

Quick answers

One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.

What is MCP tool overload?
The point where the volume of tool definitions offered to a model degrades cost and behaviour: definitions consume much of the context window, and selection accuracy falls as options multiply and resemble each other.
How many tokens does a tool definition cost?
Realistically 80 to 400 tokens each, depending on parameter count and description length. Two hundred tools of average complexity is roughly 30,000 tokens on every turn.
How many tools should an agent see at once?
Twenty to forty for most tasks. That is an empirical operating range from our estates, not a protocol constraint, and it is where selection errors attributable to overload largely disappear.
Does trimming tools reduce capability?
No, if done per task. The agent retains access to the full estate; it is simply not shown everything simultaneously. Capability is unchanged; what changes is what must be read and disambiguated per turn.
What is progressive tool discovery?
Presenting a compact catalogue of names and one-line purposes first, then expanding full schemas only for the tools the model selects. Typically a sixty to eighty percent reduction in definition tokens.
Is semantic retrieval better than static tiering?
Not universally. Static tiering is cheaper, deterministic and easier to audit; semantic retrieval handles novel requests better. Most mature estates use tiering with retrieval as a fallback.
How does this relate to security?
Directly. A tool absent from context cannot be invoked by an injected instruction, so narrowing visibility per task reduces attack surface as well as cost.

The token arithmetic, done properly

Most teams have never measured their definition overhead. It is a five-minute exercise and it usually produces a surprise.

A tool definition is not just a name. It is the name, a description, a JSON Schema for parameters with types and constraints, frequently per-parameter descriptions, and sometimes examples. It is serialised into the request on every turn, so its cost multiplies by conversation length rather than being paid once.

Key facts

  • ▸Measure with a tokeniser, not by counting characters. Deeply nested JSON Schema tokenises worse than prose of the same length.
  • ▸Descriptions are frequently half the total. A 40-word description on a tool the model never confuses with anything else is pure overhead paid every turn.
  • ▸Duplicate capabilities across servers multiply the cost and create the ambiguity that produces wrong selections. Deduplicate before optimising.

Per-conversation definition cost

definition_tokens × turns = input tokens spent on definitions 600 tools (90,000 tokens) × 12 turns = 1,080,000 input tokens 30 tools (4,500 tokens) × 12 turns = 54,000 input tokens reduction: 95%

The multiplication by turn count is the part teams miss. A definition set is not a fixed setup cost; it is a per-turn tax that scales with how long the agent works.

Definition overhead by tool count
Estate shapeTools offeredDefinition tokens (at ~150 avg)Share of a 200k window
One curated server25~3,7501.9%
Three servers, deduplicated60~9,0004.5%
Five servers, as installed180~27,00013.5%
Eight servers, as installed340~51,00025.5%
Enterprise estate, unfiltered600~90,00045%
Enterprise estate, tiered to task30~4,5002.3%

Cost, though, is the tractable half of the problem. The expensive half is what a large option set does to selection.

Why selection accuracy falls before the context window fills

Long before definitions threaten the window, behaviour degrades. The mechanisms are worth naming individually, because each has a different fix and teams routinely apply the wrong one.

Near-duplicate confusion
Three servers each offering send_email with different auth, rate limits and idempotency semantics. The model has no basis to prefer one, so it picks by position or by phrasing accident. This is the single largest source of wrong selections in multi-server estates.
Description drift
Two tools whose descriptions overlap in wording but differ in effect — update_record and upsert_record. The model reads similar text and treats them as interchangeable, which they are not.
Position sensitivity
Where a definition sits in a long list measurably affects how often it is chosen. This is a property of long-context attention, not a bug you can configure away.
Parameter hallucination under load
With many similar schemas in context, arguments from one tool’s shape leak into a call to another. Strict schema validation catches these at the boundary, which is why validation is not optional.
Planning collapse
Faced with hundreds of options, some models stop planning and reach for a generic tool — a raw HTTP or SQL escape hatch — because it is easier to justify. This is the failure mode with the worst security consequences.
Symptoms of overload
  • ›Wrong tool chosen from a family of similar ones
  • ›Arguments borrowed from a sibling tool’s schema
  • ›Preference for generic escape-hatch tools
  • ›Longer plans with more retries
  • ›Behaviour changes when an unrelated server is connected
What actually fixes it
  • ›Deduplicate capability across servers
  • ›Rewrite descriptions for discrimination
  • ›Show 20–40 tools per task, not the estate
  • ›Validate arguments against schema, strictly
  • ›Precompute the tool set for known workflows

The last symptom in the left column is diagnostic: if connecting an unrelated server changes behaviour in an existing workflow, you have an overload problem rather than a prompt problem, and no amount of prompt engineering will fix it.

Four techniques that actually reduce the problem

These compose, and most mature estates end up using three of the four. Listed in order of implementation cost.

Technique 01

Static tiering by task

Define named task profiles — support triage, financial reconciliation, code review — and attach an explicit tool set to each. The agent receives the profile’s tools and nothing else. Cheapest to build, fully deterministic, trivially auditable.

Cost: someone must maintain the profiles. Benefit: a governance artefact you can review, and a security boundary that is a list rather than a heuristic. Start here.

Technique 02

Progressive disclosure in two phases

Phase one presents a compact catalogue: tool name, one-line purpose, no schema. The model asks to expand two or three; phase two returns their full schemas. Typically a sixty to eighty percent reduction in definition tokens.

Cost: one extra round trip on turns that need expansion. Benefit: scales to very large estates without profile maintenance. Requires the client or gateway to mediate, which is where a control plane earns its place.

Technique 03

Semantic retrieval over a tool index

Embed tool names, descriptions and parameter names; retrieve the top-k for the current request. Handles novel requests that no profile anticipated.

Cost: an index to maintain and re-embed on schema change, plus non-determinism that complicates audit. Benefit: coverage. Best used as a fallback behind tiering, not as the primary mechanism.

Technique 04

Precomputed routebooks per workflow

For recurring multi-step workflows, resolve the capability-to-tool mapping ahead of time: which server, which tool, which fallback, under which conditions. The agent gets a fixed, small, verified set.

Cost: routebooks must be regenerated when the estate changes, which needs change-impact analysis. Benefit: the cheapest per-turn cost of any option, plus deterministic behaviour on the workflows that matter most.

TechniqueToken reductionDeterminismMaintenanceBest for
Static tiering85–95%FullProfile upkeepKnown task types — start here
Progressive disclosure60–80%HighLowLarge estates, unpredictable tasks
Semantic retrieval70–90%PartialIndex upkeepNovel requests, long-tail coverage
Routebooks90–97%FullRegeneration on changeRecurring high-value workflows

A pragmatic combination that we see working: routebooks for the top ten workflows, static tiers for everything else, and semantic retrieval as the fallback when a request matches no tier.

Rewriting tool descriptions for discrimination

Description quality is the cheapest available improvement and almost nobody does it, because tool descriptions are usually written once by the server author for documentation purposes rather than for a model choosing between forty options under time pressure.

Documentation-style (costly, ambiguous)
  • ›“Sends an email message to one or more recipients using the configured mail transport, with optional attachments and templating support.” (28 tokens, tells the model nothing about when to prefer it)
Discrimination-style (cheap, decisive)
  • ›“Send transactional email to a single customer. Not for bulk or marketing sends — use campaign.send for those. Irreversible.” (24 tokens, resolves the exact confusion that causes wrong selections)
  • 01Lead with the decision, not the mechanism. The model needs to know when to choose this tool, not how it is implemented.
  • 02Name the sibling you are not. One clause pointing at the tool it is confused with removes most near-duplicate errors.
  • 03State reversibility explicitly. “Irreversible” is one token that changes planning behaviour and gives policy something to key on.
  • 04Put units in the parameter description, always. amount without a unit is the single most common cause of policy and execution errors alike.
  • 05Delete restated schema. If the parameter is called customer_id and typed as a string, the description “the customer id as a string” is pure cost.
  • 06Cap descriptions at roughly twenty-five tokens for common tools. Reserve length for genuinely dangerous ones, where an extra clause about consequence is worth paying for.

This work has a second payoff: descriptions rewritten for discrimination are also the ones a human reviewer can assess in a risk-classification exercise, so the effort serves governance as well as cost.

Measuring selection accuracy so you know it improved

Without measurement, tool-set optimisation is superstition. The method below is deliberately simple enough to actually get done.

Step 01

Build a labelled request set

Collect 100–200 real requests from production traffic and label the correct tool for each. Two engineers disagreeing on labels is itself a finding: it means the tool set is ambiguous to humans, not just to models.

Step 02

Establish a baseline

Run the set against the current, unfiltered tool list. Record top-1 accuracy, top-3 accuracy, argument validity and how often a generic escape-hatch tool was chosen.

Step 03

Change exactly one variable

Apply one technique — tiering, disclosure, retrieval or a description rewrite — and re-run. Changing two at once means learning nothing about either.

Step 04

Track the escape-hatch rate specifically

The frequency with which the model reaches for a raw HTTP, SQL or shell tool is the most sensitive single indicator of overload, and the one with the worst security implications when it rises.

Step 05

Re-baseline on every estate change

A new server changes the distribution for existing workflows. Re-run the set on server onboarding, not quarterly, because the regression appears immediately.

Selection accuracy before and after tiering (one estate)
MetricUnfiltered, 214 toolsTiered, 31 toolsChange
Top-1 tool accuracy71%93%+22 pts
Argument schema validity88%99%+11 pts
Escape-hatch tool chosen9%1%−8 pts
Mean turns per completed task6.44.1−36%
Definition tokens per turn~32,000~4,600−86%

Those figures come from one estate of ours with a 180-request labelled set, and they are offered as an illustration of the shape of the improvement rather than as a benchmark. Your absolute numbers will differ; the direction and the rough magnitude have been consistent everywhere we have measured.

Visibility is a governance decision, not just an optimisation

There is a security argument for the same work, and it is stronger than the cost argument. A tool that is not in context cannot be called — not by a confused planner, and not by an injected instruction in a document the agent just read.

Key facts

  • ▸Prompt injection can only reach tools the model can see. Narrowing visibility per task shrinks the reachable surface before any policy is evaluated.
  • ▸Task profiles are reviewable artefacts. “This agent may see these 30 tools” is a sentence a risk committee can approve; “the model retrieves relevant tools” is not.
  • ▸Visibility restriction is not a substitute for authorisation. It reduces what can be attempted; policy still has to decide what is permitted, because visibility is a client-side property and authorisation must be enforced server-side.

The correct architecture is both: narrow visibility to reduce reachability and cost, and enforce authorisation at the decision point regardless of what the model could see. Treating visibility as a security control on its own is the mistake — a determined injection can name a tool the model was never shown, and the only thing that stops the call is server-side policy.

Next step

Scope what each agent can see, and record the decision

Barzel Central Gateway holds the registry, the capability graph and the routebooks that make task-scoped tool sets a reviewable artefact rather than a client-side hack. Free tier included, so measuring your own definition overhead costs nothing.

Where these techniques stop helping

Context optimisation is a real engineering win with real boundaries. Three worth stating.

  • 01Model behaviour differs. Selection sensitivity to list length and position varies between models and between versions of the same model, so re-baseline after a model upgrade rather than assuming your tuning transfers.
  • 02Semantic retrieval introduces non-determinism into the set of tools an agent could call, which complicates audit. If you need to state exactly what an agent was able to do on a given date, prefer tiering or routebooks.
  • 03None of this is an authorisation control. Narrowing visibility reduces what is likely to be attempted, not what is permitted. Server-side policy remains the thing that actually stops a call.

Frequently asked questions

What is MCP tool overload?

The point at which the number of tool definitions presented to a model degrades both economics and behaviour: definitions consume a large share of the context window on every turn, and selection accuracy falls as the model chooses between many similar options.

How many tokens does an MCP tool definition consume?

Typically 80 to 400 tokens, depending on parameter count and description length, with 150 a reasonable planning average. Two hundred tools of average complexity is roughly 30,000 input tokens on every turn, not once per conversation.

How many MCP tools should an agent see at once?

Twenty to forty for most tasks. That range is an empirical operating figure from instrumented estates rather than a protocol limit, and it is where selection errors attributable to choice overload largely disappear.

Does limiting visible tools reduce agent capability?

No, provided the limiting is per task. The agent retains access to the whole estate; it is simply not shown everything at once. What changes is how much the model must read and disambiguate on each turn.

What is progressive tool discovery?

A two-phase presentation: first a compact catalogue of tool names and one-line purposes with no schemas, then full schemas only for the tools the model asks to expand. It typically reduces definition tokens by sixty to eighty percent at the cost of one extra round trip.

Is semantic tool retrieval better than static tiering?

Not universally. Static tiering is cheaper, fully deterministic and easy to audit; semantic retrieval covers novel requests no profile anticipated but introduces non-determinism. Most mature estates use tiering as the primary mechanism with retrieval as a fallback.

What is an MCP routebook and how does it help?

A precomputed mapping from the capabilities a recurring workflow needs to the specific servers and tools that satisfy them, including fallbacks. It gives the agent a small verified tool set and the lowest per-turn token cost of any technique, at the cost of regeneration when the estate changes.

How do I know if tool overload is affecting my agents?

The diagnostic symptom is that connecting an unrelated server changes behaviour in an existing workflow. Others include arguments borrowed from a sibling tool’s schema, rising use of generic escape-hatch tools, and longer plans with more retries.

How should tool descriptions be written to reduce errors?

Lead with when to choose the tool rather than how it works, explicitly name the sibling tool it is confused with, state reversibility, always put units in parameter descriptions, and delete anything the schema already says. Aim for roughly twenty-five tokens on common tools.

How do I measure tool selection accuracy?

Label 100 to 200 real production requests with the correct tool, measure top-1 and top-3 accuracy plus argument validity against the current tool set, then change exactly one variable and re-measure. Track how often a generic escape-hatch tool is chosen — it is the most sensitive overload indicator.

Does reducing tool visibility improve security?

Yes, as a reduction in reachable surface: an injected instruction cannot invoke a tool the model was never shown. It is not a substitute for authorisation, because visibility is a client-side property and only server-side policy actually prevents a call.

Should duplicate tools across servers be consolidated?

Yes, and it should be done before any other optimisation. Duplicates multiply token cost and create the near-identical choices that cause most wrong selections. Consolidation requires reconciling tool contracts, which is why schema normalisation and context optimisation are the same project.

Glossary

Tool overload
Degradation of cost and selection quality caused by presenting too many tool definitions to a model at once.
Definition overhead
Input tokens consumed by tool definitions on every turn, independent of the conversation’s actual content.
Progressive tool discovery
Two-phase presentation of a compact catalogue followed by full schemas for selected tools only.
Static tiering
Named task profiles each carrying an explicit, reviewed set of visible tools.
Semantic tool retrieval
Embedding-based selection of the top-k most relevant tools for the current request.
Routebook
A precomputed capability-to-tool mapping for a recurring workflow, including fallbacks and conditions.
Escape-hatch tool
A generic capability such as raw HTTP, SQL or shell that a model reaches for when planning breaks down.
Near-duplicate confusion
Wrong selection between tools offering the same capability with different semantics across servers.
Discrimination-style description
A tool description written to help a model choose correctly rather than to document behaviour.
Top-1 accuracy
Share of labelled requests for which the model’s first tool choice is the correct one.

Standards and entities referenced

Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.

  1. 01 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  2. 02 · MCP projectModel Context Protocol — official documentation ↗Primary source for protocol structure, transports and the shape of a tools/list response.
  3. 03 · JSON SchemaJSON Schema Specification ↗How tool parameter contracts are expressed, and what a validator can enforce.
  4. 04 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
  5. 05 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  6. 06 · Simon WillisonPrompt injection — ongoing series ↗The most consistently updated practitioner record of the attack class.
  7. 07 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.
  8. 08 · SemVerSemantic Versioning 2.0.0 ↗The versioning contract tool schemas should honour but frequently do not.

Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). MCP Tool Overload: What Happens When Agents See Too Many Tools. Real Biz Digital. https://realbizdigital.net/insights/mcp-tool-overload/

Try the mechanics on a live server

To watch a real tools/list response, and see how much surface one server exposes, before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

Barzel Central Gateway is this layer, sold as a running product

Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so an evaluation costs an afternoon rather than a purchase order.

PlanPriceIncludedRight for
CommunityFree1,000 tool calls/mo · full policy engine, registry, routing, auditEvaluating the estate, or one team proving the path works
Starter$10/mo10,000 calls/mo · everything in CommunityOne or two production agents against a handful of servers
Team$79/mo100,000 calls/mo · routebooks, simulation, change impactA platform team governing an estate of 5–20 servers
Business$149/mo250,000 calls/mo · estate-wide evidence exportMulti-team governance with SIEM obligations

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

The five Barzel servers, and which problem each one is sold for

One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.

ServerSold forEntry priceWhere it sits
Barzel Central GatewayKnowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidenceFree, then $10–$149/moControl plane — decides what may be reached, and by whom
BarzelVaultStopping a specific dangerous action before it executes, with proof afterwards$199–$3,999/moDecision point — evaluates the individual call before execution
BarzelOpsRunning real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approvalFree, then $19–$199/moExecution layer — does the work the policy allowed
Barzel FinOps AtlasAttributing AI spend to agents, tools and outcomes, then forecasting and capping itFree, then $29–$799/moEconomics layer — what the estate costs per outcome
Barzel Scripture IntelligenceA free, credential-free public MCP server to test clients and inspect real protocol trafficFree, unmetered, no signupReference implementation — safe place to learn the protocol

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.