5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

Capacity · forecasting

MCP Capacity Planning: Forecasting Usage, Growth and Cost

Human traffic has a shape you can plan around. Agent traffic has a shape too — it is just a vertical line, arriving without warning, from one identity that has decided to retry.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202616 min3,043 words

The short answer

MCP capacity planning forecasts call volume, concurrency and cost for an agent estate. It differs from conventional capacity planning because agent load is bursty rather than diurnal: a single agent in a retry loop can generate more calls in ninety seconds than an entire day of intended traffic. Plan for per-identity ceilings and burst absorption first, then average throughput — and expect upstream API quotas, not your own compute, to be the binding constraint. Four demand drivers, a spreadsheet-grade volume model, and the concurrency arithmetic that tells you what to provision.

Key takeaways

  1. 01Agent load is bursty, not diurnal. Provisioning for the average and adding headroom is the wrong shape entirely.
  2. 02The binding constraint is almost never your compute. It is an upstream API quota, a database connection pool, an approval queue, or a per-tenant vendor limit.
  3. 03Calls per outcome is the number that actually drives your forecast — and it is a property of prompt and model quality, not of demand.
  4. 04Model or prompt changes move capacity more than user growth does. A model that retries twice as often doubles load with no new users.
  5. 05Per-identity ceilings are a capacity control, not just a security one. Without them, one agent’s forecast error is everyone’s outage.
  6. 06Forecast the approval queue alongside the compute. A human bottleneck fails earlier and less gracefully than a server.

Agent load has a different shape

Capacity planning for human-facing systems rests on a reliable assumption: demand follows people. There is a morning ramp, a lunch dip, an evening decline, a weekly rhythm and a seasonal curve. You measure the peak, provision above it, and the shape holds well enough to plan a year ahead.

Agent demand does not have that shape. It has three properties that break the usual approach.

It is bursty at a scale human traffic never reaches. A batch job kicks off and issues ten thousand calls in a minute. An agent hits an ambiguous response and retries four hundred times in thirty seconds. Neither is preceded by a ramp.

It is correlated across identities. When a shared upstream slows down, every agent retries at once. Human users give up; agents do not, and they synchronise on the failure.

It changes without a demand change. A prompt revision, a model upgrade or a schema tweak can double or halve call volume overnight with identical user numbers. This is the property that makes conventional trend forecasting close to useless: your history is a record of a system that no longer exists.

Key facts

  • ▸Provision for burst absorption and per-identity ceilings before average throughput.
  • ▸Agent retries synchronise on upstream failure, so load and failure arrive together.
  • ▸A model or prompt change invalidates your volume history.
10–100×Typical burst ratio for agent traffic against its own mean
<90sTime for one looping agent to exceed a day’s intended volume
0 usersDemand change needed for a prompt revision to double call volume
1stRank of upstream API quota among limits that bind, in our experience

Four demand drivers

Forecasting from last quarter’s call count extrapolates an artefact. Forecast from the things that actually generate calls, and the model survives a model upgrade.

Driver 01

Business event volume

How many of the triggering things happen — tickets raised, invoices received, reconciliations due, documents uploaded. This is the only driver your business can forecast independently, and it is the anchor for everything else.

Driver 02

Amplification factor

Tool calls generated per business event, across the whole workflow. A ticket that becomes eleven calls has an amplification of 11. Measure it per workflow; assuming one is the single most common forecasting error.

Amplification is where multi-step workflows hide. A five-step workflow with two lookups per step is ten calls before any retry.

Driver 03

Retry and abandonment rate

How often a call is repeated and how often a workflow is abandoned partway. This is the volatile term, and it is a property of prompt quality, model choice and upstream reliability rather than of demand.

Track it separately from amplification. Conflating them means you cannot tell a demand rise from a quality regression, and they need entirely different responses.

Driver 04

Discovery overhead

Every turn carries the tool list. At scale this is a material share of token spend that scales with conversation turns rather than with tool calls. Filtering discovery by identity reduces it directly — see tool discovery.

A volume model you can build in a spreadsheet

Six lines. It is deliberately crude, because a crude model you update monthly beats a sophisticated one you built once.

Worked example · Forecasting a support-refund estate one quarter out

Numbers are illustrative but the structure is the one to copy. Each row is a driver you can measure independently, which is what makes the forecast arguable rather than asserted.

business events2,400 refund-eligible tickets/month (from the ticketing system, not from call logs)
amplification9.2 calls per ticket (measured across the declared workflow, 6 steps)
retry factor×1.35 (current measured retry and abandonment overhead)
base volume2,400 × 9.2 × 1.35 ≈ 29,800 calls/month
growth assumption+18% quarter on quarter in ticket volume → ≈ 35,200 calls/month
quality assumptionretry factor improves to ×1.20 after prompt work → ≈ 31,300 calls/month
planning figure35,200/month — plan on the pessimistic quality case
burst provisionpeak 40 calls/sec observed; ceiling set at 60/sec per identity, estate cap 200/sec

Two things to notice. The quality assumption moves the forecast by 11% — comparable to a full quarter of growth — which is why retry rate belongs in the model as its own term. And the planning figure uses the pessimistic quality case, because prompt improvements are a hope and ticket growth is a trend.

Update this monthly against actuals and record the variance. After three months the variance itself tells you which driver you are worst at estimating, which is more useful than the forecast.

Concurrency, not just volume

Monthly volume tells you about cost. Concurrency tells you whether the system falls over. They are different questions and the second one is answered by arithmetic rather than by extrapolation.

Little’s Law: concurrency = arrival rate × average time in system. If calls arrive at 40 per second and each occupies a connection for 250ms, you need 10 concurrent slots at steady state. Straightforward — and it is the steady-state answer, which is not the one that matters.

Three adjustments turn it into a provisioning number.

  • 01Use the burst arrival rate, not the mean. If the mean is 4/sec and bursts hit 40/sec, provision against 40. Sizing to the mean guarantees the burst is an incident.
  • 02Include queue and approval time in ‘time in system’. A call held four hours for human approval occupies a workflow slot for four hours. Approval-heavy workflows have concurrency requirements that look absurd until you account for the wait, and they need a different resource model — durable state, not held connections.
  • 03Add the retry tail. Retries arrive while originals are still in flight, so effective arrival exceeds nominal during exactly the periods you care about.
  • 04Check the pool behind the pool. Your gateway may handle 200 concurrent calls; the database connection pool behind the third server may handle 20. The smallest pool in the chain is your real concurrency limit, and nobody has documented where it is.

Four limits that bind before compute

Teams plan compute because compute is the thing they buy. In practice compute is rarely first to break.

LimitWhy it binds firstHow it failsPlan for it by
Upstream API quotaVendors meter per tenant per minute, and agents burst429s across every agent at once, mid-workflowModelling calls per vendor per minute against the published quota, per environment
Approval queue throughputHumans process ~15 meaningful reviews per dayRequests age out, then get bulk-approvedForecasting approval volume as its own capacity line
Connection pools in backing systemsSized years ago for human-driven trafficTimeouts that look like network faultsFinding the smallest pool in each chain and documenting it
Per-tenant rate limits inside a shared serverOne tenant’s burst starves the restCross-tenant latency with no obvious causePer-identity ceilings enforced at the gateway, not per server

The pattern across all four: the binding constraint is usually outside the component you control and was sized for a different kind of traffic. Capacity planning for an MCP estate is therefore mostly an inventory exercise — find the limits, write them down, and model against the smallest one. The inventory is where this starts.

Built on this thinking

Ceilings are the capacity control, and they are per identity

Barzel Central Gateway enforces per-identity call and spend ceilings so one agent’s forecast error is not an estate-wide outage, and its call record supplies the amplification and retry figures a forecast actually needs. Plans from a free Community tier to 250,000 calls a month.

Cost: per call is the wrong denominator

The natural cost metric is per call, because that is how you are billed. It is also the metric that will mislead you, because the number of calls is not fixed — it is an output of quality.

A cheaper model that retries twice as often costs more per outcome while costing less per call. Every dashboard shows an improvement and the invoice goes up. Reporting per-call cost alone makes this invisible and, worse, makes it look like a win.

Forecast cost per outcome: total spend divided by completed business outcomes. Track per-call cost alongside it as a diagnostic, never as the headline. Cost per outcome works the arithmetic through in full, and spend ceilings covers the enforcement side.

Two cost lines are routinely missing from MCP forecasts. Discovery overhead — the tool list carried on every turn — scales with conversation turns, not tool calls, so it grows even when call volume does not. And refused calls still cost: policy evaluation, logging and the model tokens spent proposing them. A high refusal rate is a cost line as well as a signal.

Where these forecasts stop being reliable

Every number here depends on the workflow staying the same. A prompt revision, a model change or a new tool can move amplification and retry rate substantially, and when they move the forecast is not slightly wrong — its basis has changed. Re-measure amplification after any model or prompt change rather than trusting the trend through it.

Burst behaviour is also genuinely hard to forecast, because its causes are exogenous: an upstream outage, a bad deploy, a batch job somebody scheduled. The honest posture is not to predict bursts but to bound them. Per-identity ceilings, circuit breakers and a documented failure mode per risk class turn an unforecastable event into a bounded one, and that is the achievable goal.

Finally, none of this addresses whether the work is worth doing. An estate can be perfectly provisioned, comfortably inside every quota, and spending money on workflows that produce nothing. That is a process quality question, and capacity planning will not surface it — a well-provisioned useless workflow looks exactly like a well-provisioned useful one.

Frequently asked questions

What is MCP capacity planning?

Forecasting the call volume, concurrency, latency budget and cost of a Model Context Protocol estate, and provisioning limits and infrastructure accordingly. It differs from conventional capacity planning because agent load is bursty rather than diurnal.

Why is agent traffic harder to forecast than human traffic?

Three reasons. It bursts at 10–100 times its own mean with no ramp. Retries synchronise across identities when a shared upstream slows, so load and failure arrive together. And it changes without a demand change — a prompt or model revision can double call volume with identical user numbers, which invalidates your history.

What drives MCP call volume?

Four drivers: business event volume, which is the only one your business forecasts independently; amplification factor, meaning tool calls per business event across the whole workflow; retry and abandonment rate, which is a property of prompt and model quality; and discovery overhead, which scales with conversation turns rather than calls.

How do you calculate concurrency for an MCP estate?

Little’s Law — concurrency equals arrival rate times average time in system — with three adjustments: use the burst arrival rate rather than the mean, include queue and human-approval time in time-in-system, and add the retry tail. Then check the smallest connection pool in each chain, which is your real limit.

What limit usually binds before compute in an MCP estate?

Upstream API quotas. Vendors meter per tenant per minute and agents burst, so every agent receives 429s simultaneously mid-workflow. After that: approval queue throughput, connection pools in backing systems sized for human traffic, and per-tenant limits inside shared servers.

Should MCP cost be forecast per call or per outcome?

Per outcome. The number of calls is an output of quality, not a fixed input, so a cheaper model that retries twice as often costs more per outcome while costing less per call. Track per-call cost as a diagnostic, never as the headline.

Do refused calls cost anything?

Yes — policy evaluation, logging, and the model tokens spent proposing them. A high refusal rate is a cost line as well as a governance signal, and it is routinely missing from MCP cost forecasts.

How do you plan for bursts you cannot predict?

Bound them rather than forecast them. Per-identity call and spend ceilings, circuit breakers, and a documented failure mode per risk class convert an unforecastable event into a bounded one. Ceilings must be per identity, or one agent’s error consumes everyone’s headroom.

How often should an MCP capacity forecast be updated?

Monthly against actuals, recording the variance. After three months the variance itself tells you which driver you estimate worst, which is more useful than the forecast. Re-measure amplification after any model or prompt change rather than trusting a trend through it.

Does approval workflow affect capacity?

Substantially. A call held four hours for human approval occupies a workflow slot for four hours, so approval-heavy workflows need durable state rather than held connections — and the approval queue itself is a capacity line, bounded at roughly fifteen meaningful reviews per approver per day.

Glossary

MCP capacity planning
Forecasting the call volume, concurrency, latency budget and cost of a Model Context Protocol estate, and provisioning limits and infrastructure accordingly.
Calls per outcome
The average number of tool calls consumed to complete one business outcome, including retries and abandoned attempts.
Burst ratio
Peak calls per second divided by mean calls per second, which for agent estates is routinely two orders of magnitude higher than for human traffic.
Binding constraint
The first limit reached as load grows &mdash; frequently an upstream API quota or an approval queue rather than compute.
Amplification factor
The multiplier between one triggering business event and the tool calls it ultimately generates across a workflow.

Sources and further reading

The queueing arithmetic and SLO framing are standard and cited below. The burst-first argument, the demand-driver model and the claim that upstream quotas usually bind first are our own operating judgement from running five MCP servers.

  1. 01 · WikipediaLittle’s Law and queueing fundamentals ↗Arrival rate, concurrency and latency — the arithmetic behind capacity planning.
  2. 02 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
  3. 03 · FinOps FoundationFinOps Framework ↗Inform, optimise, operate — the phases capacity and cost planning map onto.
  4. 04 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.
  5. 05 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  6. 06 · Kubernetes projectKubernetes — cluster architecture ↗The canonical control-plane / data-plane separation, and the closest well-understood analogue.
  7. 07 · JSON-RPC Working GroupJSON-RPC 2.0 Specification ↗The request/response envelope every MCP tool call travels in.
  8. 08 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.

Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). MCP Capacity Planning: Forecasting Usage, Growth and Cost. Real Biz Digital. https://realbizdigital.net/insights/mcp-capacity-planning/

Try the mechanics on a live server

To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

Barzel Central Gateway is this layer, sold as a running product

Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so an evaluation costs an afternoon rather than a purchase order.

PlanPriceIncludedRight for
CommunityFree1,000 tool calls/mo · full policy engine, registry, auditEvaluating the estate, or a single team proving the path works
Starter$10/mo10,000 calls/moOne or two production agents against a handful of servers
Team$79/mo100,000 calls/moA platform team governing an estate of 5–20 servers
Business$149/mo250,000 calls/moEstate-wide governance with SIEM evidence and multi-team routing

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.