Capacity · forecasting
MCP Capacity Planning: Forecasting Usage, Growth and Cost
Human traffic has a shape you can plan around. Agent traffic has a shape too — it is just a vertical line, arriving without warning, from one identity that has decided to retry.
The short answer
MCP capacity planning forecasts call volume, concurrency and cost for an agent estate. It differs from conventional capacity planning because agent load is bursty rather than diurnal: a single agent in a retry loop can generate more calls in ninety seconds than an entire day of intended traffic. Plan for per-identity ceilings and burst absorption first, then average throughput — and expect upstream API quotas, not your own compute, to be the binding constraint. Four demand drivers, a spreadsheet-grade volume model, and the concurrency arithmetic that tells you what to provision.
Key takeaways
- 01Agent load is bursty, not diurnal. Provisioning for the average and adding headroom is the wrong shape entirely.
- 02The binding constraint is almost never your compute. It is an upstream API quota, a database connection pool, an approval queue, or a per-tenant vendor limit.
- 03Calls per outcome is the number that actually drives your forecast — and it is a property of prompt and model quality, not of demand.
- 04Model or prompt changes move capacity more than user growth does. A model that retries twice as often doubles load with no new users.
- 05Per-identity ceilings are a capacity control, not just a security one. Without them, one agent’s forecast error is everyone’s outage.
- 06Forecast the approval queue alongside the compute. A human bottleneck fails earlier and less gracefully than a server.
Agent load has a different shape
Capacity planning for human-facing systems rests on a reliable assumption: demand follows people. There is a morning ramp, a lunch dip, an evening decline, a weekly rhythm and a seasonal curve. You measure the peak, provision above it, and the shape holds well enough to plan a year ahead.
Agent demand does not have that shape. It has three properties that break the usual approach.
It is bursty at a scale human traffic never reaches. A batch job kicks off and issues ten thousand calls in a minute. An agent hits an ambiguous response and retries four hundred times in thirty seconds. Neither is preceded by a ramp.
It is correlated across identities. When a shared upstream slows down, every agent retries at once. Human users give up; agents do not, and they synchronise on the failure.
It changes without a demand change. A prompt revision, a model upgrade or a schema tweak can double or halve call volume overnight with identical user numbers. This is the property that makes conventional trend forecasting close to useless: your history is a record of a system that no longer exists.
Key facts
- ▸Provision for burst absorption and per-identity ceilings before average throughput.
- ▸Agent retries synchronise on upstream failure, so load and failure arrive together.
- ▸A model or prompt change invalidates your volume history.
Four demand drivers
Forecasting from last quarter’s call count extrapolates an artefact. Forecast from the things that actually generate calls, and the model survives a model upgrade.
Business event volume
How many of the triggering things happen — tickets raised, invoices received, reconciliations due, documents uploaded. This is the only driver your business can forecast independently, and it is the anchor for everything else.
Amplification factor
Tool calls generated per business event, across the whole workflow. A ticket that becomes eleven calls has an amplification of 11. Measure it per workflow; assuming one is the single most common forecasting error.
Amplification is where multi-step workflows hide. A five-step workflow with two lookups per step is ten calls before any retry.
Retry and abandonment rate
How often a call is repeated and how often a workflow is abandoned partway. This is the volatile term, and it is a property of prompt quality, model choice and upstream reliability rather than of demand.
Track it separately from amplification. Conflating them means you cannot tell a demand rise from a quality regression, and they need entirely different responses.
Discovery overhead
Every turn carries the tool list. At scale this is a material share of token spend that scales with conversation turns rather than with tool calls. Filtering discovery by identity reduces it directly — see tool discovery.
A volume model you can build in a spreadsheet
Six lines. It is deliberately crude, because a crude model you update monthly beats a sophisticated one you built once.
Numbers are illustrative but the structure is the one to copy. Each row is a driver you can measure independently, which is what makes the forecast arguable rather than asserted.
Two things to notice. The quality assumption moves the forecast by 11% — comparable to a full quarter of growth — which is why retry rate belongs in the model as its own term. And the planning figure uses the pessimistic quality case, because prompt improvements are a hope and ticket growth is a trend.
Update this monthly against actuals and record the variance. After three months the variance itself tells you which driver you are worst at estimating, which is more useful than the forecast.
Concurrency, not just volume
Monthly volume tells you about cost. Concurrency tells you whether the system falls over. They are different questions and the second one is answered by arithmetic rather than by extrapolation.
Little’s Law: concurrency = arrival rate × average time in system. If calls arrive at 40 per second and each occupies a connection for 250ms, you need 10 concurrent slots at steady state. Straightforward — and it is the steady-state answer, which is not the one that matters.
Three adjustments turn it into a provisioning number.
- 01Use the burst arrival rate, not the mean. If the mean is 4/sec and bursts hit 40/sec, provision against 40. Sizing to the mean guarantees the burst is an incident.
- 02Include queue and approval time in ‘time in system’. A call held four hours for human approval occupies a workflow slot for four hours. Approval-heavy workflows have concurrency requirements that look absurd until you account for the wait, and they need a different resource model — durable state, not held connections.
- 03Add the retry tail. Retries arrive while originals are still in flight, so effective arrival exceeds nominal during exactly the periods you care about.
- 04Check the pool behind the pool. Your gateway may handle 200 concurrent calls; the database connection pool behind the third server may handle 20. The smallest pool in the chain is your real concurrency limit, and nobody has documented where it is.
Four limits that bind before compute
Teams plan compute because compute is the thing they buy. In practice compute is rarely first to break.
| Limit | Why it binds first | How it fails | Plan for it by |
|---|---|---|---|
| Upstream API quota | Vendors meter per tenant per minute, and agents burst | 429s across every agent at once, mid-workflow | Modelling calls per vendor per minute against the published quota, per environment |
| Approval queue throughput | Humans process ~15 meaningful reviews per day | Requests age out, then get bulk-approved | Forecasting approval volume as its own capacity line |
| Connection pools in backing systems | Sized years ago for human-driven traffic | Timeouts that look like network faults | Finding the smallest pool in each chain and documenting it |
| Per-tenant rate limits inside a shared server | One tenant’s burst starves the rest | Cross-tenant latency with no obvious cause | Per-identity ceilings enforced at the gateway, not per server |
The pattern across all four: the binding constraint is usually outside the component you control and was sized for a different kind of traffic. Capacity planning for an MCP estate is therefore mostly an inventory exercise — find the limits, write them down, and model against the smallest one. The inventory is where this starts.
Built on this thinking
Ceilings are the capacity control, and they are per identity
Barzel Central Gateway enforces per-identity call and spend ceilings so one agent’s forecast error is not an estate-wide outage, and its call record supplies the amplification and retry figures a forecast actually needs. Plans from a free Community tier to 250,000 calls a month.
Cost: per call is the wrong denominator
The natural cost metric is per call, because that is how you are billed. It is also the metric that will mislead you, because the number of calls is not fixed — it is an output of quality.
A cheaper model that retries twice as often costs more per outcome while costing less per call. Every dashboard shows an improvement and the invoice goes up. Reporting per-call cost alone makes this invisible and, worse, makes it look like a win.
Forecast cost per outcome: total spend divided by completed business outcomes. Track per-call cost alongside it as a diagnostic, never as the headline. Cost per outcome works the arithmetic through in full, and spend ceilings covers the enforcement side.
Two cost lines are routinely missing from MCP forecasts. Discovery overhead — the tool list carried on every turn — scales with conversation turns, not tool calls, so it grows even when call volume does not. And refused calls still cost: policy evaluation, logging and the model tokens spent proposing them. A high refusal rate is a cost line as well as a signal.
Where these forecasts stop being reliable
Every number here depends on the workflow staying the same. A prompt revision, a model change or a new tool can move amplification and retry rate substantially, and when they move the forecast is not slightly wrong — its basis has changed. Re-measure amplification after any model or prompt change rather than trusting the trend through it.
Burst behaviour is also genuinely hard to forecast, because its causes are exogenous: an upstream outage, a bad deploy, a batch job somebody scheduled. The honest posture is not to predict bursts but to bound them. Per-identity ceilings, circuit breakers and a documented failure mode per risk class turn an unforecastable event into a bounded one, and that is the achievable goal.
Finally, none of this addresses whether the work is worth doing. An estate can be perfectly provisioned, comfortably inside every quota, and spending money on workflows that produce nothing. That is a process quality question, and capacity planning will not surface it — a well-provisioned useless workflow looks exactly like a well-provisioned useful one.
Frequently asked questions
What is MCP capacity planning?
Forecasting the call volume, concurrency, latency budget and cost of a Model Context Protocol estate, and provisioning limits and infrastructure accordingly. It differs from conventional capacity planning because agent load is bursty rather than diurnal.
Why is agent traffic harder to forecast than human traffic?
Three reasons. It bursts at 10–100 times its own mean with no ramp. Retries synchronise across identities when a shared upstream slows, so load and failure arrive together. And it changes without a demand change — a prompt or model revision can double call volume with identical user numbers, which invalidates your history.
What drives MCP call volume?
Four drivers: business event volume, which is the only one your business forecasts independently; amplification factor, meaning tool calls per business event across the whole workflow; retry and abandonment rate, which is a property of prompt and model quality; and discovery overhead, which scales with conversation turns rather than calls.
How do you calculate concurrency for an MCP estate?
Little’s Law — concurrency equals arrival rate times average time in system — with three adjustments: use the burst arrival rate rather than the mean, include queue and human-approval time in time-in-system, and add the retry tail. Then check the smallest connection pool in each chain, which is your real limit.
What limit usually binds before compute in an MCP estate?
Upstream API quotas. Vendors meter per tenant per minute and agents burst, so every agent receives 429s simultaneously mid-workflow. After that: approval queue throughput, connection pools in backing systems sized for human traffic, and per-tenant limits inside shared servers.
Should MCP cost be forecast per call or per outcome?
Per outcome. The number of calls is an output of quality, not a fixed input, so a cheaper model that retries twice as often costs more per outcome while costing less per call. Track per-call cost as a diagnostic, never as the headline.
Do refused calls cost anything?
Yes — policy evaluation, logging, and the model tokens spent proposing them. A high refusal rate is a cost line as well as a governance signal, and it is routinely missing from MCP cost forecasts.
How do you plan for bursts you cannot predict?
Bound them rather than forecast them. Per-identity call and spend ceilings, circuit breakers, and a documented failure mode per risk class convert an unforecastable event into a bounded one. Ceilings must be per identity, or one agent’s error consumes everyone’s headroom.
How often should an MCP capacity forecast be updated?
Monthly against actuals, recording the variance. After three months the variance itself tells you which driver you estimate worst, which is more useful than the forecast. Re-measure amplification after any model or prompt change rather than trusting a trend through it.
Does approval workflow affect capacity?
Substantially. A call held four hours for human approval occupies a workflow slot for four hours, so approval-heavy workflows need durable state rather than held connections — and the approval queue itself is a capacity line, bounded at roughly fifteen meaningful reviews per approver per day.
Glossary
- MCP capacity planning
- Forecasting the call volume, concurrency, latency budget and cost of a Model Context Protocol estate, and provisioning limits and infrastructure accordingly.
- Calls per outcome
- The average number of tool calls consumed to complete one business outcome, including retries and abandoned attempts.
- Burst ratio
- Peak calls per second divided by mean calls per second, which for agent estates is routinely two orders of magnitude higher than for human traffic.
- Binding constraint
- The first limit reached as load grows — frequently an upstream API quota or an approval queue rather than compute.
- Amplification factor
- The multiplier between one triggering business event and the tool calls it ultimately generates across a workflow.
Sources and further reading
The queueing arithmetic and SLO framing are standard and cited below. The burst-first argument, the demand-driver model and the claim that upstream quotas usually bind first are our own operating judgement from running five MCP servers.
- 01 · WikipediaLittle’s Law and queueing fundamentals ↗Arrival rate, concurrency and latency — the arithmetic behind capacity planning.
- 02 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
- 03 · FinOps FoundationFinOps Framework ↗Inform, optimise, operate — the phases capacity and cost planning map onto.
- 04 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.
- 05 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 06 · Kubernetes projectKubernetes — cluster architecture ↗The canonical control-plane / data-plane separation, and the closest well-understood analogue.
- 07 · JSON-RPC Working GroupJSON-RPC 2.0 Specification ↗The request/response envelope every MCP tool call travels in.
- 08 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). MCP Capacity Planning: Forecasting Usage, Growth and Cost. Real Biz Digital. https://realbizdigital.net/insights/mcp-capacity-planning/
Try the mechanics on a live server
To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
Barzel Central Gateway is this layer, sold as a running product
Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so an evaluation costs an afternoon rather than a purchase order.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Community | Free | 1,000 tool calls/mo · full policy engine, registry, audit | Evaluating the estate, or a single team proving the path works |
| Starter | $10/mo | 10,000 calls/mo | One or two production agents against a handful of servers |
| Team | $79/mo | 100,000 calls/mo | A platform team governing an estate of 5–20 servers |
| Business | $149/mo | 250,000 calls/mo | Estate-wide governance with SIEM evidence and multi-team routing |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.