Operations · instrumentation
MCP Observability: Monitoring an Enterprise MCP Estate
Uptime tells you the servers answered. It says nothing about whether agents chose the right tools, whether refusals are rising, or whether one retry loop is quietly consuming the month’s budget.
The short answer
MCP observability means instrumenting the tool call itself, not just server health. Emit one span per call with twelve fields — including caller identity, tool, policy decision, route, parameter hash and token cost — and alert on six metrics: refusal rate by tool, wrong-tool selection rate, per-identity call rate, cost per outcome, approval latency and route health. Refusal rate is the strongest leading indicator, because it moves before errors do. Three dashboards get read; the rest get built and forgotten.
What standard monitoring misses
Point an existing observability stack at an MCP estate and you get request rate, latency, error rate and saturation per server. All useful, none of it answering the questions people actually ask about agents.
Those questions are different in kind. Why did the agent call the wrong tool. Who authorised the transfer. Why did the same request behave differently an hour later. Which agent caused this month’s bill. Whether refusals are rising because policy is wrong or because an agent is being steered. Not one of them is answerable from server-level metrics, because the interesting unit is the call and its decision, not the process that served it.
The gap is structural rather than a tooling deficiency. A service either handled a request or it did not. An agent chose an action, from a set that was itself chosen for it, on the basis of content that may have come from outside your control — and any of those three steps can be the thing that went wrong.
Twelve fields on every tool-call span
One span per attempt, including refused ones. Refused attempts are the half that server-side instrumentation cannot see, because a refused call never reaches the server.
| Field | Answers |
|---|---|
| Caller identity | Who. Not a shared service account, or the field is worthless. |
| Agent and run id | Which agent, and which task — so a chain of calls is reconstructable. |
| Tool name and version | What was attempted, and against which schema. |
| Parameter hash plus flagged values | What it was attempted with, without writing secrets into telemetry. Hash the payload; extract the few values policy cares about. |
| Policy decision and rule id | Allow, deny, transform, escalate, simulate — and which rule decided. Without the rule id, tuning is guesswork. |
| Approver, where applicable | Who said yes. The single most requested field in every audit conversation. |
| Selected route and server | Where it ran — a residency answer and the explanation for behavioural differences between identical calls. |
| Result status | Succeeded, failed, refused, timed out. Four outcomes, not two. |
| Latency, split | Policy evaluation, queue or approval wait, and execution. One number hides which layer is slow. |
| Token and call cost | Attributed at the call, or attribution becomes an invoice-reconciliation exercise later. |
| Tools offered count | How many tools were in the list for this call — the input to any tool-overload analysis. |
| Data classification observed | Whether PII or payment data was present. Needed for retention and for residency questions. |
Two of these are commonly missing and both are cheap. Rule id turns policy tuning from archaeology into a query. Tools offered count is the only way to correlate selection quality with list size, which is the analysis that justifies discovery filtering.
Six metrics worth alerting on
Alertable means somebody should do something when it moves. Everything else is a dashboard, and dashboards are for investigation rather than notification.
Refusal rate, by tool and by identity
The strongest leading indicator in the estate, because it moves before errors do. A refusal spike on one tool means policy changed or the agent’s inputs changed; either way something worth knowing happened.
Read the distribution, not the count. Concentrated in one identity within an hour: almost always a retry loop or a steered agent. Spread evenly across many identities: almost always a mis-specified rule.
Wrong-tool selection rate
Approximated by: calls that succeeded, followed within the same run by a different tool achieving the same end. Crude, and still the best available proxy for whether your tool surface is comprehensible.
Correlate against tools-offered count. If it rises with list size, you have tool overload rather than a model problem.
Calls per identity per minute
The runaway-loop detector. The most common MCP incident is not an attack; it is an agent retrying the same call ten thousand times, and per-server rate limits cannot see it as one actor.
Cost per outcome
Spend divided by completed business outcomes, not by calls. Per-call cost falling while cost per outcome rises is the signature of a cheaper model doing more retries — worked example.
Approval latency, p50 and p95
Where human-in-the-loop workflows actually fail. A p95 of six hours means the approval step is a queue, and queues get bypassed by whoever is under deadline pressure.
Route health and eligibility
Error rate and latency per route, plus how often no eligible route existed. The second number is the one nobody instruments and the one that explains mysterious refusals.
Three dashboards that get read
Build these three and stop. Every additional dashboard we have watched teams build has gone unopened within a month, because it answered a question nobody was asking that week.
Estate posture, weekly
Tools by risk class, entitlements granted and revoked, unowned servers, entries past review date. Read by whoever owns governance, once a week, for ten minutes.
Refusals, live
Refusals grouped by rule and by identity, newest first. The one screen a security engineer opens during an incident, and the one that makes policy tuning tractable.
Cost by agent and outcome
Spend attributed to agents and to business outcomes, with per-call cost alongside cost per outcome. Read by whoever gets asked about the bill.
Deliberately absent: a real-time global call-volume dashboard. It looks impressive on a wall and has never once caused a decision.
Built on this thinking
The decision point is the only place refusals are visible
Barzel Central Gateway emits the call-level record — identity, tool, decision, rule, route, result, cost — including refused attempts, in SIEM-ready form. BarzelVault adds hash-chained audit integrity and anomaly monitoring on top of the same stream.
Four instrumentation pitfalls
Logging only successes
Refusals are the more informative half and the half servers cannot see. Instrument at the decision point, not in the server.
Full parameters in telemetry
Secrets and PII end up in the observability store, which usually has weaker access controls than the systems the data came from. Hash the payload; extract only what policy needs.
Redacting at read time
A log that must be sanitised before anyone may read it is a log nobody reads. Redact at write time.
No agent or run id
Without a run id, a chain of twelve calls is twelve unrelated events, and the interesting failure is always in the chain rather than in any single call.
Retention, and the two-tier answer
Two conflicting requirements: audit evidence should be kept for years, and telemetry containing personal data should not. Resolve it by separating the streams rather than compromising on one number.
Tier one is the evidence record: identity, tool, decision, approver, route, result, timestamp, parameter hash. Small, no payloads, tamper-evident, kept for the period your compliance obligations require — frequently seven years. Tier two is diagnostic telemetry: payload extracts, full traces, retrieved-context snapshots. Large, potentially sensitive, kept for days or weeks.
The parameter hash is the join between them. It lets a seven-year-old evidence record be matched to a thirty-day diagnostic trace while it still exists, without the evidence record itself ever holding the payload. How to audit AI agent actions covers the evidence tier in detail.
Frequently asked questions
What is MCP observability?
Instrumentation of the tool call itself — spans, metrics and events at call level — rather than only server health. It exists to answer why an agent did what it did: which tool it chose, from what list, under which policy decision, on which route, at what cost.
What should an MCP tool-call span record?
Twelve fields: caller identity, agent and run id, tool name and version, parameter hash plus flagged values, policy decision and rule id, approver where applicable, selected route and server, result status, latency split by policy and queue and execution, token and call cost, tools offered count, and data classification observed.
Which MCP metric is the best leading indicator?
Refusal rate by tool and by identity. It moves before error rate does, because a refusal means either policy changed or the agent’s inputs changed. Read the distribution rather than the count: concentrated in one identity within an hour is a retry loop or a steered agent; spread evenly across identities means the rule is mis-specified.
Why is logging only successful calls insufficient?
Because a refused call never reaches the server that would have logged it, so server-side instrumentation cannot see the half of the traffic that tells you most. Instrument at the decision point instead.
Should full tool-call parameters go into telemetry?
No. Secrets and personal data end up in an observability store that usually has weaker access controls than the systems the data came from. Hash the payload, extract only the few values policy cares about, and redact at write time rather than at read time.
How long should MCP records be retained?
Use two tiers. The evidence record — identity, tool, decision, approver, route, result, timestamp, parameter hash — is small and tamper-evident and is kept for the period compliance requires, often seven years. Diagnostic telemetry with payload extracts and traces is kept for days or weeks. The parameter hash joins them.
How do you detect a runaway AI agent?
Calls per identity per minute, alerted. Per-server rate limits cannot see one actor across several servers, and the most common MCP incident is not an attack but an agent retrying the same call thousands of times.
How many MCP dashboards do you actually need?
Three: estate posture read weekly, live refusals grouped by rule and identity, and cost by agent and outcome. A real-time global call-volume dashboard looks impressive and has never caused a decision.
Sources and further reading
Span and metric design follows OpenTelemetry conventions where they exist for GenAI workloads. The twelve span fields, the six alertable metrics and the claim that refusal rate leads error rate are our own operating judgement from running five servers.
- 01 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.
- 02 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
- 03 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 04 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
- 05 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
- 06 · EU AI Act (unofficial consolidated text)EU AI Act — full text ↗Obligations around logging, human oversight and traceability for higher-risk systems.
- 07 · ISOISO/IEC 42001 — AI management systems ↗The management-system standard auditors increasingly map AI governance evidence against.
Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.
Try the mechanics on a live server
To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
Barzel Central Gateway is this layer, sold as a running product
Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so the evaluation costs an afternoon rather than a purchase order.
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.