5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

AI Infrastructure · Architecture

Building a Capability Registry

You cannot govern a tool surface you cannot enumerate. Most organisations running agents in production cannot produce the list.

By Mark Alex, Founder Published 19 Aug 2026 9 min read

Key takeaways

  • Filter at discovery, not at call time. A tool absent from the advertised list is unreachable by any amount of persuasion.
  • Ownership is a required field. An unowned tool is an unpatched tool, and it will be found during an incident rather than a review.
  • Register risk class alongside the schema. Money movement, deletion, PII access and external egress are properties of the tool, not judgements made per call.
  • Version everything. Tool descriptions are text your model trusts; a changed description is a changed capability and should require review.
  • The registry is also your cost model. Per-tool call cost is what makes agent spend attributable at all.

The short answer

A capability registry is the authoritative inventory of every tool, server and action your agents can reach — with owner, environment, region, risk class, cost, version and per-identity visibility. It is what makes discovery-time scoping possible: an agent is shown only the tools its identity is entitled to, so a tool it never saw cannot be talked into existence. Without one, “least privilege” is an aspiration nobody can evidence.

Why an inventory is the prerequisite

Every governance control assumes you know what exists. Scope an agent to the tools it needs — which tools exist? Deny bulk export — which tools can export? Review third-party servers — which ones are running? All of it collapses without an inventory, and an inventory maintained by hand in a spreadsheet is an inventory that was accurate once.

  • Agents are configured by developers in local config files, so the true tool surface grows without any central event.
  • Tool lists are advertised dynamically, so what a server offers today is not necessarily what it offered at review.
  • The same logical capability often exists three times through different servers, with different owners and different bounds.
  • Nobody notices an unowned tool until it fails or is abused.

The uncomfortable exercise

Ask for a list of every tool your agents can currently call, with an owner for each. The gap between the answer and reality is your actual attack surface, and it is usually the first genuinely useful output of a governance programme.

The twelve fields

Enough to enforce policy, attribute cost and answer an auditor. Fewer than this and one of those three fails.

FieldPurpose
Tool ID and serverStable identity, independent of display name
Version / digestPins the reviewed definition; changes are events
OwnerA named human or team accountable for it
Description hashDetects silent changes to text the model trusts
Input schemaEnables parameter validation and bounds
Risk classMoney, deletion, PII, external egress, legal
ReversibilityWhether the effect can be undone, and how
EnvironmentDevelopment, staging, production
Data regionWhere the data it touches resides
Cost per callVendor or internal cost, for attribution
EntitlementsWhich identities may discover and call it
Review recordWho approved it, when, and against what

Scroll the table horizontally on narrow screens.

Two of these do disproportionate work. Description hash is how you detect a rug pull — see securing third-party MCP servers. Cost per call is what turns agent spending from a single opaque invoice line into something attributable, which is the subject of AI agent cost tracking.

Discovery-time scoping

The registry’s highest-value use is deciding what an agent is shown, not what it is allowed to call. The difference is larger than it sounds.

Call-time checks only
  • ›Agent sees every tool in the catalogue
  • ›It plans using capabilities it cannot use
  • ›Failures surface as confusing denials mid-task
  • ›A persuaded model will keep trying the forbidden tool
Discovery-time filtering
  • ›Agent sees only its entitled subset
  • ›Plans are formed within the real boundary
  • ›Fewer denials, and each one is meaningful
  • ›An unseen tool cannot be requested at all

Keep call-time enforcement as well — discovery filtering is not a substitute, because a caller can attempt an unadvertised name directly. The point is that the two together produce both a smaller attack surface and better agent behaviour, which is a rare combination.

What the registry decides, per identity

identity: agent.support.triage

advertised:
  tickets.read          allow
  tickets.comment       allow  (bounded: own tickets)
  customer.read         allow  (bounded: 1 record/call)

not advertised:
  customer.export       — risk: PII, bulk
  billing.refund        — risk: money movement
  customer.delete       — risk: irreversible

Detecting drift

A registry that is only written at onboarding decays. Three diffs, run continuously, keep it honest.

  1. 01Advertised versus registeredPoll what each server actually offers and compare to the registry. New tools, removed tools and renamed tools all appear here first.
  2. 02Description hash changesA changed tool description is a changed instruction to your model. Treat it like a dependency upgrade: alert, review the diff, then accept.
  3. 03Called versus entitledTools an agent is entitled to but never calls should lose the entitlement. This is how scope stays tight instead of accumulating permissions forever.

The one to alert on loudly

A server that begins advertising tools it did not advertise at review — particularly write-capable ones — is the runtime rug-pull pattern. It is invisible unless something is diffing the advertised list, and nobody discovers it by reading logs.

Building it without a project

Step 01

Discover, do not design

Point a gateway at your existing servers in observe-only mode and let it enumerate. A registry generated from reality beats one designed in a document.

Step 02

Assign owners before anything else

An inventory without owners cannot be acted on. This is a people exercise and it is the step most likely to stall — do it while the list is still short.

Step 03

Classify risk in four buckets

Money, deletion, PII, external egress. Resist finer taxonomies at the start; four buckets you can apply beat twelve you argue about.

Step 04

Derive entitlements from observed use

Two weeks of traffic tells you which agents use which tools. Grant that plus a margin, then remove the rest.

Step 05

Wire the registry into the enforcement path

A registry no policy layer reads is documentation. It has to be the thing that decides what gets advertised.

The product

The registry, the graph and the routing in one place

Barzel Central Gateway registers and inventories MCP servers, virtual servers, tools, owners, environments and regions, builds a capability graph, risk-scores tools and identifies duplicate, missing or unhealthy routes. Free tier at 1,000 calls a month.

Key terms

Capability registry
The authoritative inventory of every tool and action agents can reach, with the metadata needed to govern and attribute it.
Discovery-time scoping
Filtering the advertised tool list per identity, so an agent never learns of capabilities it may not use.
Capability graph
A representation of tools, servers, owners and routes that makes duplication, gaps and unhealthy paths visible.
Description hash
A fingerprint of a tool’s description text, used to detect changes to instructions the model will trust.
Risk class
A property of a tool indicating the kind of harm it can cause — money movement, deletion, PII access or external egress.

Frequently asked questions

What is a capability registry?

The authoritative inventory of every tool, server and action your AI agents can reach, recording owner, version, schema, risk class, environment, region, cost per call and which identities may discover or call it.

Why filter tools at discovery rather than at call time?

A tool that was never advertised to an agent cannot be requested by it, regardless of how the model is persuaded. Discovery filtering also produces better agent behaviour, because plans are formed inside the real boundary instead of hitting denials mid-task.

Is an API catalogue the same thing?

No. An API catalogue documents endpoints for developers. A capability registry drives runtime decisions — what gets advertised to which identity, with what bounds — and carries risk, reversibility and cost fields a catalogue does not.

How do we keep it accurate?

Continuously diff three things: advertised versus registered tools, description hashes against their reviewed values, and tools called versus tools entitled. Each diff catches a different kind of drift.

What is the minimum to start?

Tool ID, server, owner and risk class. Those four make the inventory actionable; add schema, cost and entitlements as the enforcement path comes online.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.

  1. 01 · MCP project Model Context Protocol — specification ↗ Normative protocol behaviour, including capability negotiation and authorization.
  2. 02 · Anthropic / MCP project Model Context Protocol — documentation ↗ Primary source for protocol structure, transports and tool definitions.
  3. 03 · NIST SP 800-207: Zero Trust Architecture ↗ Origin of the policy-enforcement-point and policy-decision-point separation.
  4. 04 · OpenSSF SLSA — Supply-chain Levels for Software Artifacts ↗ Provenance and integrity levels for dependencies you did not write.
  5. 05 · Reference definition Principle of least privilege ↗ The 1975 Saltzer and Schroeder formulation this all descends from.
  6. 06 · CNCF OpenTelemetry ↗ Standard for the traces and spans an agent execution record should emit.

Last reviewed 19 August 2026. External links open in a new tab; we do not control their content.

Reference documentation

Want the specification rather than the argument?

Asset registration, capability graphs and routebooks as the server actually exposes them. A registry is not a gateway — the difference, and why it matters.

Free tier · 1,000 calls/mo · on the MCPize marketplace

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.