5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

Risk · classification

MCP Tool Risk Scoring: Classifying Dangerous Capabilities

Two questions decide almost everything: can the action be taken back, and how far does it reach. Numeric scores out of a hundred feel rigorous and produce arguments about whether something is a 62 or a 68.

By Mark Alex, FounderPublished 24 Aug 2026Updated 2 Sep 202612 min

The short answer

MCP tool risk scoring classifies each tool by two properties: reversibility — can the action be undone, and by whom — and reach, meaning how many records, systems or people it can affect in one call. Those two produce five classes, from read-only through irreversible external effect, and each class carries fixed entitlement, approval and logging requirements. Parameter values change the class, so scoring at tool level alone is incomplete. Five classes, two questions, and the reason numeric scores out of a hundred make things worse.

Why numeric scores fail here

The instinct is a score: rate each tool 1 to 100 on likelihood and impact, multiply, sort. It fails for three reasons that show up within a month.

It produces arguments rather than decisions. Nobody can defend 62 over 68, so scoring meetings become negotiations, and the outcome tracks whoever cared most rather than what the tool does. It also implies false comparability — a 40 and a 45 are treated as adjacent when one deletes records and the other reads them.

And it does not connect to any action. A score of 71 does not tell you whether to require approval. Classes do: each class carries fixed entitlement, approval and logging requirements, so classification is the decision rather than an input to a later one.

So: small fixed set of classes, two questions to assign them, accept the loss of nuance. The nuance was never real.

The two questions

Ask them in this order, about a single call, assuming the parameters are as bad as the schema permits.

Can it be taken back — and by whom?

A database write reversible by the owning team is different from an email to a customer, which is not reversible by anyone. The by whom clause matters: reversible-in-principle-by-a-DBA-in-four-hours is not reversible for practical purposes during an incident.

How far does one call reach?

One record or a hundred thousand. One system or every system downstream of it. Internal only, or visible to a customer, a regulator or the public. Reach is what turns a mistake into an incident.

Two supplementary questions occasionally change a class and are worth asking on the tools that sit near a boundary: does it move money, and does it change who has access. Both tend to promote a tool by one class regardless of the first two answers.

Five classes

Each class states its own entitlement, approval and logging requirement. That is the point — the class is not a label, it is a set of consequences.

ClassWhat it isEntitlementApproval
R0 — Read, non-sensitiveReads public or internal non-sensitive data. Reversible trivially: nothing changed.Default-open to entitled agentsNone
R1 — Read, sensitiveReads PII, financial, health or otherwise regulated data. Nothing changes, but disclosure is the risk.Named agents onlyNone, but retention and redaction rules apply
W1 — Bounded writeChanges state, reversible by the owning team, reach limited to a small number of records.Named agents onlyAbove a parameter threshold
W2 — High-reach or hard-to-reverse writeBulk update, schema change, access change, anything reversible only with effort or reaching many records at once.Named agents, restricted environmentAlways
X — Irreversible or externally visiblePayments, emails and messages to third parties, deletions without recovery, public publication, anything a regulator could see.Named agents, production-restricted, per-callAlways, with the parameters shown to the approver

Two calibration notes from applying this. Most estates find 60–75% of tools land in R0 or R1, which is reassuring and also why blanket controls feel so expensive: they are being paid on the majority to protect the minority. And the X class is almost always smaller than teams expect — usually under a dozen tools — which makes per-call approval on X affordable in a way blanket approval never is.

Parameters change the class

This is the part most schemes miss, and it is where the real risk lives. A transfer tool is a different tool at 10 and at 10 million. A query tool returning 5 rows is R1; the same tool returning the whole table is a bulk export and belongs in W2 territory.

So the class is a function of the tool and its arguments. In practice that means two things. Assign the tool a base class from the worst case its schema allows — if a schema permits an unbounded amount, the tool is X until the schema says otherwise. Then declare promotion thresholds: the parameter values at which a call is treated as a higher class for approval purposes.

The useful side effect is that tightening a schema demotes a tool, and demotion is a real reward. A team that bounds its transfer tool at 5,000 gets W1 handling instead of X handling, which is faster for them. Risk classification stops being a tax and becomes something engineers do to make their own lives easier.

This is why a policy engine needs parameter values as an input. A policy keyed only on tool names cannot express any of this.

Scoring a tool you did not write

Most estates contain tools nobody in the building implemented, and their descriptions are the only documentation. Three rules keep this honest.

  • 01Score from the schema, not the description. The description is marketing written by the tool author and read by your model. The schema is what the tool will accept.
  • 02Assume the widest behaviour the schema permits. An unbounded string parameter reaching a shell or a query is a code-execution tool regardless of what the name suggests.
  • 03Score the credential too. A read-only-looking tool holding a read-write credential is one implementation change away from being a write tool, and you will not be told when that change ships.
  • 04Re-score on version change. Diff the tool descriptions and schemas on every upgrade. A changed description is a changed instruction to your model — see securing third-party MCP servers.

Built on this thinking

Classes are only useful if something enforces them per call

Barzel Central Gateway carries risk class and per-tool entitlement in the registry and evaluates parameter thresholds at call time, so a promotion from W1 to X handling happens on the call rather than in a review. BarzelVault holds the approval workflow and the tamper-evident record behind it.

Making it stick

Classification schemes decay unless assignment happens at a moment when somebody is already paying attention. Three mechanisms, in order of effectiveness.

Assign at intake, never later

Class is a required field before a route opens. Retroactive classification campaigns are the ones that stall in week three.

Default to the worse class

Unclassified means X until somebody argues otherwise. This inverts the incentive: the team that wants faster handling does the classification work.

Re-score on schema change, automatically

Diff schemas on upgrade and flag the changed tools for re-scoring. Nobody remembers to do this manually, and the changes that matter are exactly the ones nobody announces.

Four mistakes worth naming

Scoring servers instead of tools

One server routinely contains R0 reads and X writes. Server-level risk is the average of things that should never be averaged.

Risk as free text

Prose risk cannot be queried, sorted or enforced by a policy engine. Fixed classes or nothing.

Ignoring the credential

A tool is capped by what its credential permits. Score both, and the credential usually turns out to be the binding constraint.

Classes with no consequences attached

If R1 and W1 lead to identical handling, you have labels rather than classes, and the labels will stop being maintained.

Frequently asked questions

How do you score the risk of an MCP tool?

With two questions asked about a single call, assuming parameters are as bad as the schema permits: can the action be taken back, and by whom; and how far does one call reach in records, systems and external visibility. Those two answers assign one of five fixed classes, each carrying its own entitlement, approval and logging requirements.

What are the five MCP tool risk classes?

R0 read non-sensitive, R1 read sensitive, W1 bounded write, W2 high-reach or hard-to-reverse write, and X irreversible or externally visible. Most estates find 60–75% of tools in R0 or R1, and the X class is usually under a dozen tools — which is what makes per-call approval on X affordable.

Why not use a numeric risk score out of 100?

Because it produces arguments instead of decisions, implies false comparability between unlike actions, and connects to no action — a score of 71 does not tell you whether to require approval. Fixed classes each carry consequences, so classification is the decision rather than an input to a later one.

Do parameter values change a tool’s risk class?

Yes, and this is where most of the real risk lives. A transfer tool is a different tool at 10 and at 10 million. Assign a base class from the worst case the schema permits, then declare promotion thresholds — parameter values at which a call is handled as a higher class.

How do you score an MCP tool you did not write?

From the schema rather than the description, assuming the widest behaviour the schema permits, and score the credential as well as the tool. An unbounded string parameter reaching a shell or a query is a code-execution tool whatever the name suggests. Re-score on every version change.

Should risk be scored per server or per tool?

Per tool. One server routinely contains harmless reads and irreversible writes, so a server-level score is the average of things that should never be averaged.

What is the fastest way to make classification stick?

Assign the class at intake as a required field before a route opens, and default unclassified tools to the worst class until somebody argues otherwise. That inverts the incentive: the team that wants faster handling does the classification work.

Does tightening a tool schema reduce its risk class?

Yes, and it should be treated as a reward. A transfer tool bounded at 5,000 earns bounded-write handling instead of irreversible-action handling, which is faster for the team that owns it. That is what turns classification from a tax into something engineers do for themselves.

Sources and further reading

Risk classification practice generally follows the frameworks cited below. The two-question method and the five classes are our own, developed because generic likelihood-times-impact scoring produced classifications engineers would not apply consistently.

  1. 01 · NISTNIST AI Risk Management Framework ↗Govern-map-measure-manage; the vocabulary most enterprise AI risk programmes are written against.
  2. 02 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
  3. 03 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  4. 04 · MITREMITRE ATLAS ↗Adversary technique knowledge base for AI systems, useful for naming what a risk score is scoring.
  5. 05 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
  6. 06 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  7. 07 · Center for Internet SecurityCIS Critical Security Controls ↗Control 1 and 2 — inventory of assets and software — restated here for MCP servers and tools.

Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.

Try the mechanics on a live server

To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

Barzel Central Gateway is this layer, sold as a running product

Twenty-five tools covering identity-aware policy, tool routing, risk scoring, approvals, routebooks, workflow simulation and SIEM evidence. Ten policy inputs, six enforcement outcomes, per-user OAuth/OIDC. The Community tier is free, so the evaluation costs an afternoon rather than a purchase order.

CommunityFree1,000 calls/mo
Starter$10/mo10,000 calls/mo
Team$79/mo100,000 calls/mo
Business$149/mo250,000 calls/mo

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.