Checklist · 18 practices
MCP Security Best Practices: The Enterprise Checklist
Most published checklists are lists of virtues. This one attaches a failure and a test to every line, because a practice you cannot test is a practice you cannot claim.
The short answer
The eighteen MCP security practices that matter group into five areas: identity, meaning short-lived per-agent tokens and preserved caller authority; surface, meaning fewest tools, split read from write, filtered discovery; policy, meaning parameter bounds and approval for irreversible calls; evidence, meaning one record per attempt including refusals; and supply chain, meaning pinned versions and diffed tool descriptions. Each has a test that expects a refusal. Roughly two weeks of engineering for the first fifteen; the last three need decisions rather than code.
Identity — four practices
Everything else on this page degrades to guesswork without attribution. Do these first; they are also the cheapest.
| Practice | Failure it prevents | Effort | Test |
|---|---|---|---|
| Short-lived per-agent tokens, never a shared static key | Unattributable calls; inability to revoke one caller without breaking all of them | 1–2 days | Call with another agent’s token: refused, and the attempt attributed correctly |
| Validate the token on every request, not at connection | A revoked agent continuing to work for the life of a long-lived session | Hours | Revoke mid-session; the next call fails |
| Preserve the original caller’s authority across the hop | Confused deputy: a low-privilege request executing at service-account privilege | 2–3 days | A low-privilege user’s request is refused for a high-privilege action |
| Per-user OAuth rather than a pooled credential where the backing system supports it | One credential’s compromise exposing every user’s data | 2–4 days | Two users’ calls appear in the backing system’s own audit as two identities |
Practice three is the one teams skip and the one auditors ask about. MCP authentication patterns compares what each option proves.
Surface — four practices
Reducing what exists beats filtering what is reachable. These four are the only practices on the page that make the worst case smaller rather than less likely.
| Practice | Failure it prevents | Effort | Test |
|---|---|---|---|
| Scope the server’s credential to the narrowest workable permission | Excessive agency at the root: one wrong call being disproportionate | Days of engineering, weeks of agreement | Use the credential directly against the backing system for an out-of-scope action: refused there, not by your gateway |
| Expose the fewest tools that work; delete unused ones | Attack surface and selection error, together | 1 day | tools/list matches the intended list exactly |
| Split read tools from write tools into separate servers with separate credentials | A read-only deployment being talked into mutating state | 2–3 days | The read server’s credential cannot write, verified against the backing system |
| Filter tools/list by caller identity | Whole classes of injection outcome; free reconnaissance for an attacker | 1–2 days | An unentitled identity’s tool list omits the tool entirely, and calling it anyway is refused |
Policy — four practices
Where most estates have the largest gap, because tool-level permission feels like it should be enough and is not.
| Practice | Failure it prevents | Effort | Test |
|---|---|---|---|
| Bound every parameter: ceilings, allow-lists, row caps, enumerations | The dangerous call to a permitted tool | 2–4 days | A value one increment out of range is rejected with the attempt recorded |
| Reject out-of-range values rather than clamping them | Losing the signal: a clamp hides the attempt | Hours | The rejected value appears in the record; nothing was silently modified |
| Require human approval for irreversible calls, showing tool, parameters and caller | Rubber-stamp approval, and unreviewed irreversible actions | 3–5 days | The same call below and above threshold: one proceeds, one holds with full context shown |
| Per-identity rate and spend ceilings | Runaway loops, and one agent starving every other agent’s budget | 1–2 days | Burst past the limit: throttled for that identity only |
The genuinely irreversible set is usually under a dozen tools, which is what makes per-call approval affordable. Designing approval thresholds covers where to draw the line.
Evidence — three practices
None of these prevent anything. They are what make the other fifteen defensible, tunable and worth having claimed.
| Practice | Failure it prevents | Effort | Test |
|---|---|---|---|
| One structured record per attempt, including refusals | Having no evidence that any control ever worked | 2–3 days | A refused call from last week is retrievable with identity, tool, parameters, rule and result |
| Redact secrets at write time, not read time | A log nobody is permitted to read, which is a log nobody reads | 1 day | No secret material appears in the store, verified by search |
| Send it where the security team already queries | Evidence that requires an engineer to grep a container | 1–2 days | A security engineer answers an audit question without asking for help |
Supply chain — three practices
For the servers nobody in the building wrote, which is most estates. You cannot harden their internals, so the goal is containment and change detection.
| Practice | Failure it prevents | Effort | Test |
|---|---|---|---|
| Pin versions; no auto-update in production | A third-party upgrade changing what your model can do without review | Hours | The running version matches the registry entry |
| Diff tool descriptions and schemas on every upgrade | Tool poisoning and shadowing via changed descriptions — text your model trusts | 1 day to automate | A deliberately altered description is flagged before it reaches production |
| Isolate each third-party server: own container, own narrow credential, restricted egress | Exfiltration through a server you cannot audit | 2–3 days | The container cannot reach any host except its one backing system |
Securing third-party MCP servers works through the review questions in full, and tool poisoning covers what the diff is looking for.
Built on this thinking
Eleven of the eighteen are cheaper at one enforcement point
Implemented per server, most of these get re-implemented every time you add a server — and the weakest implementation sets your real posture. Barzel Central Gateway is the single identity, entitlement, routing and audit point; BarzelVault holds parameter policy, approvals and tamper-evident evidence behind it.
What to do in which order
The eighteen are not equally urgent, and doing them in list order wastes effort. Four phases.
- 01Week 1 — credential scope and tool reduction. The two practices that shrink the worst case. Everything after this is cheaper.
- 02Week 2 — identity, all four practices. Cheap, fast, and the prerequisite for anything per-agent.
- 03Weeks 3–4 — parameter bounds, approval, ceilings, and the evidence record. The largest single block of engineering and the one that produces visible results.
- 04Ongoing — supply chain. Pinning is immediate; description diffing runs on every upgrade forever. This is the practice that decays if it is not automated in week four.
Four practices that are widely recommended and do not help
Hardening the system prompt
No wording reliably beats an instruction found in retrieved content. Useful as a preference, not as a control.
Scanning tool output for injection strings
Catches the naive version and creates confidence disproportionate to coverage. Keep it; do not rely on it.
Reviewing every action manually at first
Collapses into rubber-stamping within a week and manufactures a record of oversight that did not happen.
A second model judging whether a call is safe
Moves judgement back inside the component that can be persuaded, and makes decisions unreproducible.
Frequently asked questions
What are the most important MCP security best practices?
Eighteen across five groups: identity — short-lived per-agent tokens, per-request validation, preserved caller authority, per-user OAuth; surface — scoped credentials, fewest tools, read/write separation, filtered discovery; policy — parameter bounds, reject rather than clamp, approval for irreversible calls, per-identity ceilings; evidence — one record per attempt including refusals, write-time redaction, a queryable store; supply chain — pinned versions, diffed tool descriptions, isolated containers.
Which MCP security practice should be done first?
Scoping the server’s credential to the narrowest workable permission. It is the only practice that shrinks the worst case rather than reducing the probability of reaching it, and every subsequent control is cheaper to get right against a smaller worst case.
How long does implementing MCP security best practices take?
Roughly two weeks of engineering for fifteen of the eighteen, assuming a server you own and a team that has done it once. Credential scoping and tool reduction take longer in calendar time than in engineering time, because both need the system owner to agree.
Why reject out-of-range parameters rather than clamping them?
Because a clamp hides the attempt, and the attempt is the signal worth recording. A rejected value that appears in the log tells you either the policy is wrong or the agent is being steered; a silently clamped value tells you nothing and produces a confusing debugging session weeks later.
Should read and write tools be in the same MCP server?
Preferably not. Splitting them lets each server hold a credential matching its own worst case, so a read-only deployment cannot be talked into mutating state regardless of what reaches the model.
Is scanning tool output for prompt injection worth doing?
Keep it, do not rely on it. It catches the naive version of the attack and creates confidence disproportionate to its coverage. The control that holds is that the resulting call is refused by entitlement and parameter policy regardless of the reasoning behind it.
What supply-chain practices matter for MCP servers?
Three: pin versions with no auto-update in production, diff tool descriptions and schemas on every upgrade because those descriptions are text your model trusts, and isolate each third-party server in its own container with its own narrow credential and restricted egress.
How do you prove an MCP security practice is actually working?
Attach a test to each one that expects a refusal, and assert the refusal reaches the record. A practice with no observable failure mode cannot be claimed — and if a test cannot be written for a control, that is itself the finding.
Sources and further reading
Where a practice restates an existing standard, that standard is cited. Where it is our own operating judgement — the ordering, the effort estimates and several of the tests — the text says so. The effort figures assume a server you own and a team that has done this once.
- 01 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 02 · MCP projectModel Context Protocol — official documentation ↗Primary source for protocol structure, transports and the shape of a tools/list response.
- 03 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
- 04 · OWASPOWASP Application Security Verification Standard ↗Input-validation, authorization and logging requirements restated here in MCP terms.
- 05 · IETFRFC 8693 — OAuth 2.0 Token Exchange ↗The mechanism for preserving caller authority across a gateway hop.
- 06 · OpenSSFSLSA — Supply-chain Levels for Software Artifacts ↗Provenance and reproducibility framing for third-party server intake.
- 07 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
- 08 · Simon WillisonPrompt injection — ongoing series ↗The most consistently updated practitioner record of the attack class.
Last reviewed 24 August 2026. External links open in a new tab; we do not control their content.
Try the mechanics on a live server
To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
BarzelVault is the pre-execution decision point, sold as a running product
Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic policy outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.