Data protection · DLP
MCP Data Protection: DLP, Secrets and PII in Agent Estates
Conventional DLP watches data crossing a network boundary. An agent with a read tool and a write tool exfiltrates without crossing one — every call is authorised, every destination is internal, and nothing looks wrong.
The short answer
MCP data protection means controlling sensitive data across six surfaces: tool parameters, tool responses, model context, telemetry, the audit record and the composition of legitimate tools. Conventional DLP misses most of these because it inspects network egress, while an agent moves data between two authorised tools without crossing a boundary. The controls that work are classification at call time, redaction at the enforcement point and entitlement review of tool pairs. Six surfaces, four control types, and the exfiltration pair that no per-tool review has ever flagged.
Key takeaways
- 01The exfiltration path in an agent estate is usually two authorised tools, not one unauthorised transfer.
- 02Classify at call time. Data sensitivity is a property of the payload in front of you, not of the tool that carried it.
- 03Redact at write time, never at read time. A store that must be sanitised before anyone may read it is a store nobody reads.
- 04Tool responses are the most under-protected surface: they enter model context, where they may be summarised into anywhere.
- 05Secrets in tool parameters end up in telemetry with broader read access than the systems they unlock.
- 06Review entitlement as sets. Read-sensitive plus write-external is a risk neither tool has alone.
Why conventional DLP misses agents
Data loss prevention was built around a boundary. Data at rest is classified, data in motion is inspected as it crosses egress, and policy blocks what should not leave. It works because the interesting movement crosses a line somebody drew.
An agent estate breaks the assumption in three ways.
The movement is internal. An agent reads customer records through an authorised tool and writes them into a ticket, a document, a chat message or a support reply. Every hop is inside your perimeter until the last one, which is a legitimate business function.
Every call is authorised. There is no anomaly to detect. The reading tool was approved, the writing tool was approved, the identity was valid. The composition was never reviewed because nothing reviews compositions.
The payload is transformed. The model summarises, paraphrases and reformats. Pattern-matching for a card number fails when the model has written it out in words, split it across a sentence, or described the customer well enough to identify them without quoting a single field.
So the control has to move from the boundary to the call, and from pattern-matching to entitlement design.
Key facts
- ▸Agent exfiltration typically uses two authorised tools and crosses no network boundary.
- ▸Pattern-based DLP degrades against a component that paraphrases.
- ▸The reviewable unit is the entitlement set, not the individual tool.
Six exposure surfaces
Each holds sensitive data at some point, and each needs its own control. Ranked by how often they are left unprotected.
| Surface | What sits there | Usual state | Control |
|---|---|---|---|
| Tool responses | Whatever the backing system returned, in full | Unprotected | Classify and redact at the enforcement point before the model sees it |
| Model context | Everything the agent has read this session | Unprotected and unbounded | Cap response size; fresh context per tenant and per task |
| Diagnostic telemetry | Full payloads, headers, traces | Broad read access, long retention | Hash payloads; redact at write time |
| Tool parameters | Values the model supplies, sometimes secrets | Partially protected | Reject credential-shaped parameters; classify before forwarding |
| The audit record | Identity, decision, policy-relevant values | Usually fine if designed | Policy-relevant values only, never full payloads |
| Composition of tools | No data at rest — a capability | Never reviewed | Entitlement review of pairs, not tools |
The first row is the one to fix first. Tool responses are where the largest volume of sensitive data enters the system, they are almost never inspected, and once a response is in model context it can be summarised into any tool the agent can reach. Inspecting responses at the enforcement point is the highest-leverage data-protection change available in most estates.
Classify at call time, not by tool
The tempting shortcut is to classify the tool: this one returns PII, that one does not. It fails immediately, because the same tool returns different things depending on its arguments. A customer-search tool returns nothing sensitive for a miss and a full record for a hit.
Classify the payload as it passes. Four categories are enough, and more than four stops being applied consistently.
- None
- No regulated or confidential content. The majority of traffic, and it should flow without friction.
- Personal data
- Identifiers, contact details, anything attributable to a person. Triggers residency routing and retention rules.
- Regulated financial or health data
- Payment card data, account details, clinical records. Triggers stricter handling and frequently an external audit obligation.
- Secrets
- Credentials, tokens, keys. Never legitimate in a parameter or a response; presence is a defect, not a classification.
Three implementation notes. Classification must be deterministic and fast — a regex and dictionary pass, not a model call, because a model call makes every decision unreproducible and adds latency to every request. Record the classification in the audit record as a field, because it is what a residency or retention question will be answered from later. And treat the secrets category as an alert rather than a policy branch: a token in a tool response means something upstream is broken.
Four controls, and when each applies
Blocking is the crudest option and rarely the right one. Most sensitive data flowing through an agent estate is flowing legitimately.
Redact
Remove the sensitive values and let the call proceed. Right when the agent needs the record but not the sensitive fields — a support agent needs the order status, not the card number.
The most under-used control and usually the correct one. It preserves the workflow while removing the exposure, and it is the reason a policy engine needs a transform outcome rather than only allow and deny.
Tokenise
Replace the value with a reversible reference. Right when a downstream step must operate on the value without the agent ever holding it — passing a payment method to a processor without the model seeing the number.
More work than redaction and it needs a token vault, but it is the only option when the value must survive the round trip intact.
Route
Send the call to an instance permitted to handle this classification. Right when the issue is jurisdiction or environment rather than the data itself — EU personal data to the EU instance.
Residency is an eligibility filter, never a preference. See routebook eligibility.
Refuse
Stop the call. Right when no transformation makes it acceptable — secrets in a parameter, regulated data heading somewhere with no lawful basis.
Refusal should be the smallest category. A data-protection programme whose only tool is refusal gets routed around, and then protects nothing.
Built on this thinking
Redaction is a policy outcome, applied before the model sees it
BarzelVault classifies payloads at call time and applies guardrail data protection as a transform outcome — redacting sensitive fields from tool responses before they reach model context, rather than blocking the call or detecting the leak afterwards. Classification is recorded in the hash-chained audit line.
The exfiltration pair problem
This is the part conventional review structurally cannot catch, and it is worth stating precisely.
Tool A reads customer records. Approved — the support agent needs it. Tool B sends email. Approved — the support agent needs that too. Neither review was wrong. But an agent holding both can read a thousand customer records and email them anywhere, and every individual action is within policy.
Per-tool review cannot see this because the risk is not in either tool. It is in the set. So the review unit has to change: for every agent, ask what the complete entitlement set makes possible that no individual grant would have permitted.
The pairs worth checking explicitly.
- 01Read-sensitive + write-external. The canonical pair. Any tool reaching PII or financial data, alongside any tool that emails, messages, posts or publishes.
- 02Read-sensitive + create-share-link. Worse than email, because the artefact persists and is often unauthenticated. Document tools that can generate public URLs belong in a category of their own.
- 03Read-any-tenant + write-anywhere. Cross-tenant reach combined with any write is a disclosure engine. See tenant isolation.
- 04Bulk-export + any egress. A tool with an unbounded row count changes the magnitude of every other pair it appears in.
- 05Read-sensitive + code execution. Any tool that runs code, evaluates expressions or renders templates is an egress path, whether or not it looks like one.
Three ways to break a pair, in order of preference: split the agent so no single identity holds both capabilities; bound the reading tool so bulk reads are impossible and the pair’s magnitude collapses; or require approval on the writing tool when the session has recently read sensitive data, which requires recent-behaviour as a policy input.
Splitting the agent is usually cheaper than it sounds and is the only option that removes the capability rather than constraining it.
Secrets are a separate problem
Personal and financial data flow through an estate legitimately and need handling. Secrets do not flow legitimately at all, so their presence anywhere in the agent path is a defect rather than a case to manage.
Four places they appear, and what each means.
| Where | What it means | Action |
|---|---|---|
| In a tool parameter | A server is asking the model to carry a credential | Reject at the gateway; treat the server as a design defect and isolate it |
| In a tool response | A backing system is echoing configuration or an error containing a token | Redact before the model sees it; alert; fix upstream |
| In model context | Already too late for that session | Rotate the credential; investigate how it arrived |
| In telemetry or the audit record | A credential now sits in a store with broad read access | Rotate; fix write-time redaction; audit who could have read it |
The rule that closes most of this: credentials belong in the transport layer, held by the enforcement point, never in the message body. A server that requires a credential as a parameter is asking your model to be a secret manager, and models are poor secret managers — they repeat things.
What data protection at this layer cannot do
Classification is pattern-based and therefore imperfect. It reliably catches structured identifiers — card numbers, national insurance numbers, email addresses — and reliably misses a sentence that identifies someone without containing any of them. Treat detection as a floor that catches the routine cases, and entitlement design as the control that bounds the rest.
Redaction can also break workflows in ways that are hard to predict. An agent that needs a field you removed will either fail or, worse, invent a plausible substitute. Redact fields the workflow demonstrably does not need, verify against real traffic in simulation before enforcing, and expect a tuning period.
And none of this addresses the model provider. If tool responses containing personal data reach a third-party inference API, that is a processing relationship with its own lawful-basis, residency and retention questions, entirely separate from anything your enforcement point controls. It is a contract and architecture question, and it is the one most likely to be discovered late.
Frequently asked questions
Why does conventional DLP miss AI agents?
Because it inspects data crossing a network boundary, and an agent moves data between two authorised tools without crossing one. Every call is permitted, the destination is internal until the last hop, and the model paraphrases the payload — which defeats pattern matching.
Where does sensitive data move in an MCP estate?
Across six surfaces: tool responses, model context, diagnostic telemetry, tool parameters, the audit record, and the composition of legitimate tools. Tool responses are the largest volume and the least protected.
Should data be classified by tool or at call time?
At call time. The same tool returns different things depending on its arguments — a customer search returns nothing sensitive for a miss and a full record for a hit — so tool-level classification is wrong in both directions.
How many data classification categories should you use?
Four: none, personal data, regulated financial or health data, and secrets. More than four stops being applied consistently. Classification must be deterministic and fast, so a regex and dictionary pass rather than a model call.
What is an exfiltration pair?
Two individually-approved tools whose combination allows data to leave — typically one that reads sensitive records and one that writes externally. Per-tool review cannot detect it because the risk is in the set, not in either tool.
How do you break an exfiltration pair?
Three options, in order of preference: split the agent so no single identity holds both capabilities; bound the reading tool so bulk reads are impossible, which collapses the pair’s magnitude; or require approval on the writing tool when the session has recently read sensitive data.
Which control is most under-used in agent data protection?
Redaction. Most sensitive data flows legitimately, and removing the fields the workflow does not need preserves the work while eliminating the exposure — which is why a policy engine needs a transform outcome rather than only allow and deny.
What should happen if a secret appears in a tool parameter?
Reject the call at the gateway and treat the server as a design defect. Credentials belong in the transport layer, held by the enforcement point, never in the message body. A server requiring a credential as a parameter is asking the model to be a secret manager, and models repeat things.
Should audit records contain full tool payloads?
No. Store policy-relevant values plus a hash of the payload. Full payloads move personal data and secrets into an observability store that usually has broader read access and longer retention than the systems the data came from.
What does data protection at the enforcement point not cover?
Sentences that identify someone without containing structured identifiers, which pattern matching misses. And the model provider relationship: tool responses reaching a third-party inference API create a processing relationship with its own lawful-basis and residency questions that no enforcement point controls.
Glossary
- Exposure surface
- A place in an MCP estate where sensitive data comes to rest or passes through, each requiring its own control.
- Exfiltration pair
- Two individually-approved tools whose combination allows sensitive data to leave — typically one that reads sensitive records and one that writes externally.
- Call-time classification
- Determining the sensitivity of a specific payload as it passes the enforcement point, rather than inferring it from which tool was called.
- Tokenisation
- Replacing a sensitive value with a reversible reference so downstream systems can operate on it without holding the original.
- Write-time redaction
- Removing sensitive values before they are stored, so no downstream reader ever has access to them.
Sources and further reading
Data-protection obligations come from the regulations and standards cited below, which are authoritative on requirements. The six-surface model, the exfiltration-pair framing and the control comparison are our own operating judgement.
- 01 · European UnionGDPR — Regulation (EU) 2016/679 ↗Lawful basis, data minimisation and processing records that agent estates inherit.
- 02 · PCI Security Standards CouncilPCI DSS v4.0 ↗Where payment-adjacent agent actions inherit real, externally audited requirements.
- 03 · NISTNIST SP 800-53 Rev. 5 ↗Access control and audit control families that MCP-layer controls have to satisfy.
- 04 · OWASP GenAI Security ProjectOWASP GenAI LLM Top 10 (2026) ↗Consensus risk list; excessive agency and prompt injection are the entries governance exists to bound.
- 05 · OWASPOWASP Application Security Verification Standard ↗Input-validation, authorization and logging requirements restated here in MCP terms.
- 06 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 07 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
- 08 · EU AI Act (unofficial consolidated text)EU AI Act — full text ↗Obligations around logging, human oversight and traceability for higher-risk systems.
Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). MCP Data Protection: DLP, Secrets and PII in Agent Estates. Real Biz Digital. https://realbizdigital.net/insights/mcp-data-protection/
Try the mechanics on a live server
To watch a real tools/list response before you point a client at anything that governs production — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
BarzelVault is the pre-execution decision point, sold as a running product
Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Dev | Free | 10,000 policy decisions/mo · 9 tools, 4 outcomes, hash-chained audit | A first regulated workflow: one agent, one high-consequence system |
| Team | $199/mo | 75,000 decisions/mo · approval workflow, spend and action limits | Several agents acting on money, records or customer-visible systems |
| Business | $799/mo | 750,000 decisions/mo · exact HTTPS execution, credential isolation, emergency controls | Enterprise-wide pre-execution enforcement with audit obligations |
| Enterprise | $3,999/mo | 5,000,000 decisions/mo · everything in Business, scaled | Group-wide rollout across many teams and systems |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.