5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

AI Agent Security · Execution safety

Idempotency and Replay Safety for AI Agents: Why Retries Become Incidents

An agent that times out on a payment does not know whether the payment happened. It will try again. Whether that is a retry or a second payment is decided by design you either did or did not do.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202615 min read3,642 words

The short answer

Idempotency for AI agents means every consequential tool call carries a client-generated key derived from the intent, so that a repeated call with the same key returns the original result rather than causing a second effect. Without it, an agent that cannot tell whether its first attempt succeeded will retry, and the retry becomes a duplicate payment, a duplicate email or a duplicate record. Agents retry more than humans do, more quickly, and with less ability to check first. That combination turns an ordinary distributed-systems problem into a financial one.

Summary for readers and answer engines

Reviewed 25 Aug 2026

  • ▸Agents retry aggressively and cannot easily check whether the first attempt landed, so ambiguity resolves into duplication by default.
  • ▸The key must be derived from intent, not generated per attempt. A fresh key on each retry is not idempotency; it is a second request.
  • ▸Classify every tool into one of four retry-safety categories and give each category a different retry policy. Naturally idempotent reads and non-idempotent writes should not share retry configuration.
  • ▸Where upstream systems offer no idempotency, the governance layer has to provide it: dedupe at the decision point on an intent hash, within a declared window.
  • ▸Reconciliation is the backstop. Some duplicates will occur, and the cost of a duplicate is dominated by how long it goes undetected.

Source: Mark Alex, Real Biz Digital — Idempotency and Replay Safety for AI Agents: Why Retries Become Incidents (https://realbizdigital.net/insights/ai-agent-idempotency/). Reproduce with attribution.

Key takeaways

  1. 01Derive keys from intent: agent, capability, and the semantically significant arguments. Stable across attempts, distinct across intents.
  2. 02Never let the model generate the idempotency key. It will produce a different one when it rephrases the same intent.
  3. 03Set a replay window per capability and state it. Twenty-four hours for payments, minutes for notifications, and a documented answer either way.
  4. 04Treat timeouts as unknown rather than failed. An agent told a call failed will retry; an agent told the status is unknown can be made to reconcile first.
  5. 05Cap retries at the decision point, not only in the client. Client-side caps are bypassed by the next client.
  6. 06Build the duplicate-detection query before granting the capability. It is the only control that catches what design missed.
Part of the clusterAI Agent Security →

Quick answers

One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.

Why do agents cause duplicate actions?
Because a timeout leaves the outcome unknown, and an agent’s default response to an unknown outcome is to try again. Unlike a human, it does not pause to check first.
What is an idempotency key?
A client-supplied identifier that lets a server recognise a repeated request and return the original result instead of performing the action twice.
How should the key be derived?
From the intent: agent identity, capability, and the semantically significant arguments. It must be stable across retries of the same intent and distinct across different intents.
Can the model generate the key?
No. A model rephrasing the same intent will produce a different key, which defeats the mechanism precisely when it is needed.
What if the upstream system has no idempotency support?
Provide it at the governance layer: deduplicate on an intent hash at the decision point within a declared window, and refuse or return the prior result on a match.
What is replay protection?
Refusing a request that repeats outside its legitimate retry window, which distinguishes a genuine retry from a replayed or duplicated instruction.
Is reconciliation still necessary?
Yes. Some duplicates will occur, and the total cost is dominated by detection delay rather than by any individual duplicate.

Why agent retries are different

Retry safety is an old problem. What changed is who is retrying, how fast, and how little they check.

Two properties compound the problem. Agents interpret a timeout as a failure and act accordingly, and they can iterate faster than any monitoring interval — a retry loop can produce hundreds of attempts before a five-minute alert window closes.

Add a third property specific to language models: the same intent can be expressed differently on each attempt. Anything keyed on the exact request text will treat two attempts as two intents, which is exactly wrong.

Conventional client
  • ›Retries are coded deliberately, with backoff
  • ›Developer considered idempotency at design time
  • ›Failure paths are explicit and tested
  • ›A human notices a duplicate quickly
Autonomous agent
  • ›Retries emerge from planning, not from code
  • ›Idempotency was nobody’s decision at that moment
  • ›Failure paths are inferred from an error string
  • ›Duplicates surface in reconciliation, days later

The cost of ambiguity

timeout on a payment call P(succeeded upstream) ≈ unknown, often 0.3-0.7 agent default: retry without idempotency: expected duplicates = attempts × P(first succeeded)

The mechanism does not need to be clever. It needs to exist, because the default behaviour in the presence of ambiguity is duplication.

Key derivation, scope and lifetime

Most idempotency failures are key design failures rather than missing support. Three properties matter: stable across retries, distinct across intents, and scoped correctly.

  • 01Normalise significant arguments before hashing. An amount in minor units and the same amount in major units must produce the same key, or unit ambiguity becomes duplicate payments.
  • 02Scope by business reference where one exists. “Refund for order 10432” is a better intent boundary than any hash of the request.
  • 03Choose the intent window deliberately. A single run for one-off actions; a calendar day for recurring ones; never unbounded, or a legitimate second refund next month is refused.
  • 04Exclude anything that varies per attempt. Timestamps, attempt counters and request text all break stability, which is the property the mechanism depends on.
  • 05Never let the model generate or choose the key. Derive it deterministically at the governance layer from validated inputs.
  • 06Record the key on the decision record. Without it you cannot distinguish a duplicate that was caught from one that was not.

Intent-derived key (illustrative)

key = sha256(
  agent_id            +   # who intends it
  capability          +   # what kind of effect
  business_ref        +   # order id, invoice id, ticket id
  significant_args    +   # amount, currency, recipient - normalised
  intent_window           # e.g. date, or run id for single-run intents
)

# NOT included: attempt number, timestamp, request text,
#               model-generated free text, retry counter

A useful test: ask two engineers to write the key derivation for a refund independently. If they disagree about whether the customer id or the order id is the boundary, your intent model is ambiguous and duplicates are a matter of time.

Four retry-safety categories and their controls

Key facts

  • ▸Category three requires verification, not assumption. Test that the upstream actually honours the key by sending the same key twice and comparing results — support is sometimes partial or silently ignored.
  • ▸Category four is where most real incidents live: notification storms, duplicate ledger entries, repeated ticket creation. These systems rarely offer idempotency and are rarely treated as dangerous.
  • ▸Retry policy belongs on the capability in the registry, so it is a governed property rather than a client configuration that varies by whoever wrote the agent.
CategoryExampleRetry policyAdditional control
Naturally idempotent readFetch a record, search, listRetry freely with backoffRate cap only
Naturally idempotent writeSet a field to a value, upsert by keyRetry with backoffCap attempts; log repeats
Idempotent with a keyPayments, refunds, provisioning with a supported keyRetry only with the same keyRefuse retry if key absent; verify server honours it
Non-idempotentSend email, append to a ledger, create without a natural keyDo not retry automaticallyGovernance-layer dedupe on intent hash, or human confirmation

Classifying tools into these four categories is a half-day exercise with a disproportionate payoff, because it immediately identifies which capabilities must never be retried automatically.

Replay protection: distinguishing a retry from a repeat

Idempotency handles the honest retry. Replay protection handles the case where the same request arrives again for a reason that is not a retry — a resubmitted plan, a duplicated queue message, or an instruction injected twice.

Control 01

Per-capability replay window

A declared period during which a repeated key returns the prior result rather than executing. Twenty-four hours for payments; minutes for notifications; documented for everything.

Control 02

Outside-window behaviour

A repeat arriving after the window is a new intent, and should be treated as one — which means it faces the full policy path including any approval requirement, rather than being silently permitted as a retry.

Control 03

Attempt ceiling per key

Cap attempts on the same key. Ten attempts on one payment key is not a network problem; it is a loop, and it should trip a breaker rather than continue.

Control 04

Cross-run detection

The same intent key appearing in two different runs is usually a duplicated task rather than a retry. Worth an alert, because it points at an orchestration bug that will recur.

Control two is the one most often implemented incorrectly: a repeat outside the window that is treated as a retry produces a duplicate that the mechanism was supposed to prevent, and it will look correct in the logs.

When the upstream offers no idempotency

This is common, particularly with internal systems and older integrations. The governance layer has to supply what the upstream does not.

Key facts

  • ▸The unknown state is the crux. Most implementations collapse unknown into either success or failure, and both collapses cause incidents — one hides duplicates, the other creates them.
  • ▸A reconciliation query per capability is usually a few lines and eliminates most of the ambiguity a retry would otherwise resolve badly.
  • ▸Where escalation is the answer, make the approval request state clearly that the first attempt’s status is unknown, so the human knows which question they are answering.
  • 01Deduplicate at the decision point. Hold intent keys with their outcomes for the declared window, and on a match return the prior result rather than forwarding the call.
  • 02Distinguish committed from unknown. If the first attempt’s outcome is unknown, a dedupe that returns success is a lie. Return an explicit unknown state and require reconciliation before a further attempt.
  • 03Add a reconciliation query per capability. A cheap read that answers “did this intent already take effect?” Often a simple search by business reference, and it converts unknown into known.
  • 04Prefer check-then-act where a natural key exists. Search for the effect before creating it. Racy in theory, effective in practice at agent-scale intervals.
  • 05Require human confirmation for the residual. Where no reconciliation query exists and the effect is consequential, an unknown outcome should escalate rather than retry.
  • 06Push for upstream support. A governance-layer workaround is real mitigation and remains a workaround; the durable fix is an idempotency key the upstream honours.

Reconciliation as the backstop

Design reduces duplicates; it does not eliminate them. The total cost is dominated by detection delay, which makes reconciliation frequency the most important remaining variable.

Why frequency dominates

cost = detection_delay × duplicate_rate × unit_cost 15-min detection at 4 duplicates/hour = 1 duplicate exposed 24-hour detection at 4 duplicates/hour = 96 duplicates exposed

Improving detection from daily to fifteen-minutely usually costs less than any of the preventive engineering, which makes it the right first investment when time is short.

Duplicate detection design
Capability classDetection queryFrequencyEscalation
Payments and refundsSame business reference, same amount, within windowEvery 15 minPage; hold further attempts on that reference
Ledger and accounting writesDuplicate entry by reference and periodHourlyAlert finance owner; block capability pending review
Outbound email and messagingSame recipient, same template, within windowHourlyAlert; suppress further sends to that recipient
Record creationDuplicate by natural key or fuzzy matchDailyQueue for merge; report rate to capability owner
ProvisioningDuplicate resource by name or tagHourlyAlert platform owner; block until resolved

Write the detection query before granting the capability. If you cannot express what a duplicate looks like for a given tool, you cannot safely let an autonomous system call it repeatedly — and it will.

Next step

Deduplicate at the decision point

BarzelVault holds intent keys and attempt counts at the policy layer, so a retry on a non-idempotent capability is recognised, refused or escalated rather than executed a second time.

Residual risk

Two limits worth stating.

  • 01Idempotency cannot resolve a genuinely unknown upstream state. It can prevent a second attempt from committing a second effect; it cannot tell you whether the first one committed, which is why reconciliation queries matter more than key design in practice.
  • 02Governance-layer deduplication only covers calls that pass through the governance layer. A direct credential bypasses it, and the resulting duplicates are invisible to every control described here.

Frequently asked questions

Why do AI agents cause duplicate actions?

Because a timeout leaves the outcome unknown and an agent’s default response to an unknown outcome is to try again, quickly and repeatedly. Unlike a coded client with deliberate retry logic, the retry emerges from planning rather than from a decision anyone made about safety.

What is an idempotency key for agent actions?

A deterministic identifier supplied with a request so the receiving system can recognise a repeat and return the original result rather than performing the effect twice. For agents it must be derived from intent rather than generated per attempt.

How should an idempotency key be derived?

From agent identity, capability, a business reference such as an order or invoice id, the normalised semantically significant arguments, and a bounded intent window. Attempt counters, timestamps and request text must be excluded because they vary per attempt.

Can the language model generate the idempotency key?

No. A model rephrasing the same intent will produce a different key, which defeats the mechanism exactly when it is needed. Derive the key deterministically at the governance layer from validated inputs.

Why does argument normalisation matter for idempotency?

Because an amount expressed in minor units and the same amount in major units would otherwise hash to different keys, so unit ambiguity converts directly into duplicate payments. Normalise before hashing.

Which agent tools should never be retried automatically?

Non-idempotent ones: sending email or messages, appending ledger entries, and creating records with no natural key. These need governance-layer deduplication on an intent hash or a human confirmation step rather than an automatic retry.

What is replay protection and how does it differ from idempotency?

Idempotency handles an honest retry of the same intent. Replay protection handles a repeat that is not a retry — a resubmitted plan, a duplicated queue message, an injected instruction arriving twice — by refusing or re-deciding requests that repeat outside their declared window.

How long should a replay window be?

Per capability, and documented either way: around twenty-four hours for payments and refunds, minutes for notifications. A repeat arriving after the window is a new intent and must face the full policy path including any approval requirement, not be silently treated as a retry.

What if the upstream system does not support idempotency keys?

Provide it at the decision point: hold intent keys with outcomes for the declared window and return the prior result on a match. Critically, distinguish unknown from committed — a dedupe that reports success when the first attempt’s outcome is unknown is worse than no dedupe.

How should an unknown outcome be handled?

Explicitly, as its own state. Add a cheap reconciliation query per capability that answers whether the intent already took effect, and where no such query exists and the action is consequential, escalate to a human whose approval request clearly says the first attempt’s status is unknown.

Is reconciliation necessary if idempotency is implemented?

Yes. Some duplicates will occur, and total cost is dominated by detection delay rather than by any individual duplicate. Improving detection from daily to every fifteen minutes usually costs less than the equivalent preventive engineering.

Where should retry policy be configured?

On the capability in the registry, so it is a governed property rather than a client setting that varies with whoever wrote the agent. Attempt ceilings should be enforced at the decision point as well, because client-side caps are bypassed by the next client.

Glossary

Idempotency
The property that repeating an operation produces no additional effect beyond the first application.
Idempotency key
A client-supplied identifier allowing a server to recognise and collapse repeated requests.
Intent-derived key
A key computed from the semantic intent of a request rather than from any per-attempt value.
Intent window
The bounded period within which repeats of an intent are treated as the same intent.
Replay protection
Refusing or re-deciding a request that repeats outside its legitimate retry window.
Unknown outcome
A state in which whether an effect committed cannot be determined, distinct from success or failure.
Reconciliation query
A cheap read that determines whether an intent has already taken effect.
Retry-safety category
A classification of a tool determining whether and how it may be retried automatically.
Attempt ceiling
A hard cap on repeats of the same key, enforced at the decision point.
Cross-run duplication
The same intent key appearing in two different agent runs, indicating an orchestration fault.

Standards and entities referenced

Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.

  1. 01 · WikipediaIdempotence ↗Why safe retries require this property rather than hope.
  2. 02 · microservices.ioSaga pattern ↗Compensating transactions, which is all you get when distributed rollback does not exist.
  3. 03 · Michael NygardRelease It! — stability patterns ↗Bulkheads, circuit breakers and timeouts, from the book that named them.
  4. 04 · PCI Security Standards CouncilPCI DSS v4.0 ↗Where payment-adjacent agent actions inherit real, externally audited requirements.
  5. 05 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
  6. 06 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  7. 07 · JSON-RPC Working GroupJSON-RPC 2.0 Specification ↗The request/response envelope every MCP tool call travels in.
  8. 08 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.

Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). Idempotency and Replay Safety for AI Agents: Why Retries Become Incidents. Real Biz Digital. https://realbizdigital.net/insights/ai-agent-idempotency/

Try the mechanics on a live server

To see what a tool-call envelope actually looks like before you write a policy that has to decide about one — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

BarzelVault is the pre-execution decision point, sold as a running product

Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.

PlanPriceIncludedRight for
DevFree10,000 policy decisions/mo · 9 tools, 4 outcomes, hash-chained auditA first regulated workflow: one agent, one high-consequence system
Team$199/mo75,000 decisions/mo · approval workflow, spend and action limitsSeveral agents acting on money, records or customer-visible systems
Business$799/mo750,000 decisions/mo · exact HTTPS execution, credential isolation, emergency controlsEnterprise-wide pre-execution enforcement under audit
Enterprise$3,999/mo5,000,000 decisions/mo · everything in Business, scaledGroup-wide rollout across many teams and systems

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

The five Barzel servers, and which problem each one is sold for

One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.

ServerSold forEntry priceWhere it sits
Barzel Central GatewayKnowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidenceFree, then $10–$149/moControl plane — decides what may be reached, and by whom
BarzelVaultStopping a specific dangerous action before it executes, with proof afterwards$199–$3,999/moDecision point — evaluates the individual call before execution
BarzelOpsRunning real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approvalFree, then $19–$199/moExecution layer — does the work the policy allowed
Barzel FinOps AtlasAttributing AI spend to agents, tools and outcomes, then forecasting and capping itFree, then $29–$799/moEconomics layer — what the estate costs per outcome
Barzel Scripture IntelligenceA free, credential-free public MCP server to test clients and inspect real protocol trafficFree, unmetered, no signupReference implementation — safe place to learn the protocol

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.