AI Agent Security · Execution safety
Idempotency and Replay Safety for AI Agents: Why Retries Become Incidents
An agent that times out on a payment does not know whether the payment happened. It will try again. Whether that is a retry or a second payment is decided by design you either did or did not do.
The short answer
Idempotency for AI agents means every consequential tool call carries a client-generated key derived from the intent, so that a repeated call with the same key returns the original result rather than causing a second effect. Without it, an agent that cannot tell whether its first attempt succeeded will retry, and the retry becomes a duplicate payment, a duplicate email or a duplicate record. Agents retry more than humans do, more quickly, and with less ability to check first. That combination turns an ordinary distributed-systems problem into a financial one.
Summary for readers and answer engines
Reviewed 25 Aug 2026
- ▸Agents retry aggressively and cannot easily check whether the first attempt landed, so ambiguity resolves into duplication by default.
- ▸The key must be derived from intent, not generated per attempt. A fresh key on each retry is not idempotency; it is a second request.
- ▸Classify every tool into one of four retry-safety categories and give each category a different retry policy. Naturally idempotent reads and non-idempotent writes should not share retry configuration.
- ▸Where upstream systems offer no idempotency, the governance layer has to provide it: dedupe at the decision point on an intent hash, within a declared window.
- ▸Reconciliation is the backstop. Some duplicates will occur, and the cost of a duplicate is dominated by how long it goes undetected.
Source: Mark Alex, Real Biz Digital — Idempotency and Replay Safety for AI Agents: Why Retries Become Incidents (https://realbizdigital.net/insights/ai-agent-idempotency/). Reproduce with attribution.
Key takeaways
- 01Derive keys from intent: agent, capability, and the semantically significant arguments. Stable across attempts, distinct across intents.
- 02Never let the model generate the idempotency key. It will produce a different one when it rephrases the same intent.
- 03Set a replay window per capability and state it. Twenty-four hours for payments, minutes for notifications, and a documented answer either way.
- 04Treat timeouts as unknown rather than failed. An agent told a call failed will retry; an agent told the status is unknown can be made to reconcile first.
- 05Cap retries at the decision point, not only in the client. Client-side caps are bypassed by the next client.
- 06Build the duplicate-detection query before granting the capability. It is the only control that catches what design missed.
Quick answers
One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.
- Why do agents cause duplicate actions?
- Because a timeout leaves the outcome unknown, and an agent’s default response to an unknown outcome is to try again. Unlike a human, it does not pause to check first.
- What is an idempotency key?
- A client-supplied identifier that lets a server recognise a repeated request and return the original result instead of performing the action twice.
- How should the key be derived?
- From the intent: agent identity, capability, and the semantically significant arguments. It must be stable across retries of the same intent and distinct across different intents.
- Can the model generate the key?
- No. A model rephrasing the same intent will produce a different key, which defeats the mechanism precisely when it is needed.
- What if the upstream system has no idempotency support?
- Provide it at the governance layer: deduplicate on an intent hash at the decision point within a declared window, and refuse or return the prior result on a match.
- What is replay protection?
- Refusing a request that repeats outside its legitimate retry window, which distinguishes a genuine retry from a replayed or duplicated instruction.
- Is reconciliation still necessary?
- Yes. Some duplicates will occur, and the total cost is dominated by detection delay rather than by any individual duplicate.
Why agent retries are different
Retry safety is an old problem. What changed is who is retrying, how fast, and how little they check.
Two properties compound the problem. Agents interpret a timeout as a failure and act accordingly, and they can iterate faster than any monitoring interval — a retry loop can produce hundreds of attempts before a five-minute alert window closes.
Add a third property specific to language models: the same intent can be expressed differently on each attempt. Anything keyed on the exact request text will treat two attempts as two intents, which is exactly wrong.
- ›Retries are coded deliberately, with backoff
- ›Developer considered idempotency at design time
- ›Failure paths are explicit and tested
- ›A human notices a duplicate quickly
- ›Retries emerge from planning, not from code
- ›Idempotency was nobody’s decision at that moment
- ›Failure paths are inferred from an error string
- ›Duplicates surface in reconciliation, days later
The cost of ambiguity
timeout on a payment call P(succeeded upstream) ≈ unknown, often 0.3-0.7 agent default: retry without idempotency: expected duplicates = attempts × P(first succeeded)
The mechanism does not need to be clever. It needs to exist, because the default behaviour in the presence of ambiguity is duplication.
Key derivation, scope and lifetime
Most idempotency failures are key design failures rather than missing support. Three properties matter: stable across retries, distinct across intents, and scoped correctly.
- 01Normalise significant arguments before hashing. An amount in minor units and the same amount in major units must produce the same key, or unit ambiguity becomes duplicate payments.
- 02Scope by business reference where one exists. “Refund for order 10432” is a better intent boundary than any hash of the request.
- 03Choose the intent window deliberately. A single run for one-off actions; a calendar day for recurring ones; never unbounded, or a legitimate second refund next month is refused.
- 04Exclude anything that varies per attempt. Timestamps, attempt counters and request text all break stability, which is the property the mechanism depends on.
- 05Never let the model generate or choose the key. Derive it deterministically at the governance layer from validated inputs.
- 06Record the key on the decision record. Without it you cannot distinguish a duplicate that was caught from one that was not.
Intent-derived key (illustrative)
key = sha256( agent_id + # who intends it capability + # what kind of effect business_ref + # order id, invoice id, ticket id significant_args + # amount, currency, recipient - normalised intent_window # e.g. date, or run id for single-run intents ) # NOT included: attempt number, timestamp, request text, # model-generated free text, retry counter
A useful test: ask two engineers to write the key derivation for a refund independently. If they disagree about whether the customer id or the order id is the boundary, your intent model is ambiguous and duplicates are a matter of time.
Four retry-safety categories and their controls
Key facts
- ▸Category three requires verification, not assumption. Test that the upstream actually honours the key by sending the same key twice and comparing results — support is sometimes partial or silently ignored.
- ▸Category four is where most real incidents live: notification storms, duplicate ledger entries, repeated ticket creation. These systems rarely offer idempotency and are rarely treated as dangerous.
- ▸Retry policy belongs on the capability in the registry, so it is a governed property rather than a client configuration that varies by whoever wrote the agent.
| Category | Example | Retry policy | Additional control |
|---|---|---|---|
| Naturally idempotent read | Fetch a record, search, list | Retry freely with backoff | Rate cap only |
| Naturally idempotent write | Set a field to a value, upsert by key | Retry with backoff | Cap attempts; log repeats |
| Idempotent with a key | Payments, refunds, provisioning with a supported key | Retry only with the same key | Refuse retry if key absent; verify server honours it |
| Non-idempotent | Send email, append to a ledger, create without a natural key | Do not retry automatically | Governance-layer dedupe on intent hash, or human confirmation |
Classifying tools into these four categories is a half-day exercise with a disproportionate payoff, because it immediately identifies which capabilities must never be retried automatically.
Replay protection: distinguishing a retry from a repeat
Idempotency handles the honest retry. Replay protection handles the case where the same request arrives again for a reason that is not a retry — a resubmitted plan, a duplicated queue message, or an instruction injected twice.
Per-capability replay window
A declared period during which a repeated key returns the prior result rather than executing. Twenty-four hours for payments; minutes for notifications; documented for everything.
Outside-window behaviour
A repeat arriving after the window is a new intent, and should be treated as one — which means it faces the full policy path including any approval requirement, rather than being silently permitted as a retry.
Attempt ceiling per key
Cap attempts on the same key. Ten attempts on one payment key is not a network problem; it is a loop, and it should trip a breaker rather than continue.
Cross-run detection
The same intent key appearing in two different runs is usually a duplicated task rather than a retry. Worth an alert, because it points at an orchestration bug that will recur.
Control two is the one most often implemented incorrectly: a repeat outside the window that is treated as a retry produces a duplicate that the mechanism was supposed to prevent, and it will look correct in the logs.
When the upstream offers no idempotency
This is common, particularly with internal systems and older integrations. The governance layer has to supply what the upstream does not.
Key facts
- ▸The unknown state is the crux. Most implementations collapse unknown into either success or failure, and both collapses cause incidents — one hides duplicates, the other creates them.
- ▸A reconciliation query per capability is usually a few lines and eliminates most of the ambiguity a retry would otherwise resolve badly.
- ▸Where escalation is the answer, make the approval request state clearly that the first attempt’s status is unknown, so the human knows which question they are answering.
- 01Deduplicate at the decision point. Hold intent keys with their outcomes for the declared window, and on a match return the prior result rather than forwarding the call.
- 02Distinguish committed from unknown. If the first attempt’s outcome is unknown, a dedupe that returns success is a lie. Return an explicit unknown state and require reconciliation before a further attempt.
- 03Add a reconciliation query per capability. A cheap read that answers “did this intent already take effect?” Often a simple search by business reference, and it converts unknown into known.
- 04Prefer check-then-act where a natural key exists. Search for the effect before creating it. Racy in theory, effective in practice at agent-scale intervals.
- 05Require human confirmation for the residual. Where no reconciliation query exists and the effect is consequential, an unknown outcome should escalate rather than retry.
- 06Push for upstream support. A governance-layer workaround is real mitigation and remains a workaround; the durable fix is an idempotency key the upstream honours.
Reconciliation as the backstop
Design reduces duplicates; it does not eliminate them. The total cost is dominated by detection delay, which makes reconciliation frequency the most important remaining variable.
Why frequency dominates
cost = detection_delay × duplicate_rate × unit_cost 15-min detection at 4 duplicates/hour = 1 duplicate exposed 24-hour detection at 4 duplicates/hour = 96 duplicates exposed
Improving detection from daily to fifteen-minutely usually costs less than any of the preventive engineering, which makes it the right first investment when time is short.
| Capability class | Detection query | Frequency | Escalation |
|---|---|---|---|
| Payments and refunds | Same business reference, same amount, within window | Every 15 min | Page; hold further attempts on that reference |
| Ledger and accounting writes | Duplicate entry by reference and period | Hourly | Alert finance owner; block capability pending review |
| Outbound email and messaging | Same recipient, same template, within window | Hourly | Alert; suppress further sends to that recipient |
| Record creation | Duplicate by natural key or fuzzy match | Daily | Queue for merge; report rate to capability owner |
| Provisioning | Duplicate resource by name or tag | Hourly | Alert platform owner; block until resolved |
Write the detection query before granting the capability. If you cannot express what a duplicate looks like for a given tool, you cannot safely let an autonomous system call it repeatedly — and it will.
Next step
Deduplicate at the decision point
BarzelVault holds intent keys and attempt counts at the policy layer, so a retry on a non-idempotent capability is recognised, refused or escalated rather than executed a second time.
Residual risk
Two limits worth stating.
- 01Idempotency cannot resolve a genuinely unknown upstream state. It can prevent a second attempt from committing a second effect; it cannot tell you whether the first one committed, which is why reconciliation queries matter more than key design in practice.
- 02Governance-layer deduplication only covers calls that pass through the governance layer. A direct credential bypasses it, and the resulting duplicates are invisible to every control described here.
Frequently asked questions
Why do AI agents cause duplicate actions?
Because a timeout leaves the outcome unknown and an agent’s default response to an unknown outcome is to try again, quickly and repeatedly. Unlike a coded client with deliberate retry logic, the retry emerges from planning rather than from a decision anyone made about safety.
What is an idempotency key for agent actions?
A deterministic identifier supplied with a request so the receiving system can recognise a repeat and return the original result rather than performing the effect twice. For agents it must be derived from intent rather than generated per attempt.
How should an idempotency key be derived?
From agent identity, capability, a business reference such as an order or invoice id, the normalised semantically significant arguments, and a bounded intent window. Attempt counters, timestamps and request text must be excluded because they vary per attempt.
Can the language model generate the idempotency key?
No. A model rephrasing the same intent will produce a different key, which defeats the mechanism exactly when it is needed. Derive the key deterministically at the governance layer from validated inputs.
Why does argument normalisation matter for idempotency?
Because an amount expressed in minor units and the same amount in major units would otherwise hash to different keys, so unit ambiguity converts directly into duplicate payments. Normalise before hashing.
Which agent tools should never be retried automatically?
Non-idempotent ones: sending email or messages, appending ledger entries, and creating records with no natural key. These need governance-layer deduplication on an intent hash or a human confirmation step rather than an automatic retry.
What is replay protection and how does it differ from idempotency?
Idempotency handles an honest retry of the same intent. Replay protection handles a repeat that is not a retry — a resubmitted plan, a duplicated queue message, an injected instruction arriving twice — by refusing or re-deciding requests that repeat outside their declared window.
How long should a replay window be?
Per capability, and documented either way: around twenty-four hours for payments and refunds, minutes for notifications. A repeat arriving after the window is a new intent and must face the full policy path including any approval requirement, not be silently treated as a retry.
What if the upstream system does not support idempotency keys?
Provide it at the decision point: hold intent keys with outcomes for the declared window and return the prior result on a match. Critically, distinguish unknown from committed — a dedupe that reports success when the first attempt’s outcome is unknown is worse than no dedupe.
How should an unknown outcome be handled?
Explicitly, as its own state. Add a cheap reconciliation query per capability that answers whether the intent already took effect, and where no such query exists and the action is consequential, escalate to a human whose approval request clearly says the first attempt’s status is unknown.
Is reconciliation necessary if idempotency is implemented?
Yes. Some duplicates will occur, and total cost is dominated by detection delay rather than by any individual duplicate. Improving detection from daily to every fifteen minutes usually costs less than the equivalent preventive engineering.
Where should retry policy be configured?
On the capability in the registry, so it is a governed property rather than a client setting that varies with whoever wrote the agent. Attempt ceilings should be enforced at the decision point as well, because client-side caps are bypassed by the next client.
Glossary
- Idempotency
- The property that repeating an operation produces no additional effect beyond the first application.
- Idempotency key
- A client-supplied identifier allowing a server to recognise and collapse repeated requests.
- Intent-derived key
- A key computed from the semantic intent of a request rather than from any per-attempt value.
- Intent window
- The bounded period within which repeats of an intent are treated as the same intent.
- Replay protection
- Refusing or re-deciding a request that repeats outside its legitimate retry window.
- Unknown outcome
- A state in which whether an effect committed cannot be determined, distinct from success or failure.
- Reconciliation query
- A cheap read that determines whether an intent has already taken effect.
- Retry-safety category
- A classification of a tool determining whether and how it may be retried automatically.
- Attempt ceiling
- A hard cap on repeats of the same key, enforced at the decision point.
- Cross-run duplication
- The same intent key appearing in two different agent runs, indicating an orchestration fault.
Standards and entities referenced
Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.
- 01 · WikipediaIdempotence ↗Why safe retries require this property rather than hope.
- 02 · microservices.ioSaga pattern ↗Compensating transactions, which is all you get when distributed rollback does not exist.
- 03 · Michael NygardRelease It! — stability patterns ↗Bulkheads, circuit breakers and timeouts, from the book that named them.
- 04 · PCI Security Standards CouncilPCI DSS v4.0 ↗Where payment-adjacent agent actions inherit real, externally audited requirements.
- 05 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
- 06 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 07 · JSON-RPC Working GroupJSON-RPC 2.0 Specification ↗The request/response envelope every MCP tool call travels in.
- 08 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). Idempotency and Replay Safety for AI Agents: Why Retries Become Incidents. Real Biz Digital. https://realbizdigital.net/insights/ai-agent-idempotency/
Try the mechanics on a live server
To see what a tool-call envelope actually looks like before you write a policy that has to decide about one — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
BarzelVault is the pre-execution decision point, sold as a running product
Nine tools, 12 static resources, 3 resource templates and 9 prompts. Four deterministic outcomes — allow, deny, dry-run, require approval — with approval workflow, hash-chained audit and guardrail data protection. Streamable HTTP, JSON-RPC 2.0.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Dev | Free | 10,000 policy decisions/mo · 9 tools, 4 outcomes, hash-chained audit | A first regulated workflow: one agent, one high-consequence system |
| Team | $199/mo | 75,000 decisions/mo · approval workflow, spend and action limits | Several agents acting on money, records or customer-visible systems |
| Business | $799/mo | 750,000 decisions/mo · exact HTTPS execution, credential isolation, emergency controls | Enterprise-wide pre-execution enforcement under audit |
| Enterprise | $3,999/mo | 5,000,000 decisions/mo · everything in Business, scaled | Group-wide rollout across many teams and systems |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
The five Barzel servers, and which problem each one is sold for
One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.
| Server | Sold for | Entry price | Where it sits |
|---|---|---|---|
| Barzel Central Gateway | Knowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidence | Free, then $10–$149/mo | Control plane — decides what may be reached, and by whom |
| BarzelVault | Stopping a specific dangerous action before it executes, with proof afterwards | $199–$3,999/mo | Decision point — evaluates the individual call before execution |
| BarzelOps | Running real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approval | Free, then $19–$199/mo | Execution layer — does the work the policy allowed |
| Barzel FinOps Atlas | Attributing AI spend to agents, tools and outcomes, then forecasting and capping it | Free, then $29–$799/mo | Economics layer — what the estate costs per outcome |
| Barzel Scripture Intelligence | A free, credential-free public MCP server to test clients and inspect real protocol traffic | Free, unmetered, no signup | Reference implementation — safe place to learn the protocol |
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.