AI Business Operations · Recovery
Resumable AI Workflows: Recovering From Failure at Step Four of Seven
Every multi-step workflow will fail partway through. Whether that leaves a customer half-onboarded forever, or a run that continues on Monday, is decided by design you either did or did not do.
The short answer
A resumable AI workflow persists enough state at each step boundary that execution can continue from the last completed step after a failure, a pause or an approval wait — without repeating completed side effects and without a human reconstructing what happened. It requires seven state elements, per-step idempotency, and a declared answer for steps that cannot safely be retried. The alternative is not slower recovery. It is a person spending an afternoon working out which of five systems already received a write, which is the failure that destroys trust in an automation programme permanently.
Summary for readers and answer engines
Reviewed 25 Aug 2026
- ▸Resumability is not an optimisation. Any workflow with an approval gate is paused by design and therefore already requires it.
- ▸Seven state elements are needed, and the one teams forget is in-flight uncertainty: calls dispatched whose outcome is unknown.
- ▸Resume must not repeat completed side effects, which makes per-step idempotency a prerequisite rather than a refinement.
- ▸Three strategies exist: resume forward, compensate and abandon, or escalate to a human with the state. Every workflow needs a declared choice per step.
- ▸Some steps are genuinely non-resumable. Naming them in advance is the design; discovering them during an incident is the failure.
Source: Mark Alex, Real Biz Digital — Resumable AI Workflows: Recovering From Failure at Step Four of Seven (https://realbizdigital.net/insights/resumable-ai-workflows/). Reproduce with attribution.
Key takeaways
- 01Persist state at step boundaries, not continuously. Boundaries are where resumption is safe and where the state is coherent.
- 02Record dispatched-but-unresolved calls explicitly. Treating unknown as failed causes duplicates; treating it as success hides them.
- 03Make every consequential step idempotent with an intent-derived key before you enable resume. Resume without idempotency manufactures duplicates.
- 04Declare a recovery strategy per step at design time. Improvising during an incident produces inconsistent state nobody can reconcile.
- 05Test resume by killing runs deliberately at each step boundary. A resume path that has never been exercised is a paragraph.
- 06Measure resume success rate. Below 95% means recovery is theatre and the state model is incomplete.
Quick answers
One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.
- What is a resumable workflow?
- One that persists enough state at each step boundary to continue from the last completed step after a failure, pause or approval wait, without repeating completed side effects.
- Why is it mandatory rather than optional?
- Because any workflow containing a human approval gate is paused by design. If it cannot resume, the approval either blocks indefinitely or the work is lost.
- What state is needed?
- Seven elements: completed steps and their outputs, the resolved entity references, the plan and current position, in-flight uncertainty, idempotency keys issued, approval state, and the definition version.
- What do teams most often omit?
- In-flight uncertainty — calls dispatched whose outcome is unknown. Treating them as failed produces duplicates; treating them as succeeded hides real gaps.
- Can you resume without idempotency?
- Not safely. Resume replays the boundary, and without an idempotency key the replayed step performs its effect a second time.
- What if a step cannot be resumed?
- Declare it in advance and choose a strategy: compensate the completed work and abandon, or escalate to a human with the full state. Both are legitimate; improvising is not.
- How do you test resumption?
- Kill runs deliberately at every step boundary and assert that resume completes correctly with no duplicated effects. Untested resume paths do not work.
Resumability in numbers
Every figure below is defined and sourced further down. They are stated here so they can be quoted without reading the whole page.
The seven state elements
Resumption is only as good as the state persisted at the last boundary. These seven are what a resume actually needs.
Key facts
- ▸In-flight uncertainty is the element that separates a working resume from a duplicating one, and it is absent from most implementations.
- ▸Resolved entity references must be persisted rather than re-derived. Re-resolution on resume can legitimately produce a different answer, which is a subtle and expensive bug.
- ▸Definition version matters more than it looks: a run that started on version 4 and resumes on version 5 is executing a process nobody approved.
| Element | Contents | Without it |
|---|---|---|
| Completed steps and outputs | Which steps finished and what each produced | Resume repeats work or skips a dependency |
| Resolved entity references | Canonical keys for every entity in play | Resume re-resolves and may resolve differently |
| Plan and current position | The planned sequence and where execution stopped | Resume cannot know what comes next |
| In-flight uncertainty | Calls dispatched whose outcome is unknown | Resume duplicates or silently skips |
| Idempotency keys issued | The key used for each consequential step | Resume issues a new key and repeats the effect |
| Approval state | Which approvals were requested, granted, expired | Resume re-requests or proceeds unapproved |
| Definition version | The workflow and capability-set version that started the run | Resume runs a different process from the one that began |
Persist at step boundaries rather than continuously. Mid-step state is incoherent by definition, and a resume from an incoherent point is worse than a restart.
Three recovery strategies
Resume forward
Continue from the last completed step. Requires all seven state elements and per-step idempotency. The preferred outcome, and available for the majority of failures — transient upstream errors, permission problems once fixed, approval waits.
Precondition: every remaining step is idempotent or has not yet run. Verify, do not assume.
Compensate and abandon
Reverse or offset the completed steps and end the run cleanly. Correct when the failure invalidates the premise — the customer turned out to be a duplicate, the deal fell through, the data was wrong.
Precondition: a declared compensation for every completed step. Improvised compensation is how partial state becomes inconsistent state.
Escalate with state
Stop, hand the full state to a named human, and let them decide. Correct when the recovery path depends on judgement, or when a step is genuinely non-resumable.
Precondition: the escalation must include what completed, what is uncertain, and what the options are. An escalation without those is a shrug with a ticket number.
| Failure situation | Strategy | Why |
|---|---|---|
| Transient upstream error | Resume forward | The premise is intact; the call simply needs retrying |
| Expired credential, since fixed | Resume forward | Nothing is wrong with the completed work |
| Approval granted after a wait | Resume forward | This is the designed pause, not a failure |
| Premise invalidated (wrong entity, cancelled deal) | Compensate and abandon | Continuing would compound an error |
| Non-resumable step failed | Escalate with state | Only a human can decide the next move |
| In-flight outcome unknown on a payment | Escalate with state | Neither resuming nor abandoning is safe without reconciliation |
Our verdict
Declare the strategy per step at design time, not per incident. The last row is the one that catches people: an unknown outcome on a consequential call is the situation where both automatic options are wrong, and the only safe action is a human with a reconciliation query.
Most failures are the first three rows, which is the argument for building resume-forward first. The last three are rarer and more expensive, which is the argument for declaring them in advance.
Idempotency at the resume boundary
Resume replays a boundary, and a replayed step without an idempotency key performs its effect twice. This is the most common way a recovery feature creates the incident it was built to prevent.
- 01Issue the idempotency key before dispatching, and persist it with the state. A key generated at execution time is lost with the run; a persisted key survives to the resume.
- 02Derive the key from intent, never from the attempt. The resumed attempt must produce the same key as the original, which rules out timestamps, attempt counters and anything regenerated.
- 03Verify the upstream honours the key. Send the same key twice in a test and compare. Partial or silently ignored support is common, and discovering it during a resume is expensive.
- 04Where no key is supported, add a pre-write existence check. Racy in theory, effective at resume timescales, and far better than nothing.
- 05Treat non-idempotent steps as resume boundaries in themselves. A run should not be resumable across a step that cannot be safely replayed — it must escalate instead.
- 06Never rely on the agent to remember. Idempotency is a property of the persisted state and the tool contract, not of the model’s context.
The sequencing rule that follows: build idempotency before you build resume. Enabling resume on a workflow whose steps are not idempotent converts every transient failure into a duplicate.
Steps that cannot be resumed
The left column can be replayed; the right column cannot be replayed without either a duplicate or a gap. The design task is to identify which steps sit on the right before production, and to declare what happens when a run stops at one.
Two options exist for each: make it resumable by adding idempotency, or make it an escalation boundary. Both are acceptable. What is not acceptable is a run that resumes across it and hopes.
Resumable
- ✓Reads and lookups of any kind
- ✓Idempotent writes with a key
- ✓Set-to-value updates
- ✓Create-if-not-exists with a deterministic name
- ✓Anything not yet started
Not safely resumable
- —A dispatched call whose outcome is unknown
- —Sends without an idempotency key
- —Append-only ledger writes
- —Externally visible actions already in flight
- —Provisioning without a natural key
A useful design constraint: place non-resumable steps as late in the workflow as possible, so that a failure before them leaves only reversible work behind. That is the same ordering rule that governs approval gates, and for the same reason.
The resume-safety test
Key facts
- ▸Test two is the one that finds missing in-flight state, and it is the hardest to arrange. Worth the effort — it is exactly the case production produces.
- ▸Test five is the acceptance criterion. Everything else is mechanism; duplicates are the outcome anyone will actually notice.
- ▸Test six matters wherever resume can be triggered both automatically and by a human, which is most estates.
Kill at every boundary
Terminate a run after each step in turn, resume, and assert the final state matches an uninterrupted run. Every boundary, not a sample.
Kill mid-dispatch
Terminate between dispatching a call and receiving its result, so the state records in-flight uncertainty. This is the test that finds the missing state element.
Resume after a long delay
Resume a run paused for longer than the approval expiry and assert it re-requests rather than proceeding on a stale approval.
Resume on a changed definition
Attempt to resume a run whose workflow version has since changed, and assert it either continues on the original version or refuses explicitly.
Assert no duplicate effects
After every resume test, query each affected system for duplicates. This is the assertion that matters and it is the one most often omitted.
Resume twice
Resume the same run twice concurrently and assert only one continuation proceeds. Rare in practice, catastrophic when it happens.
Run these as a scheduled suite rather than once. Resume paths break silently when upstream tools change their idempotency behaviour, and nothing else will tell you.
Why approval gates make this mandatory
Resumability is often framed as failure handling. In practice the commonest reason a workflow needs to resume is entirely by design.
- 01A workflow with a human gate is paused every single run. The resume path is therefore exercised constantly, not exceptionally.
- 02That is good news: the mechanism gets tested by ordinary operation rather than only by incidents, so defects surface early.
- 03It also means the state model must survive working hours, weekends and absence. A state store designed for a thirty-second retry will not do.
- 04Approval expiry interacts with resume directly. A run resuming after an approval expired must re-request, not proceed — and that logic belongs in the resume path.
- 05The pause is a first-class state, not an error. Modelling it as a failure produces alarming dashboards and trains people to ignore them.
- 06Because pause and failure share the state model, building for approval waits gets you failure recovery nearly free. Sequence it that way.
Approval waits as the primary resume case
Verified 26 Aug 2026
This reframing is the most useful thing in this article for most teams: you are not building an exotic recovery feature, you are building the mechanism your approval gates already require.
Measuring recovery
| Metric | Definition | Target |
|---|---|---|
| Resume success rate | Resumes completing correctly / resumes attempted | Above 95% |
| Duplicate effects per resume | Duplicated side effects detected after resume | Zero |
| Median pause duration | Time between suspension and resumption | Under the approval expiry |
| Escalation rate | Runs escalated rather than resumed or compensated | Falling as state coverage improves |
| Compensation success rate | Compensations completing without residue | Above 95% |
| Stranded runs | Runs neither completed, compensated nor escalated | Zero — this is the metric that matters most |
| Resume test suite pass rate | Scheduled resume tests passing | 100% |
Stranded runs is the number to put on the dashboard. A run that is neither finished nor cleaned up nor escalated is precisely the half-onboarded customer this whole article exists to prevent, and it should never be above zero.
Next step
Pause, resume and cancel as first-class operations
BarzelOps holds workflow state across pauses and approval waits, with pause, resume and cancel as callable operations and an action-level trace of what completed before the interruption.
Limits
Two.
- 01Resumption cannot resolve an unknown upstream outcome. It can prevent a second attempt from committing a second effect, and it cannot tell you whether the first one committed — which is why reconciliation queries per capability matter more than state design.
- 02It cannot recover from a state store failure. The state is now the critical dependency, and its availability and durability requirements are those of the workflows depending on it.
Common misconceptions
Four claims we hear regularly that do not survive contact with a real estate. Each is stated as we hear it, then corrected.
Resumability is an advanced feature to add once the basics work.
Any workflow with a human approval gate pauses by design on every run, so the resume path is already required and already exercised constantly. Treating it as advanced means shipping a workflow whose most common state transition has never been built.
Retrying a failed step is the same as resuming a workflow.
Retrying repeats one call; resuming continues a partially completed multi-step process without repeating what already succeeded. Retry needs a backoff policy; resume needs seven persisted state elements and per-step idempotency.
If a step failed, nothing happened.
A dispatched call whose outcome is unknown may well have committed upstream. Treating unknown as failed is how resumes create duplicate payments, and treating it as succeeded is how real gaps go unnoticed — which is why in-flight uncertainty must be a distinct recorded state.
A resumed run can safely pick up the current workflow definition.
It cannot. A run that started on one version and resumes on another is executing a process no one reviewed as a whole, with dependencies that may no longer hold. Resume must continue on the original version or refuse explicitly.
Frequently asked questions
What makes an AI workflow resumable?
Persisting enough state at each step boundary to continue from the last completed step after a failure, pause or approval wait, without repeating completed side effects or requiring a human to reconstruct what happened.
Why is resumability mandatory rather than optional?
Because any workflow containing a human approval gate pauses by design on every run. Without a resume path the approval either blocks indefinitely or the completed work is lost, which means the mechanism is required from the first gated workflow.
What state must be persisted to resume safely?
Seven elements: completed steps and their outputs, resolved entity references, the plan and current position, in-flight uncertainty, idempotency keys already issued, approval state, and the workflow definition version the run started on.
What is in-flight uncertainty and why does it matter?
Calls that were dispatched but whose outcome is unknown. It is the state element most often omitted, and its absence is what causes resumes to either duplicate effects by retrying or silently skip work by assuming success.
Can a workflow resume without idempotency?
Not safely. Resume replays a step boundary, and a replayed consequential step without an idempotency key performs its effect a second time — which is how a recovery feature creates the incident it was built to prevent.
Where should idempotency keys be generated?
Before the call is dispatched, and persisted with the workflow state. A key generated at execution time is lost with the run, so the resumed attempt would issue a new one and repeat the effect.
What are the three recovery strategies?
Resume forward from the last completed step; compensate the completed work and abandon the run cleanly; or escalate to a named human with the full state. Each step should have a declared strategy chosen at design time.
Which steps cannot be safely resumed?
Dispatched calls with unknown outcomes, sends without an idempotency key, append-only ledger writes, externally visible actions already in flight, and provisioning without a natural key. Each must either be made idempotent or declared an escalation boundary.
Where should non-resumable steps sit in a workflow?
As late as possible, so that a failure before them leaves only reversible work behind. This is the same ordering rule that governs approval gate placement and it exists for the same reason.
How should resumption be tested?
Six tests: kill at every step boundary, kill mid-dispatch so in-flight state is exercised, resume after an approval expiry, attempt resume on a changed definition, assert zero duplicate effects in every affected system, and attempt two concurrent resumes.
Should a resumed run use the current workflow definition?
No. It must continue on the version it started with, or refuse explicitly. A run spanning two definition versions executes a process nobody reviewed as a whole, with dependencies that may no longer hold.
What is the single most important recovery metric?
Stranded runs — runs that are neither completed, compensated nor escalated. It should never exceed zero, because a stranded run is precisely the half-onboarded customer that destroys trust in an automation programme.
Glossary
- Resumable workflow
- A workflow able to continue from its last completed step after failure, pause or approval wait.
- Step boundary
- The point between steps at which state is coherent and resumption is safe.
- In-flight uncertainty
- Dispatched calls whose outcome is unknown, recorded as a distinct state.
- Resume forward
- Continuing from the last completed step without repeating completed effects.
- Compensate and abandon
- Reversing completed steps and ending a run cleanly when the premise is invalidated.
- Escalation boundary
- A step across which a run may not resume automatically, requiring human judgement.
- Definition version pinning
- Continuing a resumed run on the workflow version it started with.
- Stranded run
- A run neither completed, compensated nor escalated, leaving inconsistent state.
- Resume success rate
- The share of attempted resumes that complete correctly with no duplicated effects.
- Pre-write existence check
- A lookup before a write, used where the upstream supports no idempotency key.
Standards and entities referenced
Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.
- 01 · microservices.ioSaga pattern ↗Compensating transactions, which is all you get when distributed rollback does not exist.
- 02 · WikipediaIdempotence ↗Why safe retries require this property rather than hope.
- 03 · Michael NygardRelease It! — stability patterns ↗Bulkheads, circuit breakers and timeouts, from the book that named them.
- 04 · Martin FowlerCircuit breaker pattern ↗The original write-up of the pattern agent runaway control borrows.
- 05 · Chaos Engineering communityPrinciples of Chaos Engineering ↗Why deliberately breaking a system is the only way to know its failure behaviour.
- 06 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
- 07 · Object Management GroupBPMN 2.0 specification ↗The modelling standard business process orchestration vocabulary comes from.
- 08 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). Resumable AI Workflows: Recovering From Failure at Step Four of Seven. Real Biz Digital. https://realbizdigital.net/insights/resumable-ai-workflows/
Try the mechanics on a live server
To see a governed tool surface respond before you point an agent at your accounting system — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
BarzelOps is this execution layer, sold as a running product
Forty tools covering capability discovery, workflow planning, preview before execution, run, trace, pause, resume and cancel, approval request and resolution, credential validation and connector health — across accounting, CRM, email, calendar, documents, Slack and storage. Opinionated workflows ship with it: customer onboarding, invoice follow-up, lead-to-invoice, pipeline cleanup, monthly close preparation and weekly operations briefings. The Free tier runs 100 calls a day.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Free | Free | 100 calls/day · 10 core tools: plan, preview, run, trace | Proving one workflow end to end before anyone signs anything |
| Pro | $19/mo | 15,000 calls/mo · 26 tools including approvals and templates | One operator automating their own recurring procedures |
| Team | $49/mo | 50,000 calls/mo · the complete 40-tool surface | An operations team running cross-app workflows under approval |
| Enterprise | $199/mo | Unlimited calls · deployment and connector scope by agreement | Multi-team execution with audit and residency obligations |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
The five Barzel servers, and which problem each one is sold for
One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.
| Server | Sold for | Entry price | Where it sits |
|---|---|---|---|
| Barzel Central Gateway | Knowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidence | Free, then $10–$149/mo | Control plane — decides what may be reached, and by whom |
| BarzelVault | Stopping a specific dangerous action before it executes, with proof afterwards | $199–$3,999/mo | Decision point — evaluates the individual call before execution |
| BarzelOps | Running real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approval | Free, then $19–$199/mo | Execution layer — does the work the policy allowed |
| Barzel FinOps Atlas | Attributing AI spend to agents, tools and outcomes, then forecasting and capping it | Free, then $29–$799/mo | Economics layer — what the estate costs per outcome |
| Barzel Scripture Intelligence | A free, credential-free public MCP server to test clients and inspect real protocol traffic | Free, unmetered, no signup | Reference implementation — safe place to learn the protocol |
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.