5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

AI Business Operations · Recovery

Resumable AI Workflows: Recovering From Failure at Step Four of Seven

Every multi-step workflow will fail partway through. Whether that leaves a customer half-onboarded forever, or a run that continues on Monday, is decided by design you either did or did not do.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202616 min read4,136 words

The short answer

A resumable AI workflow persists enough state at each step boundary that execution can continue from the last completed step after a failure, a pause or an approval wait — without repeating completed side effects and without a human reconstructing what happened. It requires seven state elements, per-step idempotency, and a declared answer for steps that cannot safely be retried. The alternative is not slower recovery. It is a person spending an afternoon working out which of five systems already received a write, which is the failure that destroys trust in an automation programme permanently.

Summary for readers and answer engines

Reviewed 25 Aug 2026

  • ▸Resumability is not an optimisation. Any workflow with an approval gate is paused by design and therefore already requires it.
  • ▸Seven state elements are needed, and the one teams forget is in-flight uncertainty: calls dispatched whose outcome is unknown.
  • ▸Resume must not repeat completed side effects, which makes per-step idempotency a prerequisite rather than a refinement.
  • ▸Three strategies exist: resume forward, compensate and abandon, or escalate to a human with the state. Every workflow needs a declared choice per step.
  • ▸Some steps are genuinely non-resumable. Naming them in advance is the design; discovering them during an incident is the failure.

Source: Mark Alex, Real Biz Digital — Resumable AI Workflows: Recovering From Failure at Step Four of Seven (https://realbizdigital.net/insights/resumable-ai-workflows/). Reproduce with attribution.

Key takeaways

  1. 01Persist state at step boundaries, not continuously. Boundaries are where resumption is safe and where the state is coherent.
  2. 02Record dispatched-but-unresolved calls explicitly. Treating unknown as failed causes duplicates; treating it as success hides them.
  3. 03Make every consequential step idempotent with an intent-derived key before you enable resume. Resume without idempotency manufactures duplicates.
  4. 04Declare a recovery strategy per step at design time. Improvising during an incident produces inconsistent state nobody can reconcile.
  5. 05Test resume by killing runs deliberately at each step boundary. A resume path that has never been exercised is a paragraph.
  6. 06Measure resume success rate. Below 95% means recovery is theatre and the state model is incomplete.
Part of the clusterAI Workflow Automation →

Quick answers

One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.

What is a resumable workflow?
One that persists enough state at each step boundary to continue from the last completed step after a failure, pause or approval wait, without repeating completed side effects.
Why is it mandatory rather than optional?
Because any workflow containing a human approval gate is paused by design. If it cannot resume, the approval either blocks indefinitely or the work is lost.
What state is needed?
Seven elements: completed steps and their outputs, the resolved entity references, the plan and current position, in-flight uncertainty, idempotency keys issued, approval state, and the definition version.
What do teams most often omit?
In-flight uncertainty — calls dispatched whose outcome is unknown. Treating them as failed produces duplicates; treating them as succeeded hides real gaps.
Can you resume without idempotency?
Not safely. Resume replays the boundary, and without an idempotency key the replayed step performs its effect a second time.
What if a step cannot be resumed?
Declare it in advance and choose a strategy: compensate the completed work and abandon, or escalate to a human with the full state. Both are legitimate; improvising is not.
How do you test resumption?
Kill runs deliberately at every step boundary and assert that resume completes correctly with no duplicated effects. Untested resume paths do not work.

Resumability in numbers

Every figure below is defined and sourced further down. They are stated here so they can be quoted without reading the whole page.

7state elements resumption requires
3recovery strategies
1state element teams always forget: in-flight uncertainty
above 95%target resume success rate
0completed side effects a resume may repeat
Anystep behind an approval gate makes resumability mandatory

The seven state elements

Resumption is only as good as the state persisted at the last boundary. These seven are what a resume actually needs.

Key facts

  • ▸In-flight uncertainty is the element that separates a working resume from a duplicating one, and it is absent from most implementations.
  • ▸Resolved entity references must be persisted rather than re-derived. Re-resolution on resume can legitimately produce a different answer, which is a subtle and expensive bug.
  • ▸Definition version matters more than it looks: a run that started on version 4 and resumes on version 5 is executing a process nobody approved.
Workflow resumption state elements
ElementContentsWithout it
Completed steps and outputsWhich steps finished and what each producedResume repeats work or skips a dependency
Resolved entity referencesCanonical keys for every entity in playResume re-resolves and may resolve differently
Plan and current positionThe planned sequence and where execution stoppedResume cannot know what comes next
In-flight uncertaintyCalls dispatched whose outcome is unknownResume duplicates or silently skips
Idempotency keys issuedThe key used for each consequential stepResume issues a new key and repeats the effect
Approval stateWhich approvals were requested, granted, expiredResume re-requests or proceeds unapproved
Definition versionThe workflow and capability-set version that started the runResume runs a different process from the one that began

Persist at step boundaries rather than continuously. Mid-step state is incoherent by definition, and a resume from an incoherent point is worse than a restart.

Three recovery strategies

Strategy 01

Resume forward

Continue from the last completed step. Requires all seven state elements and per-step idempotency. The preferred outcome, and available for the majority of failures — transient upstream errors, permission problems once fixed, approval waits.

Precondition: every remaining step is idempotent or has not yet run. Verify, do not assume.

Strategy 02

Compensate and abandon

Reverse or offset the completed steps and end the run cleanly. Correct when the failure invalidates the premise — the customer turned out to be a duplicate, the deal fell through, the data was wrong.

Precondition: a declared compensation for every completed step. Improvised compensation is how partial state becomes inconsistent state.

Strategy 03

Escalate with state

Stop, hand the full state to a named human, and let them decide. Correct when the recovery path depends on judgement, or when a step is genuinely non-resumable.

Precondition: the escalation must include what completed, what is uncertain, and what the options are. An escalation without those is a shrug with a ticket number.

Choosing a strategy
Failure situationStrategyWhy
Transient upstream errorResume forwardThe premise is intact; the call simply needs retrying
Expired credential, since fixedResume forwardNothing is wrong with the completed work
Approval granted after a waitResume forwardThis is the designed pause, not a failure
Premise invalidated (wrong entity, cancelled deal)Compensate and abandonContinuing would compound an error
Non-resumable step failedEscalate with stateOnly a human can decide the next move
In-flight outcome unknown on a paymentEscalate with stateNeither resuming nor abandoning is safe without reconciliation

Our verdict

Declare the strategy per step at design time, not per incident. The last row is the one that catches people: an unknown outcome on a consequential call is the situation where both automatic options are wrong, and the only safe action is a human with a reconciliation query.

Most failures are the first three rows, which is the argument for building resume-forward first. The last three are rarer and more expensive, which is the argument for declaring them in advance.

Idempotency at the resume boundary

Resume replays a boundary, and a replayed step without an idempotency key performs its effect twice. This is the most common way a recovery feature creates the incident it was built to prevent.

  • 01Issue the idempotency key before dispatching, and persist it with the state. A key generated at execution time is lost with the run; a persisted key survives to the resume.
  • 02Derive the key from intent, never from the attempt. The resumed attempt must produce the same key as the original, which rules out timestamps, attempt counters and anything regenerated.
  • 03Verify the upstream honours the key. Send the same key twice in a test and compare. Partial or silently ignored support is common, and discovering it during a resume is expensive.
  • 04Where no key is supported, add a pre-write existence check. Racy in theory, effective at resume timescales, and far better than nothing.
  • 05Treat non-idempotent steps as resume boundaries in themselves. A run should not be resumable across a step that cannot be safely replayed — it must escalate instead.
  • 06Never rely on the agent to remember. Idempotency is a property of the persisted state and the tool contract, not of the model’s context.

The sequencing rule that follows: build idempotency before you build resume. Enabling resume on a workflow whose steps are not idempotent converts every transient failure into a duplicate.

Steps that cannot be resumed

The left column can be replayed; the right column cannot be replayed without either a duplicate or a gap. The design task is to identify which steps sit on the right before production, and to declare what happens when a run stops at one.

Two options exist for each: make it resumable by adding idempotency, or make it an escalation boundary. Both are acceptable. What is not acceptable is a run that resumes across it and hopes.

Resumable

  • ✓Reads and lookups of any kind
  • ✓Idempotent writes with a key
  • ✓Set-to-value updates
  • ✓Create-if-not-exists with a deterministic name
  • ✓Anything not yet started

Not safely resumable

  • —A dispatched call whose outcome is unknown
  • —Sends without an idempotency key
  • —Append-only ledger writes
  • —Externally visible actions already in flight
  • —Provisioning without a natural key

A useful design constraint: place non-resumable steps as late in the workflow as possible, so that a failure before them leaves only reversible work behind. That is the same ordering rule that governs approval gates, and for the same reason.

The resume-safety test

Key facts

  • ▸Test two is the one that finds missing in-flight state, and it is the hardest to arrange. Worth the effort — it is exactly the case production produces.
  • ▸Test five is the acceptance criterion. Everything else is mechanism; duplicates are the outcome anyone will actually notice.
  • ▸Test six matters wherever resume can be triggered both automatically and by a human, which is most estates.
Six tests before enabling resume in production
Test 1

Kill at every boundary

Terminate a run after each step in turn, resume, and assert the final state matches an uninterrupted run. Every boundary, not a sample.

Test 2

Kill mid-dispatch

Terminate between dispatching a call and receiving its result, so the state records in-flight uncertainty. This is the test that finds the missing state element.

Test 3

Resume after a long delay

Resume a run paused for longer than the approval expiry and assert it re-requests rather than proceeding on a stale approval.

Test 4

Resume on a changed definition

Attempt to resume a run whose workflow version has since changed, and assert it either continues on the original version or refuses explicitly.

Test 5

Assert no duplicate effects

After every resume test, query each affected system for duplicates. This is the assertion that matters and it is the one most often omitted.

Test 6

Resume twice

Resume the same run twice concurrently and assert only one continuation proceeds. Rare in practice, catastrophic when it happens.

Run these as a scheduled suite rather than once. Resume paths break silently when upstream tools change their idempotency behaviour, and nothing else will tell you.

Why approval gates make this mandatory

Resumability is often framed as failure handling. In practice the commonest reason a workflow needs to resume is entirely by design.

  • 01A workflow with a human gate is paused every single run. The resume path is therefore exercised constantly, not exceptionally.
  • 02That is good news: the mechanism gets tested by ordinary operation rather than only by incidents, so defects surface early.
  • 03It also means the state model must survive working hours, weekends and absence. A state store designed for a thirty-second retry will not do.
  • 04Approval expiry interacts with resume directly. A run resuming after an approval expired must re-request, not proceed — and that logic belongs in the resume path.
  • 05The pause is a first-class state, not an error. Modelling it as a failure produces alarming dashboards and trains people to ignore them.
  • 06Because pause and failure share the state model, building for approval waits gets you failure recovery nearly free. Sequence it that way.

Approval waits as the primary resume case

Verified 26 Aug 2026

Hours to daystypical approval wait duration
100%of gated workflows that pause by design
1state model serving both pause and failure
0difference between resuming a pause and resuming a failure

This reframing is the most useful thing in this article for most teams: you are not building an exotic recovery feature, you are building the mechanism your approval gates already require.

Measuring recovery

Recovery metrics
MetricDefinitionTarget
Resume success rateResumes completing correctly / resumes attemptedAbove 95%
Duplicate effects per resumeDuplicated side effects detected after resumeZero
Median pause durationTime between suspension and resumptionUnder the approval expiry
Escalation rateRuns escalated rather than resumed or compensatedFalling as state coverage improves
Compensation success rateCompensations completing without residueAbove 95%
Stranded runsRuns neither completed, compensated nor escalatedZero — this is the metric that matters most
Resume test suite pass rateScheduled resume tests passing100%

Stranded runs is the number to put on the dashboard. A run that is neither finished nor cleaned up nor escalated is precisely the half-onboarded customer this whole article exists to prevent, and it should never be above zero.

Next step

Pause, resume and cancel as first-class operations

BarzelOps holds workflow state across pauses and approval waits, with pause, resume and cancel as callable operations and an action-level trace of what completed before the interruption.

Limits

Two.

  • 01Resumption cannot resolve an unknown upstream outcome. It can prevent a second attempt from committing a second effect, and it cannot tell you whether the first one committed — which is why reconciliation queries per capability matter more than state design.
  • 02It cannot recover from a state store failure. The state is now the critical dependency, and its availability and durability requirements are those of the workflows depending on it.

Common misconceptions

Four claims we hear regularly that do not survive contact with a real estate. Each is stated as we hear it, then corrected.

Myth

Resumability is an advanced feature to add once the basics work.

Actually

Any workflow with a human approval gate pauses by design on every run, so the resume path is already required and already exercised constantly. Treating it as advanced means shipping a workflow whose most common state transition has never been built.

Myth

Retrying a failed step is the same as resuming a workflow.

Actually

Retrying repeats one call; resuming continues a partially completed multi-step process without repeating what already succeeded. Retry needs a backoff policy; resume needs seven persisted state elements and per-step idempotency.

Myth

If a step failed, nothing happened.

Actually

A dispatched call whose outcome is unknown may well have committed upstream. Treating unknown as failed is how resumes create duplicate payments, and treating it as succeeded is how real gaps go unnoticed — which is why in-flight uncertainty must be a distinct recorded state.

Myth

A resumed run can safely pick up the current workflow definition.

Actually

It cannot. A run that started on one version and resumes on another is executing a process no one reviewed as a whole, with dependencies that may no longer hold. Resume must continue on the original version or refuse explicitly.

Frequently asked questions

What makes an AI workflow resumable?

Persisting enough state at each step boundary to continue from the last completed step after a failure, pause or approval wait, without repeating completed side effects or requiring a human to reconstruct what happened.

Why is resumability mandatory rather than optional?

Because any workflow containing a human approval gate pauses by design on every run. Without a resume path the approval either blocks indefinitely or the completed work is lost, which means the mechanism is required from the first gated workflow.

What state must be persisted to resume safely?

Seven elements: completed steps and their outputs, resolved entity references, the plan and current position, in-flight uncertainty, idempotency keys already issued, approval state, and the workflow definition version the run started on.

What is in-flight uncertainty and why does it matter?

Calls that were dispatched but whose outcome is unknown. It is the state element most often omitted, and its absence is what causes resumes to either duplicate effects by retrying or silently skip work by assuming success.

Can a workflow resume without idempotency?

Not safely. Resume replays a step boundary, and a replayed consequential step without an idempotency key performs its effect a second time — which is how a recovery feature creates the incident it was built to prevent.

Where should idempotency keys be generated?

Before the call is dispatched, and persisted with the workflow state. A key generated at execution time is lost with the run, so the resumed attempt would issue a new one and repeat the effect.

What are the three recovery strategies?

Resume forward from the last completed step; compensate the completed work and abandon the run cleanly; or escalate to a named human with the full state. Each step should have a declared strategy chosen at design time.

Which steps cannot be safely resumed?

Dispatched calls with unknown outcomes, sends without an idempotency key, append-only ledger writes, externally visible actions already in flight, and provisioning without a natural key. Each must either be made idempotent or declared an escalation boundary.

Where should non-resumable steps sit in a workflow?

As late as possible, so that a failure before them leaves only reversible work behind. This is the same ordering rule that governs approval gate placement and it exists for the same reason.

How should resumption be tested?

Six tests: kill at every step boundary, kill mid-dispatch so in-flight state is exercised, resume after an approval expiry, attempt resume on a changed definition, assert zero duplicate effects in every affected system, and attempt two concurrent resumes.

Should a resumed run use the current workflow definition?

No. It must continue on the version it started with, or refuse explicitly. A run spanning two definition versions executes a process nobody reviewed as a whole, with dependencies that may no longer hold.

What is the single most important recovery metric?

Stranded runs — runs that are neither completed, compensated nor escalated. It should never exceed zero, because a stranded run is precisely the half-onboarded customer that destroys trust in an automation programme.

Glossary

Resumable workflow
A workflow able to continue from its last completed step after failure, pause or approval wait.
Step boundary
The point between steps at which state is coherent and resumption is safe.
In-flight uncertainty
Dispatched calls whose outcome is unknown, recorded as a distinct state.
Resume forward
Continuing from the last completed step without repeating completed effects.
Compensate and abandon
Reversing completed steps and ending a run cleanly when the premise is invalidated.
Escalation boundary
A step across which a run may not resume automatically, requiring human judgement.
Definition version pinning
Continuing a resumed run on the workflow version it started with.
Stranded run
A run neither completed, compensated nor escalated, leaving inconsistent state.
Resume success rate
The share of attempted resumes that complete correctly with no duplicated effects.
Pre-write existence check
A lookup before a write, used where the upstream supports no idempotency key.

Standards and entities referenced

Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.

  1. 01 · microservices.ioSaga pattern ↗Compensating transactions, which is all you get when distributed rollback does not exist.
  2. 02 · WikipediaIdempotence ↗Why safe retries require this property rather than hope.
  3. 03 · Michael NygardRelease It! — stability patterns ↗Bulkheads, circuit breakers and timeouts, from the book that named them.
  4. 04 · Martin FowlerCircuit breaker pattern ↗The original write-up of the pattern agent runaway control borrows.
  5. 05 · Chaos Engineering communityPrinciples of Chaos Engineering ↗Why deliberately breaking a system is the only way to know its failure behaviour.
  6. 06 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
  7. 07 · Object Management GroupBPMN 2.0 specification ↗The modelling standard business process orchestration vocabulary comes from.
  8. 08 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.

Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). Resumable AI Workflows: Recovering From Failure at Step Four of Seven. Real Biz Digital. https://realbizdigital.net/insights/resumable-ai-workflows/

Try the mechanics on a live server

To see a governed tool surface respond before you point an agent at your accounting system — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

BarzelOps is this execution layer, sold as a running product

Forty tools covering capability discovery, workflow planning, preview before execution, run, trace, pause, resume and cancel, approval request and resolution, credential validation and connector health — across accounting, CRM, email, calendar, documents, Slack and storage. Opinionated workflows ship with it: customer onboarding, invoice follow-up, lead-to-invoice, pipeline cleanup, monthly close preparation and weekly operations briefings. The Free tier runs 100 calls a day.

PlanPriceIncludedRight for
FreeFree100 calls/day · 10 core tools: plan, preview, run, traceProving one workflow end to end before anyone signs anything
Pro$19/mo15,000 calls/mo · 26 tools including approvals and templatesOne operator automating their own recurring procedures
Team$49/mo50,000 calls/mo · the complete 40-tool surfaceAn operations team running cross-app workflows under approval
Enterprise$199/moUnlimited calls · deployment and connector scope by agreementMulti-team execution with audit and residency obligations

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

The five Barzel servers, and which problem each one is sold for

One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.

ServerSold forEntry priceWhere it sits
Barzel Central GatewayKnowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidenceFree, then $10–$149/moControl plane — decides what may be reached, and by whom
BarzelVaultStopping a specific dangerous action before it executes, with proof afterwards$199–$3,999/moDecision point — evaluates the individual call before execution
BarzelOpsRunning real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approvalFree, then $19–$199/moExecution layer — does the work the policy allowed
Barzel FinOps AtlasAttributing AI spend to agents, tools and outcomes, then forecasting and capping itFree, then $29–$799/moEconomics layer — what the estate costs per outcome
Barzel Scripture IntelligenceA free, credential-free public MCP server to test clients and inspect real protocol trafficFree, unmetered, no signupReference implementation — safe place to learn the protocol

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.