5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

Workflow automation

Approval checkpoints: what happens to the workflow while it waits

Human approval is a commodity feature. Every serious automation platform has an approval step, and anyone selling you one as a differentiator is selling you a checkbox. The engineering question is what happens to the execution during the wait: whether the pause holds an in-memory context that a deploy destroys, or a checkpoint written to durable storage that rehydrates afterwards. That distinction decides how long you can wait, whether a restart loses the run, and whether the resume posts your invoice twice.

How Barzel applies here Start free with BarzelOps

Why is “we have approval checkpoints” not a differentiator?

Because everyone has them. n8n ships send-and-wait operations across Slack, Teams, Gmail and Outlook. Zapier’s human-in-the-loop step is generally available, with approve, decline and change-data responses, timeouts and reminders. Power Automate has a Start and wait for an approval action. Camunda has BPMN user tasks, Temporal has signals, Step Functions has the task-token pattern. Table stakes.

The uncomfortable proof is Relay.app. It had the richest human-in-the-loop model in the mainstream market — four step types, dynamic role assignment, interactive Slack approvals, per-step escalation — and it is being switched off at 23:59 PT on 14 September 2026. Depth at the approval step did not save it, and will not distinguish anyone else.

So this is not about approval features. It is about the hours or days between the request and the click — the part nobody demos, and where the failures live. Whether you can take the workflow elsewhere afterwards is a separate question.

What happens to a workflow while it waits for approval?

One of two things, and the builder canvas looks identical either way.

Pause-and-wait keeps the execution alive: the run holds its variables in process memory, occupies a worker slot, or sits inside a suspended function invocation. The wait is a property of a running process, so anything that ends the process ends the wait — a deploy, an autoscaler scaling in, an out-of-memory kill, a node replacement.

Checkpoint-and-resume ends the execution deliberately. The engine serialises where it had reached — position in the graph, variable state, a correlation identifier — writes that to durable storage, and releases the worker. Nothing is running. When the approval arrives carrying that identifier, the engine reads the record back and reconstitutes a materially identical run. The wait is a row.

n8n publishes the clearest example of both in one product. Its Wait node documentation draws the line at 65 seconds: shorter waits are held in memory, and beyond that n8n saves the execution to the database and restores it when the wait ends. Same node, two different engineering artefacts — and the boundary is a number in someone else’s documentation, not anything visible on your canvas.

Why the distinction bites specifically on approvals

Because approval windows are long. The finance lead approving a batch of invoices is at lunch, on a plane, or on leave until Tuesday. Realistic windows run from hours to days, with a longer tail than anyone plans for.

That length crosses infrastructure events. If you deploy weekly and approvals take three days, some proportion of your waits will span a deploy — not as an edge case but as arithmetic. Add vendor incidents, which arrive on their own schedule, and the conclusion is unavoidable: anything that must survive a deploy has to be durable state, not a suspended process. No configuration makes an in-memory pause outlive its process.

Pause-and-wait versus checkpoint-and-resume

PropertyPause-and-waitCheckpoint-and-resume
Survives a deploy or restart No. The execution is the process. Yes. The execution is a row that outlives every process.
Survives a vendor incident or worker eviction No, and you find out from an absent result rather than an alert. Yes, if storage survived. The wait is not what is at risk.
Practical maximum wait Until the next restart — unpredictable, and shorter than your deploy interval. Documented run-duration and retention limits. Step Functions Standard Workflows document one year for the task-token pattern.
Duplicate risk on resume High. No stable identity survives the failure, so a retry is a new request downstream. Bounded — but only if you did the work. Durability does not give you idempotency.
What the approver saw Re-rendered live at click time, unrecoverable afterwards. A snapshot in the checkpoint, hashable and citable months later.
Evidence produced A log line saying the run was approved. Payload hash, approver, timestamp, decision and dispatch outcome, linked.

The fourth row is the one to re-read. Teams moving to a durable engine assume persistence implies exactly-once behaviour. It does not, and the vendors say so: Temporal documents activity executions as at-least-once, Camunda’s Zeebe engine documents at-least-once job handling. Durable state guarantees the run comes back. It guarantees nothing about how many times the thing at the end of it happens.

Why do approvals cause duplicate invoices?

The failure in one sentence: at-least-once execution plus a non-idempotent downstream endpoint equals duplicates. Both halves are ordinary. Creating an invoice, a payment or a purchase order is an HTTP POST, and RFC 9110 is explicit that POST is not idempotent — two identical requests are two invoices, by design, because sometimes you do want to bill twice.

The dangerous window is narrow and unavoidable. Your workflow dispatches the request; the downstream accepts and commits; your process dies before the response is recorded. The checkpoint now says the step was attempted, not whether it succeeded, so the resume cannot distinguish not sent from sent, response lost. Something has to break the tie, and if nothing does, the safe-looking default — retry — double-posts. What that costs downstream is a separate conversation with finance.

The key must be assigned before the first attempt

This is the most common implementation error in the area, and it survives code review because it reads correctly.

The wrong version generates the idempotency key inside the dispatch function, so every attempt mints a fresh one. The retry presents a different key, the downstream sees a request it has never encountered, and creates a second record. The mechanism is present, configured, and doing precisely nothing.

The right version derives the key deterministically from the run and step identifiers, writes it into the checkpoint before the first attempt, and never regenerates it. Any resume — first retry, third retry, a rehydration two days after a crash — presents the identical value. The key belongs to the workflow’s durable state. In the retry path it is decoration.

There are two clocks, and only one of them is yours

A correctly assigned key still has a shelf life, and the retention window belongs to the receiving system. Stripe documents that idempotency keys are removed automatically after 24 hours — sensible for ordinary API retries, short for a workflow whose approval queue runs to Tuesday.

The clocks differ. Your approval window runs from checkpoint to click. The provider’s deduplication window runs from the first request it actually sees, after the approval. Usually that keeps you inside it — until the downstream has a bad day: approval granted, first attempt fails, outage runs thirty hours, resume presents a key nobody remembers.

A provider key is therefore necessary and not sufficient. The complement is your own record: store the downstream identifier from the response before marking the step complete, and on any retry outside the deduplication window, query the downstream for an existing artefact carrying your reference before writing a new one. Where neither holds, the remaining honest position is to write anyway and reconcile — building the detection and reversal path deliberately, rather than discovering it by accident. That is a legitimate choice; believing you are protected when you are actually reconciling is not.

What did the approver actually see?

An approval interface that renders live data at click time is not evidence of what was approved.

The scenario is mundane. The workflow pauses on a payment of £4,200. Somebody amends the underlying record. The approver opens the request next morning, sees £6,900 re-queried at render, and clicks approve. The audit record says the payment was approved, which is true, and implies the approved figure was £4,200, which is false. Nothing is lying; the record never captured what mattered.

The fix is cheap and belongs in the checkpoint. Snapshot the payload, hash it, render the approval from the snapshot rather than a live query, and store the hash with the approver identity, timestamp and decision. Then re-verify at dispatch: if the payload no longer hashes to what was approved, that is a new decision rather than a resumption, and should fail closed or re-request.

This is the reconstruction standard from building an audit trail: not whether a record exists, but whether someone can rebuild the decision without a developer. It is also where the engineering acquires a compliance consequence. UK GDPR Article 22C, in force since 5 February 2026, gives a person the right to obtain human intervention on a significant automated decision — and that right presupposes an intervention you can evidence.

What else breaks at resume?

Four things a same-session test never surfaces, because it resumes in seconds.

  • Credentials expire during the wait. An OAuth token acquired before a three-day approval will not be valid after it, and the resume fails at the last step rather than the first. Multi-tenant setups make this worse — see credentials as the real boundary.
  • The upstream record moved. The deal closed-lost, the invoice was voided, the customer was merged. The checkpoint is a snapshot of the world at pause time, and the world did not hold still.
  • The workflow definition changed. A run checkpointed on version 12 resuming onto version 15 is undefined behaviour unless the engine versions in-flight executions explicitly. Temporal ships versioning APIs for this; without an equivalent, every deploy gambles with whatever is paused.
  • Nobody approves at all. Expiry has three plausible behaviours — fail closed, escalate, or proceed unapproved — and if you have not chosen one, a vendor default is making a business rule.

How to design an approval checkpoint that survives a restart

  1. Declare what the checkpoint contains first: run and step identifiers, dispatch payload, upstream record versions, approver, expiry behaviour. Anything off that list will not survive the wait.
  2. Persist the checkpoint before you ask anyone. The other order leaves a window where an approver can act on a pause the system has no record of.
  3. Assign the idempotency key at checkpoint time, derived from run and step identity, stored in the record, never regenerated.
  4. Snapshot what the approver will see and hash it. Render from the snapshot, not a live query.
  5. Make the wait a correlation, not a held process — an inbound signal carrying an identifier matched against storage. A run that owns a connection or a thread has tied its approval window to process uptime.
  6. Re-validate at resume: payload hash still matches, credentials still valid, upstream record not moved. A mismatch is a new decision.
  7. Dispatch through an idempotent envelope. Send the stored key, record the downstream identifier before marking the step complete, and check for an existing artefact before any late retry.
  8. Test by restarting during the wait. Pause a run, redeploy, approve. One correct completion and exactly one downstream artefact means the checkpoint is real. Anything else is pause-and-wait wearing a durable-sounding name.

Step eight is the only one that produces an answer rather than an intention, and the one that gets skipped. It takes ten minutes.

Where we sit

In practice

BarzelOps is built around this shape of problem: cross-system workflow automation with durable state on PostgreSQL, crash recovery and idempotency, human approval checkpoints with resumable execution, tenant isolation with signed evidence receipts, and a portable Governed Workflow Manifest intended to run in cloud, VPC, on-premises or air-gapped deployments. To be explicit about that term: Governed Workflow Manifest is Real Biz Digital’s own name for a portable definition format, and it is not an industry standard — no such standard exists, which is the point. The approval step is commodity here as everywhere else; the durable wait behind it is what to test.

How Barzel applies here

Frequently asked questions

What happens to a workflow while it waits for approval?

Either it holds context in process memory, so a restart destroys the wait, or the engine writes state to durable storage and reconstitutes the run on approval. Only the second survives a wait measured in days.

Why do approvals cause duplicate invoices?

At-least-once execution plus a non-idempotent endpoint. POST is not idempotent under RFC 9110, and a crash between dispatch and recording the response leaves the resume unable to tell not sent from sent-but-unrecorded. Without a stable key, it re-sends.

When should an idempotency key be generated?

At checkpoint time, before the first attempt, derived from run and step identity. Never inside the retry path — that mints a fresh key per attempt and gives no protection at all.

Is an approval step the same as an audit trail?

No. An approval step records that a decision was made; an audit trail records what it was made about. Data rendered live at click time leaves no trace of what changed.

How long can a workflow wait for approval?

An in-memory pause lasts until the next restart. A durable checkpoint is bounded by documented run-duration and retention limits — Step Functions Standard Workflows document one year for the task-token pattern. Look up yours.

What should happen if nobody approves?

Fail closed, escalate, or proceed unapproved. Only the first two are defensible for anything that moves money, and if you have not configured it, a vendor default is deciding.

Where this leaves you

Approval checkpoints are not a feature to shop for; they are a property to test. Two questions separate a demo from a system: does a run in the approval state survive a restart, and does a resume after a crash produce one downstream record or two. Both are answerable in an afternoon, before you commit — and neither appears on a comparison page, because the honest answer is usually unflattering.

In practice

Automation you can leave, with the approvals and the evidence intact.

A workflow platform is only as safe as your ability to leave it. BarzelOps runs cross-system workflows with durable state, human checkpoints and signed evidence, exports them as a portable manifest, and migrates Relay workflows with a dry run before cutover.

BarzelOps

Governed workflow automation across the systems that run the business.

  • Durable, idempotent execution: a timeout is retried once, never filed twice.
  • Human approval checkpoints that pause the workflow and resume it.
  • Isolation per entity or client, signed evidence receipts and a portable manifest; HubSpot, Xero, Gmail, Google Drive and Slack.

Free tier: 100 calls a dayPaid plans from $19 a monthLive on MCPize

Start free Ask by emailProduct pageDocumentation

BarzelVault

The AI action firewall: decide what an agent may do before it does it.

  • Approval thresholds and policy checks enforced before execution; human approvals that expire and escalate.
  • Cryptographically signed audit receipts: trigger, inputs, policy version, approver, outcome.
  • Credential isolation, spend and action limits, and an emergency kill switch.

Free tier: 10,000 calls a monthPaid plans from $199 a monthLive on MCPize

Start free Ask by emailProduct pageDocumentation

Enterprise: written quote by email within two business days. No sales call.


Related

Sources

  1. IETF, RFC 9110: HTTP Semantics, section on idempotent methods. The Idempotency-Key header itself remains an IETF internet-draft rather than a ratified standard.
  2. n8n documentation, Wait node — in-memory versus database-backed waits and the 65-second threshold; Human in the loop send-and-wait operations.
  3. Temporal documentation — durable execution, signals, at-least-once activity execution, and workflow versioning for in-flight executions.
  4. Camunda documentation — BPMN user tasks and at-least-once job handling in the Zeebe engine.
  5. AWS, Step Functions documentation — the waitForTaskToken service integration pattern and Standard Workflow duration limits.
  6. Stripe API documentation, Idempotent requests — key retention window.
  7. Zapier, Request approval to keep your workflow running with Human in the Loop; Microsoft, Start and wait for an approval action reference.
  8. relay.app and docs.relay.app — shutdown notice and human-in-the-loop step documentation, as published.
  9. Data (Use and Access) Act 2025, section 80, inserting UK GDPR Article 22C, in force 5 February 2026.

This article is for information and does not constitute legal, engineering or procurement advice. Platform limits, retention windows and product behaviour are as documented at 2 September 2026 and change without notice; verify each against the vendor’s own current documentation before relying on it in a design.