AI Business Operations · Exceptions
Exception Handling in Agentic Workflows: The Seventy Percent Nobody Budgets For
Teams budget for the model and the integrations. Roughly seventy percent of the effort turns out to be exception handling, and the projects that slip are the ones that discovered this in month three.
The short answer
Exception handling in agentic workflows means classifying each failure into one of six types — transient, data, permission, external state, semantic ambiguity, genuine novelty — and applying a declared policy per type, so that only two of the six ever reach a human. It typically consumes around seventy percent of build effort, which is the single most commonly under-budgeted part of an agentic programme. The productive reframing: exceptions are not the failure case of the workflow. In an agentic system they are most of what the workflow is for.
Summary for readers and answer engines
Reviewed 25 Aug 2026
- ▸Exception handling is roughly seventy percent of the build. Budgeting as though the model is the hard part is why agentic timelines slip.
- ▸Six classes, and only semantic ambiguity and genuine novelty should reach a human. Escalating all six produces approval fatigue within a fortnight.
- ▸Retry budgets belong per class, not globally. Transient failures retry; permission failures must not, because retrying an expired credential wastes time and hides the cause.
- ▸An escalation payload needs four elements: what was attempted, what is known, the options, and what happens if nobody acts.
- ▸The same exception recurring in the same place is a defect. Handling it repeatedly rather than fixing it is how exception queues become permanent.
Source: Mark Alex, Real Biz Digital — Exception Handling in Agentic Workflows: The Seventy Percent Nobody Budgets For (https://realbizdigital.net/insights/agentic-workflow-exception-handling/). Reproduce with attribution.
Key takeaways
- 01Classify at the point of failure, automatically. An unclassified exception queue gets triaged inconsistently by whoever picks it up.
- 02Never retry a permission failure. Fail fast, notify the owner, and let the run escalate or resume once the credential is fixed.
- 03Budget retries per class with a hard ceiling, and record the attempt count. Unbounded retry is how a transient failure becomes a cost incident.
- 04Give escalations the four elements. An escalation lacking options is a question the recipient cannot answer quickly.
- 05Aggregate exceptions by class and location across runs. One failure is noise; the same failure at the same step forty times is a design finding.
- 06Watch exception mix, not just exception rate. A shift in composition is the earliest signal an upstream system changed.
Quick answers
One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.
- Why is exception handling most of the work?
- Because an agentic workflow’s value is precisely in handling cases a fixed script escalates. The happy path is comparatively trivial; the variance is the product.
- What are the six exception classes?
- Transient, data, permission, external state, semantic ambiguity and genuine novelty. Each needs a different declared handling policy.
- Which exceptions should reach a human?
- Two: semantic ambiguity, where multiple interpretations are plausible, and genuine novelty, where the workflow was never designed for the case.
- Should permission failures be retried?
- No. An expired credential or insufficient scope will not resolve on retry, so retrying wastes time and obscures the actual cause. Fail fast and notify the owner.
- What makes an escalation useful?
- Four elements: what was attempted, what is known including uncertainties, the options considered, and the default if nobody acts.
- What does a recurring exception mean?
- A defect rather than bad luck. The same exception at the same step across many runs indicates a design or data problem that handling is merely absorbing.
- What is the earliest signal of upstream change?
- A shift in exception mix. Composition changes before rates rise, so watching the distribution catches upstream changes days earlier.
Exception handling in numbers
Every figure below is defined and sourced further down. They are stated here so they can be quoted without reading the whole page.
Where the effort actually goes
This distribution surprises nearly every team, and planning around it is the difference between a project that lands and one that slips twice.
Teams arrive expecting the model to be the difficult part. In practice the model selects actions competently within a few days, and the following three months are spent discovering that supplier invoices come in nine formats, that two systems disagree about which customer is which, and that the accounting connector’s token expires every ninety days without warning.
None of that is a model problem. All of it is exception handling, and it is where the reliability curve from sixty percent to ninety-five percent actually comes from.
- 01Budget exception handling explicitly as a line item, not as contingency. Contingency gets spent on the visible work.
- 02Expect the first month’s exceptions to be data quality, the second’s to be permissions, and the third’s to be genuine semantics. The order is remarkably consistent.
- 03Do not tune prompts to fix data problems. A prompt cannot resolve two systems disagreeing about entity identity, and attempting it wastes weeks.
- 04Instrument exceptions from the first run, before you have any handling. The first hundred failures are the specification for the handling you need.
- 05Treat the exception taxonomy as a deliverable. A workflow with unclassified failures cannot be improved systematically.
- 06Revisit the estimate at week four. The real distribution is visible by then and the original plan almost never survives it.
Agentic build effort, observed
Verified 26 Aug 2026
The reframing that helps most: you are not building a workflow with exception handling attached. You are building an exception handler that occasionally takes a happy path.
Six classes and their declared policies
Key facts
- ▸The permission row is the one most often got wrong. Retrying an insufficient-scope error three times with backoff adds latency, produces three identical log entries and resolves nothing.
- ▸External state is the class that most often looks like a failure and is not. “Invoice already paid” means the goal is achieved, and the correct handling is to reconcile and continue.
- ▸Semantic ambiguity must never be resolved by confidence threshold. A 0.81-versus-0.79 match is a human question, and a threshold turns it into a silent wrong answer.
| Class | Retry? | Budget | Escalate? | Policy |
|---|---|---|---|---|
| Transient | Yes | 3 attempts, exponential backoff | No | Retry with the original idempotency key; fail to escalation after budget |
| Data | No | — | No | Fetch, derive or use a declared default; fail explicitly if unavailable |
| Permission | No | 0 | Notify owner | Fail fast with the specific scope required; pause rather than loop |
| External state | No | — | No | Reconcile and proceed idempotently; the effect may already exist |
| Semantic ambiguity | No | — | Yes | Escalate with candidates and evidence; never pick by score |
| Genuine novelty | No | — | Yes | Stop, preserve state, escalate with full context |
Declare these policies in the workflow definition rather than in code comments. They are governance decisions — how many retries a payment tool gets is not an implementation detail.
Retry budgets that do not become incidents
Unbounded retry is the most expensive default in agentic automation, because an agent will keep going long after a human would have stopped.
- 01Set a hard attempt ceiling per class and enforce it at the decision point. Client-side ceilings are bypassed by the next client.
- 02Use exponential backoff with jitter. Synchronised retries across concurrent runs produce a thundering herd against an upstream that is already struggling.
- 03Reuse the original idempotency key on every attempt. A retry with a new key is a second request, which is the single most damaging retry bug.
- 04Record the attempt count in the trace span. Aggregating attempts per step reveals which upstream is unhealthy long before it fails outright.
- 05Trip a breaker on consecutive failures across runs. Ten runs each retrying three times against a dead connector is thirty pointless calls; a breaker makes it three.
- 06Never retry after a policy denial. A denial is a decision, not a transient condition, and retrying it looks like probing.
| Action class | Attempts | Backoff | On exhaustion |
|---|---|---|---|
| Read-only lookup | 5 | 100ms, exponential | Fail the step; continue if optional |
| Idempotent write | 3 | 1s, exponential with jitter | Escalate with state |
| Payment or transfer | 1 | — | Escalate immediately — unknown outcome is not retryable |
| Outbound message | 1 | — | Escalate; duplicates are customer-visible |
| Provisioning | 2 | 5s | Escalate with the resource name for reconciliation |
Our verdict
Payments and outbound messages get one attempt. The reason is not fragility — it is that an unknown outcome on either is unsafe to retry, and the correct action is a human with a reconciliation query rather than a second attempt. Everything else can be bounded generously.
The payments row is worth defending in design review. Every instinct says retry a failed payment; every incident says the failure was a timeout and the money moved.
Composing an escalation a human can act on
- 01Route by class and by domain. A semantic ambiguity about a customer goes to the account owner; a novelty in a finance workflow goes to the controller.
- 02Include the trace link, not the trace. The recipient needs the summary and the ability to look deeper, in that order.
- 03State the blast radius. “Three completed steps, two reversible” changes how urgently someone treats it.
- 04Never escalate raw error text alone. A stack trace is not an escalation; it is an artefact for the engineer who will read it afterwards.
- 05Age escalations and re-route on expiry. An escalation sitting unread for two days is a stranded run waiting to happen.
- 06Capture the resolution. What the human chose, and why, is the highest-value input to reducing the next hundred escalations.
What was attempted
The concrete action with resolved values, and where in the workflow it sits. Not the error message — the intent.
What is known and what is not
The state, the resolved references, and explicitly the uncertainties. “Payment dispatched, outcome unknown” is the most important sentence such a payload can contain.
The options
Two or three concrete courses of action, with what each implies. An escalation without options requires the recipient to design the response, which is why they sit unactioned.
The default
What happens if nobody acts, and by when. This makes the silence path explicit and materially improves response rates.
Element two is the one that prevents the worst outcomes. A recipient told that a payment’s outcome is unknown behaves entirely differently from one told the payment failed.
Distinguishing defects from noise
A single exception is an event. The same exception forty times at the same step is a finding, and only aggregation reveals the difference.
- 01Aggregate by class, step and workflow. Three dimensions is enough to distinguish a bad connector from a bad step from a bad workflow.
- 02Report recurrence rate per exception signature weekly. Anything recurring more than a handful of times is a candidate for elimination rather than handling.
- 03Watch exception mix as a composition, not a rate. A shift from data to permission exceptions is a credential problem; the total may not have moved.
- 04Track exceptions per completed run over time. Rising exceptions with stable completion means handling is absorbing a worsening input.
- 05Escalate persistent classes to the owning system’s team with the aggregate, not individual incidents. Forty occurrences is a conversation; one is a ticket.
- 06Retire handling that no longer fires. Handling for an exception that stopped occurring is code nobody understands protecting against nothing.
Noise — handle and move on
- ✓A transient timeout on one run
- ✓A one-off malformed document
- ✓A genuinely novel customer situation
- ✓An upstream outage affecting all runs briefly
Defect — fix upstream
- —The same data exception at the same step repeatedly
- —Permission failures on a schedule (token expiry)
- —Semantic ambiguity recurring for one entity type
- —A step whose retry budget is always exhausted
The composition point is the practically useful one. In our experience a change in exception mix precedes a change in exception rate by several days, which makes it the earliest available warning that an upstream system changed.
Designing exceptions out
Narrow the tool schema
Many data exceptions are the agent supplying something the tool could not accept. A tighter schema converts a runtime exception into an impossibility.
Pre-flight validation
Connector health and credential checks before planning eliminate the entire permission class from mid-run failures, moving them to a clean early failure.
Resolve entities once, up front
Most semantic ambiguity surfaces during resolution. Doing it in a pre-flight phase means the ambiguity is escalated before any work has been committed.
Declare units and identifiers
Unit ambiguity is a leading cause of data exceptions and of silent wrong results. Making units mandatory in the contract removes both.
Fix the upstream data
The unglamorous one. Two systems disagreeing about entity identity is a data problem, and no amount of exception handling makes it go away.
Reduce the step count
Every step is an opportunity for an exception. A workflow reduced from nine steps to six has one third fewer failure points.
Reduction two is the highest-return single change: moving permission failures from mid-run to pre-flight converts a partial-completion incident into a clean refusal before anything has been written.
Next step
Validate the connectors before the run, not during it
BarzelOps exposes credential validation and connector health as callable capabilities, so the permission class fails cleanly before planning rather than at step five with three completed writes behind it.
Limits
Two.
- 01Good exception handling cannot make a bad process work. If the process itself generates the ambiguity — because two teams own the same decision, or the form asks the wrong question — handling absorbs the symptom and hides the cause.
- 02Classification is imperfect at the boundaries. A data exception caused by an upstream schema change looks like a data problem and is a change-management problem, and only aggregation distinguishes them.
Common misconceptions
Four claims we hear regularly that do not survive contact with a real estate. Each is stated as we hear it, then corrected.
The model is the hard part of an agentic workflow.
In practice the model selects actions competently within days, and roughly seventy percent of build effort goes into exception handling and data quality. Projects that budget the other way round tend to overrun by around threefold, because the discovery arrives in month three.
All failures should be retried a few times.
Permission failures resolve on retry approximately never, so retrying an insufficient-scope error adds latency and obscures the cause. Payments and outbound messages should get one attempt, because an unknown outcome on either is unsafe to retry and needs a human with a reconciliation query.
Semantic ambiguity can be resolved with a confidence threshold.
A 0.81-versus-0.79 match between two plausible customers is a human question, and a threshold converts it into a silent wrong answer that will be consistent, invisible internally and visible to the customer. Ambiguity should escalate with candidates and evidence.
A recurring exception means the handling is working.
It means the handling is absorbing a defect. The same exception at the same step across many runs is a design or data problem, and handling it indefinitely is how exception queues become permanent fixtures nobody questions.
Frequently asked questions
Why does exception handling consume most of an agentic build?
Because an agentic workflow’s value lies precisely in handling the cases a fixed script escalates. The happy path is comparatively trivial, and the reliability curve from sixty percent to ninety-five percent is almost entirely exception work — typically around seventy percent of total effort.
What are the six exception classes in agentic workflows?
Transient, data, permission, external state, semantic ambiguity and genuine novelty. Each requires a different declared handling policy, and only the last two should reach a human.
Should permission failures be retried?
No. An expired credential or insufficient scope will not resolve on retry, so retrying adds latency, produces duplicate log entries and obscures the cause. Fail fast with the specific scope required and notify the owner.
How many retry attempts should a payment get?
One. An unknown outcome on a payment is unsafe to retry, because the failure was frequently a timeout on a call that succeeded upstream. The correct response is escalation to a human with a reconciliation query.
Why must retries reuse the original idempotency key?
Because a retry with a newly generated key is a second request rather than a repeat of the first. This is the single most damaging retry bug in agentic automation and it produces duplicate payments and duplicate messages.
What is external state as an exception class?
A failure indicating the goal is already achieved — the invoice is already paid, the record already updated. It looks like a failure and is not, and the correct handling is to reconcile and proceed idempotently rather than to retry or escalate.
Can semantic ambiguity be resolved automatically?
It should not be. Two plausible candidate matches at similar confidence is a human question, and resolving it by threshold produces a silent wrong answer that is consistent, invisible internally and eventually visible to the customer.
What makes an escalation actionable?
Four elements: the concrete action attempted with resolved values, what is known and explicitly what is not, two or three concrete options with their implications, and the default if nobody acts and by when.
What is the most important sentence an escalation can contain?
That an outcome is unknown. A recipient told a payment’s status is unknown behaves entirely differently from one told the payment failed, and conflating the two is how duplicates and gaps both occur.
How do you distinguish an exception defect from noise?
By aggregating across runs by class, step and workflow. A single transient timeout is noise; the same data exception at the same step forty times is a design or data defect that handling is merely absorbing.
What is the earliest warning that an upstream system changed?
A shift in exception mix rather than a rise in exception rate. Composition changes typically precede rate changes by several days, so watching the distribution catches upstream changes earlier than watching totals.
What is the highest-return way to reduce exceptions?
Pre-flight validation of connector health and credentials. It moves the entire permission class from mid-run failures — which strand partially completed work — to a clean refusal before anything has been written.
Glossary
- Exception class
- One of six categories of workflow failure, each with a declared handling policy.
- Retry budget
- The bounded number of attempts permitted for a class of action before escalation.
- Escalation payload
- The structured information handed to a human when a workflow cannot proceed.
- External state exception
- A failure indicating the intended effect already exists.
- Exception signature
- The combination of class, step and workflow used to detect recurrence.
- Exception mix
- The composition of exceptions by class, whose change precedes rate change.
- Thundering herd
- Synchronised retries across concurrent runs overwhelming a struggling upstream.
- Pre-flight validation
- Checking connectors and credentials before planning, eliminating mid-run permission failures.
- Recurrence rate
- How often a given exception signature repeats, distinguishing defects from noise.
- Handling retirement
- Removing exception handling for a class that no longer occurs.
Standards and entities referenced
Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.
- 01 · Michael NygardRelease It! — stability patterns ↗Bulkheads, circuit breakers and timeouts, from the book that named them.
- 02 · Martin FowlerCircuit breaker pattern ↗The original write-up of the pattern agent runaway control borrows.
- 03 · WikipediaIdempotence ↗Why safe retries require this property rather than hope.
- 04 · microservices.ioSaga pattern ↗Compensating transactions, which is all you get when distributed rollback does not exist.
- 05 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
- 06 · Chaos Engineering communityPrinciples of Chaos Engineering ↗Why deliberately breaking a system is the only way to know its failure behaviour.
- 07 · Google CloudDORA metrics ↗Precedent for measuring a delivery process rather than its output.
- 08 · Object Management GroupBPMN 2.0 specification ↗The modelling standard business process orchestration vocabulary comes from.
Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). Exception Handling in Agentic Workflows: The Seventy Percent Nobody Budgets For. Real Biz Digital. https://realbizdigital.net/insights/agentic-workflow-exception-handling/
Try the mechanics on a live server
To see a governed tool surface respond before you point an agent at your accounting system — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
BarzelOps is this execution layer, sold as a running product
Forty tools covering capability discovery, workflow planning, preview before execution, run, trace, pause, resume and cancel, approval request and resolution, credential validation and connector health — across accounting, CRM, email, calendar, documents, Slack and storage. Opinionated workflows ship with it: customer onboarding, invoice follow-up, lead-to-invoice, pipeline cleanup, monthly close preparation and weekly operations briefings. The Free tier runs 100 calls a day.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Free | Free | 100 calls/day · 10 core tools: plan, preview, run, trace | Proving one workflow end to end before anyone signs anything |
| Pro | $19/mo | 15,000 calls/mo · 26 tools including approvals and templates | One operator automating their own recurring procedures |
| Team | $49/mo | 50,000 calls/mo · the complete 40-tool surface | An operations team running cross-app workflows under approval |
| Enterprise | $199/mo | Unlimited calls · deployment and connector scope by agreement | Multi-team execution with audit and residency obligations |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
The five Barzel servers, and which problem each one is sold for
One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.
| Server | Sold for | Entry price | Where it sits |
|---|---|---|---|
| Barzel Central Gateway | Knowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidence | Free, then $10–$149/mo | Control plane — decides what may be reached, and by whom |
| BarzelVault | Stopping a specific dangerous action before it executes, with proof afterwards | $199–$3,999/mo | Decision point — evaluates the individual call before execution |
| BarzelOps | Running real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approval | Free, then $19–$199/mo | Execution layer — does the work the policy allowed |
| Barzel FinOps Atlas | Attributing AI spend to agents, tools and outcomes, then forecasting and capping it | Free, then $29–$799/mo | Economics layer — what the estate costs per outcome |
| Barzel Scripture Intelligence | A free, credential-free public MCP server to test clients and inspect real protocol traffic | Free, unmetered, no signup | Reference implementation — safe place to learn the protocol |
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.