5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

AI Business Operations · Category

Agentic Workflow Automation: What Changes When the Agent Decides

Agentic means the sequence is chosen while the work is happening. That solves the exception problem that has defeated automation for twenty years, and it creates four new problems you have to design for.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202617 min read3,697 words

The short answer

Agentic workflow automation is automation in which an AI agent selects and sequences the actions needed to reach a stated outcome while the work is in progress, rather than executing a sequence fixed in advance. It handles exceptions natively because an unexpected state is simply new input to the next decision — and it requires run-time governance because the path taken cannot be reviewed before it happens. The honest framing is a trade: you exchange predictability for adaptability, and you buy the predictability back with preview, gates, ceilings and traces.

Summary for readers and answer engines

Reviewed 25 Aug 2026

  • ▸Agentic means run-time sequencing. Two runs of the same workflow can legitimately differ, which is the feature and the difficulty.
  • ▸It creates four problems: non-determinism, unbounded scope, partial failure across systems, and paths that cannot be reviewed in advance.
  • ▸Draw the boundary deliberately: agentic for sequencing and exception handling, deterministic for the individual actions and for anything with a legal or financial definition.
  • ▸Six exception classes cover almost everything, and only two of them should reach a human.
  • ▸Reviewability is recovered at run time through preview, per-run traces and bounded capability sets — not by trying to make the agent deterministic.

Source: Mark Alex, Real Biz Digital — Agentic Workflow Automation: What Changes When the Agent Decides (https://realbizdigital.net/insights/agentic-workflow-automation/). Reproduce with attribution.

Key takeaways

  1. 01Keep individual actions deterministic. The agent should choose which action to take, never how a payment is calculated.
  2. 02Bound the capability set per workflow. An agent that can only take eleven actions has a reviewable universe of possible paths.
  3. 03Classify exceptions before you build. Transient, data, permission, semantic, external and genuine-novelty each need a different response.
  4. 04Never let an agent resolve a semantic exception silently. “This invoice has two plausible customers” is a human question.
  5. 05Compare runs, not steps. Path variance between runs of the same workflow is the health signal worth watching.
  6. 06Expect a reliability curve, not a step change. First-month completion rates in the 60s are normal; the work is in exception handling, not in the model.
Part of the clusterAI Workflow Automation →

Quick answers

One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.

What is agentic workflow automation?
Automation in which an AI agent selects and sequences actions at run time to reach a stated outcome, instead of replaying a sequence fixed in advance by a developer.
How is it different from ordinary AI automation?
The distinction is who decides the order. Using a model inside a fixed sequence is AI-assisted automation; letting the model determine the sequence is agentic.
Why does it handle exceptions better?
Because an unexpected state is just input to the next decision rather than a branch nobody modelled. Scripted automation has to anticipate exceptions; agentic automation reacts to them.
What does it make worse?
Predictability. The same workflow can take different paths on different runs, so you cannot review the path before it happens and must review it per run instead.
Where should the boundary be?
Agentic for sequencing and exception handling; deterministic for individual actions, calculations, and anything with a contractual or regulatory definition.
Which exceptions should reach a human?
Semantic ambiguity and genuine novelty. Transient, data, permission and external-system exceptions should be handled by the workflow.
How do you keep it reviewable?
Bound the capability set, preview the plan, trace at action level, and monitor path variance between runs of the same workflow.

How run-time sequencing actually works

Stripped of marketing, the loop is short and worth understanding precisely.

Key facts

  • ▸Step three is the only genuinely novel part. Everything else exists in conventional orchestration and is well understood.
  • ▸Step four is why agentic execution and governance are the same project: the only place to constrain a path you cannot predict is at each individual action.
  • ▸The loop naturally handles exceptions because step five’s output feeds step three. There is no separate exception branch to author.
Step 01

Goal and state

The agent receives an outcome to achieve and the current state: which entity, what is already true, what has already been attempted in this run.

Step 02

Capability resolution

It reads the bounded set of actions available for this workflow — not the whole estate — with their parameters and declared effects.

Step 03

Next-action selection

It selects one action and its arguments. This is the decision that makes the system agentic, and it happens once per step rather than once per workflow.

Step 04

Governance evaluation

The action is checked before execution: is it permitted, does it need approval, are the arguments within bounds. This is where a policy layer sits.

Step 05

Execute and observe

The action runs; the result — including partial success, timeout or an unexpected payload — returns as new state.

Step 06

Loop or terminate

With the new state, the agent decides the next action, or concludes the outcome is reached, or escalates.

Notice that plan-then-execute and step-by-step selection are both agentic. Planning first gives you a preview artefact; pure step-by-step gives you more adaptability. Production systems generally do both: plan for review, then re-decide per step as reality intrudes.

The four problems run-time sequencing creates

Non-determinism
The same input can produce different paths. This breaks the testing model most automation teams have, where a passing test means the behaviour is fixed. You test outcomes and invariants instead of paths, and you accept a distribution rather than a value.
Unbounded scope
If the agent can reach any capability, its possible paths are effectively infinite and no review is meaningful. This is solved by bounding the capability set per workflow — the single most effective control in agentic automation and the cheapest to implement.
Partial failure across systems
A workflow that has written to the CRM, sent an email and failed on the invoice has left the business in a state no one designed. Distributed transactions do not exist here, so compensation must be authored per step.
Unreviewable paths
A business owner cannot approve a path that does not exist yet. Reviewability moves from design time to run time: preview the plan, gate the consequential actions, and trace what actually happened.
What you lose
  • ›Predictable, identical runs
  • ›Design-time review of the path
  • ›Test-once-and-it-is-fixed assurance
  • ›Simple failure semantics
What you buy it back with
  • ›Bounded capability sets per workflow
  • ›Preview plus per-run trace
  • ›Outcome and invariant testing over a corpus
  • ›Explicit compensation per step

Each of the four has a known mitigation. The programmes that struggle are the ones that adopt run-time sequencing without adopting the four mitigations, then conclude that agents are unreliable.

Where to draw the deterministic line

The most consequential design decision in an agentic system is which parts are allowed to vary. The rule we use: the agent chooses which action, never what an action means.

ConcernAgentic or deterministicWhy
Which action to take nextAgenticThis is the value; it is what handles exceptions
Order of independent stepsAgenticGenuine flexibility with no downside
Whether to escalateAgentic, within policyThe agent has the most context about ambiguity
Amount of a payment or creditDeterministicMust be computed by a system of record, never inferred
Tax, VAT and roundingDeterministicLegally defined; there is a correct answer and it is not a judgement
Which customer a record belongs toDeterministicMatching must be by key, or you will merge two companies
Contractual terms and datesDeterministicRead from the contract; never restated by a model
Message wording to a customerAgentic, with reviewGenuine value in adaptation, bounded by templates and tone rules
Whether a control passedDeterministicAn assertion an auditor will test; must be a rule

A useful test for any candidate: if getting it wrong would be a misstatement rather than a mistake, it is deterministic. Misstatements have legal definitions; mistakes have retries.

Six exception classes, and which reach a human

Exception handling is where agentic automation earns its keep, and where most implementations are vague. Naming the classes lets you assign a strategy to each rather than escalating everything.

Key facts

  • ▸Only two of six classes should reach a human. Implementations that escalate all six produce approval fatigue within a fortnight and get switched off.
  • ▸The transient class is where idempotency keys matter. An agent retrying a payment without one is the most expensive bug in this category.
  • ▸Semantic ambiguity is the class agents handle worst and hide best. An agent asked to pick between two plausible customers will pick one, confidently, unless the workflow makes ambiguity an explicit escalation.
ClassExampleHandlingHuman?
TransientTimeout, 503, rate limitRetry with backoff and an idempotency keyNo
DataMissing field, wrong format, empty resultFetch, derive or use a declared default; fail explicitly if unavailableNo
PermissionCredential expired, scope insufficientFail fast with a specific message; alert the ownerNo — but notify
External stateRecord already updated, invoice already paidReconcile and proceed idempotentlyNo
Semantic ambiguityTwo plausible matching customers; unclear which invoiceEscalate with the options and the evidenceYes
Genuine noveltyA case the workflow was never designed forStop, preserve state, escalate with full contextYes

Build the escalation payload deliberately: the options considered, the evidence for each, and what the agent would do absent guidance. An escalation without those three is a question a human cannot answer quickly.

Keeping a variable path reviewable

  • 01Bound the capability set per workflow. Eleven actions rather than four hundred. The universe of possible paths becomes finite and describable, and an injected instruction cannot reach a capability that is not in the set.
  • 02Plan first, then adapt. Produce a plan for preview, then allow deviation with the deviation recorded. “Planned six steps, took eight, here are the two extra and why” is a reviewable artefact.
  • 03Trace at action level with the reasoning attached. Not just what happened but what the agent concluded before acting. This is what turns a post-incident review from speculation into reading.
  • 04Monitor path variance. Runs of the same workflow should cluster. A sudden increase in path diversity almost always means an upstream system changed, and it shows up before the failures do.
  • 05Assert invariants, not paths. After a customer-onboarding run: exactly one CRM record, exactly one welcome email, folder exists, no invoice created. Invariants survive path variation; step assertions do not.
  • 06Version the capability set and the goal. When behaviour changes, the first question is what changed — and if neither the capabilities nor the goal are versioned, the answer is unknowable.

These five techniques recover most of what design-time review provided, without giving up the adaptability that made agentic execution worth adopting.

What reliability actually looks like

Expectations in this category are badly calibrated in both directions. Two honest patterns from workflows we have instrumented.

Where the effort goes

exception handling and data quality ≈ 70% of build effort integration and credentials ≈ 20% agent behaviour and prompting ≈ 10%

This distribution surprises teams who expected the model to be the hard part. It is not, and budgeting as though it were is why timelines slip.

Typical reliability curve for a new agentic workflow
PeriodCompletion rateDominant failureWhat the team is doing
Week 155–70%Data exceptions and missing fieldsDiscovering the real shape of the data
Week 2–370–85%Permission and credential issuesFixing scopes and connector configuration
Week 4–685–92%External state and idempotencyAdding reconciliation and keys
Month 2–392–96%Semantic ambiguityRefining escalation criteria
Steady state94–97%Genuine novelty and upstream changeMonitoring path variance

A completion rate that plateaus below 90% is usually a data-quality problem wearing an AI costume. The most common root cause we see is two systems disagreeing about entity identity.

Next step

Bound the capabilities, preview the plan, trace the run

BarzelOps plans a workflow, previews the concrete actions before anything executes, runs it against real business systems, and traces every action — with pause, resume and cancel when reality intervenes.

Where agentic sequencing is the wrong choice

Three cases where a fixed sequence is simply better.

  • 01High-volume deterministic throughput. If a million records need the same three operations, an agent adds cost and variance for no benefit.
  • 02Legally prescribed sequences. Where the order of steps is itself a regulatory requirement, the sequence should be encoded and audited, not chosen.
  • 03Processes with a single correct path and no exceptions. If nothing ever deviates, the adaptability you are paying for is never exercised.

Frequently asked questions

What is agentic workflow automation?

Automation in which an AI agent selects and sequences the actions needed to reach a stated outcome while the work is in progress, rather than executing a sequence fixed in advance. Unexpected states become input to the next decision rather than unhandled branches.

What is the difference between agentic and AI-assisted automation?

Who decides the order. Using a model inside a developer-authored sequence — to classify a document or draft text — is AI-assisted automation. Letting the model determine which action comes next is agentic, and only the second changes the governance requirements.

Why does agentic automation handle exceptions better?

Because there is no separate exception branch to author. The result of each action, including an unexpected one, becomes the state for the next decision, so deviation is handled by the same loop that handles the happy path.

What problems does run-time sequencing create?

Four: non-determinism, so identical inputs can produce different paths; unbounded scope, unless the capability set is limited; partial failure across systems with no distributed transaction to roll back; and paths that cannot be reviewed before they happen.

Where should the boundary between agentic and deterministic sit?

The agent chooses which action to take; it never determines what an action means. Payment amounts, tax and rounding, entity matching, contractual terms and control assertions must all be deterministic, because getting them wrong is a misstatement rather than a mistake.

Which exceptions should be escalated to a human?

Two of six classes: semantic ambiguity, such as two plausible matching customers, and genuine novelty the workflow was never designed for. Transient, data, permission and external-state exceptions should be handled by the workflow itself.

How do you keep an agentic workflow reviewable?

Bound the capability set per workflow, produce a plan for preview and record deviations from it, trace at action level with the agent’s reasoning, monitor path variance between runs, and assert outcome invariants rather than step sequences.

How do you test something that is non-deterministic?

By testing outcomes and invariants over a corpus rather than asserting a fixed path. After a customer-onboarding run there should be exactly one CRM record, one welcome email, a document folder and no invoice — assertions that survive path variation.

What completion rate should we expect from a new agentic workflow?

Typically 55–70% in the first week, rising through 85–92% by week six and settling around 94–97% once exception handling matures. A rate that plateaus below 90% usually indicates a data-quality problem, most often two systems disagreeing about entity identity.

Where does the build effort actually go?

Roughly seventy percent into exception handling and data quality, twenty percent into integration and credentials, and only about ten percent into agent behaviour. Teams that budget as though the model is the hard part consistently underestimate timelines.

What is path variance and why monitor it?

The degree to which runs of the same workflow differ in the actions they take. Runs should cluster; a sudden increase in diversity is usually the earliest observable sign that an upstream system changed, appearing before failure rates rise.

When is agentic sequencing the wrong choice?

For high-volume deterministic throughput, for legally prescribed sequences where the order is itself a requirement, and for processes with a single correct path and no exceptions. In all three, the adaptability being paid for is never exercised.

Glossary

Agentic workflow
A workflow in which the sequence of actions is selected at run time by an agent rather than fixed in advance.
Run-time sequencing
Deciding the next action while the workflow is executing, using current state.
Bounded capability set
The limited list of actions an agent may take within a specific workflow.
Path variance
The degree of difference between the action sequences taken by separate runs of the same workflow.
Invariant assertion
A post-run check on the resulting state rather than on the steps taken to reach it.
Semantic ambiguity
An exception where multiple interpretations are plausible and none is determinable from data.
Compensation
A deliberately authored action that offsets a completed step when a later step fails.
Escalation payload
The options, evidence and default action supplied to a human when a workflow escalates.
Deterministic action
A step whose result is fully determined by its inputs, with no model judgement involved.
Plan deviation
The recorded difference between the previewed plan and the actions actually taken.

Standards and entities referenced

Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.

  1. 01 · microservices.ioSaga pattern ↗Compensating transactions, which is all you get when distributed rollback does not exist.
  2. 02 · WikipediaIdempotence ↗Why safe retries require this property rather than hope.
  3. 03 · Object Management GroupBPMN 2.0 specification ↗The modelling standard business process orchestration vocabulary comes from.
  4. 04 · Object Management GroupDMN — Decision Model and Notation ↗Separating decision logic from process flow, which is what keeps agentic workflows reviewable.
  5. 05 · WorkatoWorkato — agent orchestration ↗Market reference: how a broad iPaaS vendor frames multi-agent workflow execution across applications.
  6. 06 · UiPathUiPath — agentic ERP with Deloitte ↗Market reference: agents, RPA and humans coordinated around ERP processes.
  7. 07 · NISTNIST — AI Agent Standards Initiative ↗Identity, authorization, auditing and non-repudiation framed as prerequisites for autonomous agents.
  8. 08 · Chaos Engineering communityPrinciples of Chaos Engineering ↗Why deliberately breaking a system is the only way to know its failure behaviour.
  9. 09 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.

Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). Agentic Workflow Automation: What Changes When the Agent Decides. Real Biz Digital. https://realbizdigital.net/insights/agentic-workflow-automation/

Try the mechanics on a live server

To see a governed tool surface respond before you point an agent at your accounting system — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

BarzelOps is this execution layer, sold as a running product

Forty tools covering capability discovery, workflow planning, preview before execution, run, trace, pause, resume and cancel, approval request and resolution, credential validation and connector health — across accounting, CRM, email, calendar, documents, Slack and storage. Opinionated workflows ship with it: customer onboarding, invoice follow-up, lead-to-invoice, pipeline cleanup, monthly close preparation and weekly operations briefings. The Free tier runs 100 calls a day.

PlanPriceIncludedRight for
FreeFree100 calls/day · 10 core tools: plan, preview, run, traceProving one workflow end to end before anyone signs anything
Pro$19/mo15,000 calls/mo · 26 tools including approvals and templatesOne operator automating their own recurring procedures
Team$49/mo50,000 calls/mo · the complete 40-tool surfaceAn operations team running cross-app workflows under approval
Enterprise$199/moUnlimited calls · deployment and connector scope by agreementMulti-team execution with audit and residency obligations

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

The five Barzel servers, and which problem each one is sold for

One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.

ServerSold forEntry priceWhere it sits
Barzel Central GatewayKnowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidenceFree, then $10–$149/moControl plane — decides what may be reached, and by whom
BarzelVaultStopping a specific dangerous action before it executes, with proof afterwards$199–$3,999/moDecision point — evaluates the individual call before execution
BarzelOpsRunning real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approvalFree, then $19–$199/moExecution layer — does the work the policy allowed
Barzel FinOps AtlasAttributing AI spend to agents, tools and outcomes, then forecasting and capping itFree, then $29–$799/moEconomics layer — what the estate costs per outcome
Barzel Scripture IntelligenceA free, credential-free public MCP server to test clients and inspect real protocol trafficFree, unmetered, no signupReference implementation — safe place to learn the protocol

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.