AI Business Operations · Observability
Agent Workflow Tracing: Reconstructing What an Agent Actually Did
“Workflow 4192 completed” is not a trace. When a customer disputes what happened, you need the actions, the arguments, the results and the reasoning that connected them — and you need it without storing the model’s entire context.
The short answer
Agent workflow tracing records, per workflow run, every action attempted with its arguments and result, the decision that selected it, and the state at each point — structured as spans under one run identifier, so any completed run can be reconstructed without inference. It differs from application logging in capturing the reasoning link between actions, and from model observability in being organised around business actions rather than tokens. The design constraint that makes it practical: capture the decision, not the deliberation. A summarised rationale per action is useful for years; a stored context window is a liability within weeks.
Summary for readers and answer engines
Reviewed 25 Aug 2026
- ▸A trace is organised around business actions, not model tokens. Token-level observability answers a different question and does not substitute.
- ▸Fourteen fields per span. The four most often missing are the decision rationale, the pre-action state summary, the resolved entity reference and the plan position.
- ▸Capture the decision, not the deliberation. A one-line rationale per action is durable; the full context window is a data-retention problem.
- ▸One run identifier must survive every hop, including into systems the agent calls. Without it, reconstruction requires joining logs by timestamp and guessing.
- ▸Sample routine reads if you must. Never sample actions with business consequence, approvals, failures or recovery steps.
Source: Mark Alex, Real Biz Digital — Agent Workflow Tracing: Reconstructing What an Agent Actually Did (https://realbizdigital.net/insights/agent-workflow-tracing/). Reproduce with attribution.
Key takeaways
- 01Trace at the action level and attach the reasoning to the action, not to the run. A run-level narrative cannot answer why step four chose that account.
- 02Record the state the agent saw before deciding, summarised. Half of all post-incident questions are actually about what the agent knew.
- 03Propagate the run identifier as a first-class requirement of every tool contract. Retrofitting correlation is far harder than requiring it.
- 04Summarise, hash and classify arguments rather than storing them raw. This is what makes long retention affordable and lawful.
- 05Separate the decision timestamp from the upstream latency. A slow run is either the agent thinking or the upstream responding, and only the split tells you which.
- 06Map to OpenTelemetry GenAI conventions at emission. Custom trace formats age badly and cannot be read by anything you buy later.
Quick answers
One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.
- What is agent workflow tracing?
- Recording every action a workflow run attempted, with arguments, results, the decision that selected it and the state at that point, structured as spans under one run identifier.
- How is it different from application logging?
- Application logs record what a system did. A trace records why the next action was chosen, which is the question that matters when the sequence itself was decided at run time.
- Is model observability the same thing?
- No. Token-level observability tells you about inference cost and latency; a workflow trace tells you which business actions occurred and why. Both are useful and neither substitutes.
- What are the essential fields?
- Run identifier, plan position, action, resolved arguments summary, decision rationale, pre-action state summary, result, decision and upstream latency separately, and outcome classification.
- Should the full model context be stored?
- No. Store a summarised rationale per action. A retained context window is expensive, frequently contains personal data, and is rarely what anyone actually needs.
- Can traces be sampled?
- Routine read-only spans, yes. Never sample actions with business consequence, approvals, failures or recovery steps — that is the entire signal.
- How long should traces be kept?
- Ninety days hot for operations, with consequential actions retained in a cold immutable tier for as long as the relevant obligation requires.
Tracing in numbers
Every figure below is defined and sourced further down. They are stated here so they can be quoted without reading the whole page.
The four questions a trace must answer
Design backwards from these. A trace that cannot answer all four will be rebuilt after the first incident.
Key facts
- ▸Question two is the one that distinguishes agentic tracing from ordinary observability, and it is the field most commonly absent.
- ▸Question three is answerable cheaply if you capture a state summary per span, and effectively unanswerable if you do not.
- ▸Question four usually lives in a separate system. That is fine, provided the run identifier is shared — and a disaster if it is not.
What happened, in order?
Every action attempted, its arguments, its result, and the sequence. This is the part most systems get right and it is the least interesting of the four.
Why was each action chosen?
The decision rationale at each step. In a deterministic workflow the answer is ‘because the code says so’; in an agentic one it is the whole question, and it cannot be reconstructed later.
What did the agent know at that point?
The state it was working from: resolved references, prior results, what it had already attempted. Half of post-incident questions turn out to be about this.
What was permitted, and by whom?
Policy outcomes, approvals, the human principal. This is the audit view, and it is why traces and decision records must share a correlation identifier.
A useful acceptance test before a workflow goes to production: pick a run from a week ago and answer all four questions from the trace alone, without asking the engineer who built it.
The fourteen-field span
Key facts
- ▸
decision.rationalemust be authored at decision time by whatever made the decision. Reconstructing it afterwards from a model is a fabrication, however plausible. - ▸
plan.positionmarkedunplannedis how undisclosed steps become visible — the link between preview and trace. - ▸Splitting
latency.decision_msfromlatency.upstream_mscosts nothing and settles the most common performance argument immediately.
| Field | Example | Answers |
|---|---|---|
run.id | ULID per workflow run | Correlation across every system |
span.id, parent.id | Span tree structure | Nesting of sub-steps under actions |
plan.position | 4 of 7 planned, or unplanned | Whether this action was in the preview |
action | accounting.customer.create | What was attempted |
args.summary | Classified, thresholded, hashed where sensitive | Q1, without a retention liability |
entity.refs | {crm:4471, accounting:A-1102} | Which records were touched |
decision.rationale | One line, authored at decision time | Q2 — the field most often missing |
state.summary | What was resolved and already attempted | Q3 |
policy.outcome | allow / approve / transform / deny | Q4, linking to the decision record |
approval.ref | Approval id and approver | Q4 |
result.status | success / partial / failed / skipped | Q1, and the basis of completion metrics |
result.summary | What came back, summarised | Diagnosis without storing payloads |
latency.decision_ms, latency.upstream_ms | 820, 412 | Whether the agent or the upstream was slow |
error.class | One of the six exception classes | Aggregation across runs |
Fourteen fields at roughly 300–400 bytes per span keeps a busy workflow’s tracing affordable. The raw alternative — full arguments and full context — is five to ten times larger and mostly unread.
Capturing reasoning without keeping the context
This is the design problem unique to agentic tracing, and both extremes are wrong.
- 01Author the rationale at decision time, in the component that decided. Anything generated afterwards is reconstruction.
- 02Constrain it to one sentence. A rationale field that accepts paragraphs becomes a place to dump context, which defeats the purpose.
- 03Include the discriminating fact, not the conclusion. “Chose A-1102: matched on VAT number” is useful; “chose the correct customer” is not.
- 04Keep the fuller context in a short-retention, access-controlled store, referenced by identifier from the span. Investigations occasionally need it; audits do not.
- 05Never put the rationale in the argument summary. They serve different readers and have different retention.
- 06Classify the rationale for personal data like any other field. Rationales quote source content and therefore leak.
| Approach | Storage per span | Usefulness at 6 months | Verdict |
|---|---|---|---|
| Store nothing | 0 | None — Q2 unanswerable | Insufficient for any agentic workflow |
| Store full context window | 10–200 KB | Low — too much to read, and a data liability | Expensive, unlawful in places, rarely used |
| Store a one-line rationale per action | under 200 bytes | High — exactly what a reader needs | Correct default |
| Rationale plus a reference to archived context | under 250 bytes | High, with escalation available | Best where investigations occasionally need depth |
Our verdict
Store a one-line rationale per action, authored at decision time, with an optional reference to a short-retention archive of the fuller context. Nobody reads a context window during an incident; they read the rationale, and the rationale is what makes the sequence intelligible. Retaining full context is a cost and a privacy exposure that buys very little.
The one-sentence constraint is the part teams resist and the part that makes traces readable. A hundred spans with a hundred one-line rationales can be read in five minutes; a hundred context windows cannot be read at all.
Mapping to OpenTelemetry
Emit standard attributes at the boundary. Custom trace formats age badly and cannot be consumed by tooling you have not bought yet.
| Canonical field | OpenTelemetry | Note |
|---|---|---|
run.id | trace_id | The natural fit; a run is a trace |
span.id, parent.id | span_id, parent_span_id | Direct |
action | span name plus gen_ai.tool.name | GenAI semantic conventions cover tool calls |
args.summary | gen_ai.tool.call.arguments (summarised) | Never emit raw arguments here |
decision.rationale | span event, or a custom attribute | No standard attribute exists yet; namespace it explicitly |
latency.decision_ms | child span duration | Model the decision as its own span |
latency.upstream_ms | child span duration | And the upstream call as another |
policy.outcome | custom attribute plus a link to the decision record | Bridges tracing and governance |
result.status | span status plus error.class | Standard status is too coarse alone |
Modelling the decision and the upstream call as separate child spans is the cleanest way to get the latency split for free, and it makes flame graphs immediately meaningful.
Sampling without losing the signal
- 01Decide sampling per span, not per run. Sampling a whole run away means losing the consequential spans inside it.
- 02Use tail-based sampling where available: keep the whole run if any span in it failed or carried consequence. This is the single best sampling policy for agentic workloads.
- 03Compensate for sampled reads with pre-aggregated counters. You lose per-span detail and keep the volume baseline.
- 04Never sample on consequential workflows at all. A payments workflow generating a few hundred spans a day does not need sampling and cannot afford gaps.
- 05Record the sampling decision in the span. A reader needs to know whether an absence means nothing happened or nothing was kept.
- 06Review the policy quarterly. Sampling rules set at pilot volume are usually wrong at production volume in both directions.
Safe to sample
- ✓Read-only lookups on high-volume workflows
- ✓Health checks and connector probes
- ✓Repeated identical polling
- ✓Capability discovery calls
- ✓Idempotent no-op writes
Never sample
- —Any action with business consequence
- —Policy denials, approvals and transforms
- —Failures and recovery actions
- —First observation of an action for a workflow
- —Anything on a run that later failed
Tail-based sampling changes the economics substantially: you keep every interesting run in full and discard the uneventful majority, which is exactly the distribution agentic workloads produce.
Retention tiering
- 01Route by consequence, not by workflow. A single run can contain spans destined for three different tiers.
- 02Keep the payload archive short and separately access-controlled. It is the tier most likely to contain personal data and the least often read.
- 03Never let cold-tier spans depend on the payload archive for meaning. A span retained for seven years must be intelligible after the archive has expired.
- 04Aggregate before you expire. Losing ninety days of spans without retaining the counters destroys your baselines.
- 05State the retention position per tier in writing. It is the first question a privacy or audit reviewer asks.
- 06Test restoration from the cold tier. Evidence you cannot restore is evidence you do not have.
| Tier | Contents | Duration | Purpose |
|---|---|---|---|
| Hot, queryable | All retained spans, summarised | 90 days | Debugging, tuning, operational review |
| Payload archive | Fuller context and raw arguments, access-controlled | 7–30 days | Deep investigation, referenced by span id |
| Cold immutable | Spans for consequential actions only | Per obligation, commonly 7 years | Evidence, dispute resolution, audit |
| Aggregates | Counters and percentiles per workflow and action | Indefinite | Baselines, capacity, trend analysis |
The dependency rule in the third point is the one most often violated: a cold-tier span whose rationale says “see archived context 4471” becomes meaningless the moment the archive expires.
Against application logs and model observability
| Capability | Organised around | Answers | Does not answer |
|---|---|---|---|
| Application logs | System events | What a service did | Why the next action was chosen |
| Model observability | Tokens and inference | Cost, latency, prompt performance | Which business actions occurred |
| Governance decision records | Policy evaluations | What was permitted, and by whom | The sequence and its reasoning |
| Workflow tracing | Business actions in a run | All four reconstruction questions | Inference-level cost detail |
Our verdict
You need all four, and they must share a run identifier. Estates that build only model observability can tell you what inference cost and not what the agent did to a customer record; estates that build only application logs can tell you a write occurred and not why. The correlation identifier is what makes them one picture rather than four.
If you have to sequence these, build workflow tracing first. It is the only one of the four that answers the question you will be asked in an incident.
Next step
An action-level trace for every run
BarzelOps traces each workflow at the action level with the plan position, the result and the decision context — alongside pause, resume and cancel, so a failed run can be understood before it is retried.
What tracing cannot give you
Three limits.
- 01It cannot recover a rationale that was never authored. This is the one gap that cannot be closed retrospectively, which is why the field has to exist before the first production run.
- 02It cannot explain a model’s internal reasoning, only the rationale it reported. Those are different things, and a trace should not be presented as an explanation of why a model believed something.
- 03It cannot substitute for governance decision records. A trace shows what happened; whether it was permitted, under which policy version, and with whose approval belongs in a tamper-evident record that a trace store is not.
Common misconceptions
Four claims we hear regularly that do not survive contact with a real estate. Each is stated as we hear it, then corrected.
Application logs are sufficient if they are detailed enough.
Detail is not the missing property. Application logs record what a service did; in an agentic workflow the sequence itself was chosen at run time, so the question is why the next action was selected — and no amount of log verbosity captures a decision that was never recorded as one.
Model observability covers agent workflow tracing.
They are organised around different units. Token-level observability answers what inference cost and how prompts performed; it does not tell you which customer record was modified or why that account was chosen. Both are necessary and neither substitutes for the other.
The full model context should be retained so nothing is lost.
In practice nobody reads a context window during an incident, and retaining them is expensive and frequently contains personal data that creates a lawful-basis problem. A one-line rationale authored at decision time is what readers actually use and is durable for years.
A rationale can be reconstructed later by asking the model.
It cannot, and doing so produces a plausible fabrication rather than a record. The rationale must be authored at decision time by the component that made the decision, or question two is permanently unanswerable for that run.
Frequently asked questions
What is agent workflow tracing?
Recording, per workflow run, every action attempted with its arguments and result, the decision that selected it, and the state at that point — structured as spans under a single run identifier so any completed run can be reconstructed without inference.
How does it differ from application logging?
Application logs record what a system did. In an agentic workflow the sequence was chosen at run time, so the important question is why the next action was selected — and that decision only appears in a trace if it was recorded as one at the time.
Is model observability the same as workflow tracing?
No. Model observability is organised around tokens and inference, answering cost and latency questions. Workflow tracing is organised around business actions, answering which records were touched and why. Both are needed and they share a run identifier.
Which trace fields are most commonly missing?
Four: the decision rationale, the pre-action state summary, the resolved entity references, and the plan position indicating whether the action appeared in the preview. Their absence is what makes post-incident reconstruction guesswork.
Should the full model context window be stored?
No. Store a one-line rationale per action authored at decision time. Retained context windows run to tens or hundreds of kilobytes per span, frequently contain personal data, and are almost never read during an actual investigation.
Can a decision rationale be reconstructed after the fact?
No. Asking a model to explain a past decision produces a plausible fabrication rather than a record. The rationale must be authored at decision time by the component that decided, which is why the field has to exist before the first production run.
What should a good rationale contain?
The discriminating fact rather than the conclusion, in one sentence. ‘Chose A-1102: matched on VAT number’ is useful; ‘chose the correct customer’ conveys nothing and cannot be checked.
How should agent traces be sampled?
Per span rather than per run, using tail-based sampling that keeps an entire run when any span in it failed or carried business consequence. Never sample consequential actions, policy denials, approvals, transforms, failures or recovery steps.
Why split decision latency from upstream latency?
Because a slow run is either the agent deciding or the upstream responding, and only the split identifies which. Modelling each as its own child span costs nothing and settles the most common performance argument immediately.
How long should traces be retained?
Four tiers: ninety days hot and queryable, seven to thirty days for a separately access-controlled payload archive, per-obligation cold immutable storage for consequential actions, and aggregates indefinitely so baselines survive expiry.
Why must cold-tier spans be self-contained?
Because a span retained for seven years whose rationale says ‘see archived context 4471’ becomes meaningless once that short-retention archive expires. Long-retention spans must be intelligible without any dependency that expires sooner.
Does tracing replace governance decision records?
No. A trace shows what happened and why an action was chosen; whether it was permitted, under which policy version and with whose approval belongs in a tamper-evident decision record. They must share a correlation identifier but serve different purposes.
Glossary
- Workflow trace
- The structured record of a single workflow run, organised as spans under one run identifier.
- Span
- One traced unit of work: an action, a decision, or an upstream call.
- Run identifier
- The correlation id shared by every span, decision record and log entry for one workflow run.
- Decision rationale
- A one-line explanation authored at decision time for why an action was selected.
- State summary
- A condensed record of what the agent had resolved and attempted before an action.
- Plan position
- Whether an action was previewed, and where in the plan it sat.
- Tail-based sampling
- Retaining a whole trace based on properties observed after it completes, such as failure or consequence.
- Payload archive
- A short-retention, access-controlled store for fuller context and raw arguments, referenced by span.
- Self-contained span
- A span intelligible without reference to any store with a shorter retention period.
- Reconstruction question
- One of four questions a trace must be able to answer about a completed run.
Standards and entities referenced
Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.
- 01 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.
- 02 · NISTNIST SP 800-92 — Log Management ↗Baseline expectations for log content, retention and integrity.
- 03 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
- 04 · Google CloudDORA metrics ↗Precedent for measuring a delivery process rather than its output.
- 05 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
- 06 · NISTNIST — AI Agent Standards Initiative ↗Identity, authorization, auditing and non-repudiation framed as prerequisites for autonomous agents.
- 07 · European UnionGDPR — Regulation (EU) 2016/679 ↗Lawful basis, data minimisation and processing records that agent estates inherit.
- 08 · ISOISO/IEC 27001 — Information security management ↗The ISMS baseline that agent-layer controls have to fit inside rather than beside.
Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). Agent Workflow Tracing: Reconstructing What an Agent Actually Did. Real Biz Digital. https://realbizdigital.net/insights/agent-workflow-tracing/
Try the mechanics on a live server
To see a governed tool surface respond before you point an agent at your accounting system — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
BarzelOps is this execution layer, sold as a running product
Forty tools covering capability discovery, workflow planning, preview before execution, run, trace, pause, resume and cancel, approval request and resolution, credential validation and connector health — across accounting, CRM, email, calendar, documents, Slack and storage. Opinionated workflows ship with it: customer onboarding, invoice follow-up, lead-to-invoice, pipeline cleanup, monthly close preparation and weekly operations briefings. The Free tier runs 100 calls a day.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Free | Free | 100 calls/day · 10 core tools: plan, preview, run, trace | Proving one workflow end to end before anyone signs anything |
| Pro | $19/mo | 15,000 calls/mo · 26 tools including approvals and templates | One operator automating their own recurring procedures |
| Team | $49/mo | 50,000 calls/mo · the complete 40-tool surface | An operations team running cross-app workflows under approval |
| Enterprise | $199/mo | Unlimited calls · deployment and connector scope by agreement | Multi-team execution with audit and residency obligations |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
The five Barzel servers, and which problem each one is sold for
One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.
| Server | Sold for | Entry price | Where it sits |
|---|---|---|---|
| Barzel Central Gateway | Knowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidence | Free, then $10–$149/mo | Control plane — decides what may be reached, and by whom |
| BarzelVault | Stopping a specific dangerous action before it executes, with proof afterwards | $199–$3,999/mo | Decision point — evaluates the individual call before execution |
| BarzelOps | Running real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approval | Free, then $19–$199/mo | Execution layer — does the work the policy allowed |
| Barzel FinOps Atlas | Attributing AI spend to agents, tools and outcomes, then forecasting and capping it | Free, then $29–$799/mo | Economics layer — what the estate costs per outcome |
| Barzel Scripture Intelligence | A free, credential-free public MCP server to test clients and inspect real protocol traffic | Free, unmetered, no signup | Reference implementation — safe place to learn the protocol |
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.