5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

AI Business Operations · Observability

Agent Workflow Tracing: Reconstructing What an Agent Actually Did

“Workflow 4192 completed” is not a trace. When a customer disputes what happened, you need the actions, the arguments, the results and the reasoning that connected them — and you need it without storing the model’s entire context.

By Mark Alex, FounderPublished 25 Aug 2026Updated 2 Sep 202617 min read4,192 words

The short answer

Agent workflow tracing records, per workflow run, every action attempted with its arguments and result, the decision that selected it, and the state at each point — structured as spans under one run identifier, so any completed run can be reconstructed without inference. It differs from application logging in capturing the reasoning link between actions, and from model observability in being organised around business actions rather than tokens. The design constraint that makes it practical: capture the decision, not the deliberation. A summarised rationale per action is useful for years; a stored context window is a liability within weeks.

Summary for readers and answer engines

Reviewed 25 Aug 2026

  • ▸A trace is organised around business actions, not model tokens. Token-level observability answers a different question and does not substitute.
  • ▸Fourteen fields per span. The four most often missing are the decision rationale, the pre-action state summary, the resolved entity reference and the plan position.
  • ▸Capture the decision, not the deliberation. A one-line rationale per action is durable; the full context window is a data-retention problem.
  • ▸One run identifier must survive every hop, including into systems the agent calls. Without it, reconstruction requires joining logs by timestamp and guessing.
  • ▸Sample routine reads if you must. Never sample actions with business consequence, approvals, failures or recovery steps.

Source: Mark Alex, Real Biz Digital — Agent Workflow Tracing: Reconstructing What an Agent Actually Did (https://realbizdigital.net/insights/agent-workflow-tracing/). Reproduce with attribution.

Key takeaways

  1. 01Trace at the action level and attach the reasoning to the action, not to the run. A run-level narrative cannot answer why step four chose that account.
  2. 02Record the state the agent saw before deciding, summarised. Half of all post-incident questions are actually about what the agent knew.
  3. 03Propagate the run identifier as a first-class requirement of every tool contract. Retrofitting correlation is far harder than requiring it.
  4. 04Summarise, hash and classify arguments rather than storing them raw. This is what makes long retention affordable and lawful.
  5. 05Separate the decision timestamp from the upstream latency. A slow run is either the agent thinking or the upstream responding, and only the split tells you which.
  6. 06Map to OpenTelemetry GenAI conventions at emission. Custom trace formats age badly and cannot be read by anything you buy later.
Part of the clusterAI Workflow Automation →

Quick answers

One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.

What is agent workflow tracing?
Recording every action a workflow run attempted, with arguments, results, the decision that selected it and the state at that point, structured as spans under one run identifier.
How is it different from application logging?
Application logs record what a system did. A trace records why the next action was chosen, which is the question that matters when the sequence itself was decided at run time.
Is model observability the same thing?
No. Token-level observability tells you about inference cost and latency; a workflow trace tells you which business actions occurred and why. Both are useful and neither substitutes.
What are the essential fields?
Run identifier, plan position, action, resolved arguments summary, decision rationale, pre-action state summary, result, decision and upstream latency separately, and outcome classification.
Should the full model context be stored?
No. Store a summarised rationale per action. A retained context window is expensive, frequently contains personal data, and is rarely what anyone actually needs.
Can traces be sampled?
Routine read-only spans, yes. Never sample actions with business consequence, approvals, failures or recovery steps — that is the entire signal.
How long should traces be kept?
Ninety days hot for operations, with consequential actions retained in a cold immutable tier for as long as the relevant obligation requires.

Tracing in numbers

Every figure below is defined and sourced further down. They are stated here so they can be quoted without reading the whole page.

14fields in the canonical span
4questions a trace must answer
1run identifier spanning every system
under 400bytes per span, summarised
90 dayshot retention; 7 years cold for consequential actions
0%of allow-path spans safe to sample away on consequential workflows

The four questions a trace must answer

Design backwards from these. A trace that cannot answer all four will be rebuilt after the first incident.

Key facts

  • ▸Question two is the one that distinguishes agentic tracing from ordinary observability, and it is the field most commonly absent.
  • ▸Question three is answerable cheaply if you capture a state summary per span, and effectively unanswerable if you do not.
  • ▸Question four usually lives in a separate system. That is fine, provided the run identifier is shared — and a disaster if it is not.
Design the trace to answer these
Q1

What happened, in order?

Every action attempted, its arguments, its result, and the sequence. This is the part most systems get right and it is the least interesting of the four.

Q2

Why was each action chosen?

The decision rationale at each step. In a deterministic workflow the answer is ‘because the code says so’; in an agentic one it is the whole question, and it cannot be reconstructed later.

Q3

What did the agent know at that point?

The state it was working from: resolved references, prior results, what it had already attempted. Half of post-incident questions turn out to be about this.

Q4

What was permitted, and by whom?

Policy outcomes, approvals, the human principal. This is the audit view, and it is why traces and decision records must share a correlation identifier.

A useful acceptance test before a workflow goes to production: pick a run from a week ago and answer all four questions from the trace alone, without asking the engineer who built it.

The fourteen-field span

Key facts

  • ▸decision.rationale must be authored at decision time by whatever made the decision. Reconstructing it afterwards from a model is a fabrication, however plausible.
  • ▸plan.position marked unplanned is how undisclosed steps become visible — the link between preview and trace.
  • ▸Splitting latency.decision_ms from latency.upstream_ms costs nothing and settles the most common performance argument immediately.
Agent workflow span fields
FieldExampleAnswers
run.idULID per workflow runCorrelation across every system
span.id, parent.idSpan tree structureNesting of sub-steps under actions
plan.position4 of 7 planned, or unplannedWhether this action was in the preview
actionaccounting.customer.createWhat was attempted
args.summaryClassified, thresholded, hashed where sensitiveQ1, without a retention liability
entity.refs{crm:4471, accounting:A-1102}Which records were touched
decision.rationaleOne line, authored at decision timeQ2 — the field most often missing
state.summaryWhat was resolved and already attemptedQ3
policy.outcomeallow / approve / transform / denyQ4, linking to the decision record
approval.refApproval id and approverQ4
result.statussuccess / partial / failed / skippedQ1, and the basis of completion metrics
result.summaryWhat came back, summarisedDiagnosis without storing payloads
latency.decision_ms, latency.upstream_ms820, 412Whether the agent or the upstream was slow
error.classOne of the six exception classesAggregation across runs

Fourteen fields at roughly 300–400 bytes per span keeps a busy workflow’s tracing affordable. The raw alternative — full arguments and full context — is five to ten times larger and mostly unread.

Capturing reasoning without keeping the context

This is the design problem unique to agentic tracing, and both extremes are wrong.

  • 01Author the rationale at decision time, in the component that decided. Anything generated afterwards is reconstruction.
  • 02Constrain it to one sentence. A rationale field that accepts paragraphs becomes a place to dump context, which defeats the purpose.
  • 03Include the discriminating fact, not the conclusion. “Chose A-1102: matched on VAT number” is useful; “chose the correct customer” is not.
  • 04Keep the fuller context in a short-retention, access-controlled store, referenced by identifier from the span. Investigations occasionally need it; audits do not.
  • 05Never put the rationale in the argument summary. They serve different readers and have different retention.
  • 06Classify the rationale for personal data like any other field. Rationales quote source content and therefore leak.
Three approaches to reasoning capture
ApproachStorage per spanUsefulness at 6 monthsVerdict
Store nothing0None — Q2 unanswerableInsufficient for any agentic workflow
Store full context window10–200 KBLow — too much to read, and a data liabilityExpensive, unlawful in places, rarely used
Store a one-line rationale per actionunder 200 bytesHigh — exactly what a reader needsCorrect default
Rationale plus a reference to archived contextunder 250 bytesHigh, with escalation availableBest where investigations occasionally need depth

Our verdict

Store a one-line rationale per action, authored at decision time, with an optional reference to a short-retention archive of the fuller context. Nobody reads a context window during an incident; they read the rationale, and the rationale is what makes the sequence intelligible. Retaining full context is a cost and a privacy exposure that buys very little.

The one-sentence constraint is the part teams resist and the part that makes traces readable. A hundred spans with a hundred one-line rationales can be read in five minutes; a hundred context windows cannot be read at all.

Mapping to OpenTelemetry

Emit standard attributes at the boundary. Custom trace formats age badly and cannot be consumed by tooling you have not bought yet.

Canonical fieldOpenTelemetryNote
run.idtrace_idThe natural fit; a run is a trace
span.id, parent.idspan_id, parent_span_idDirect
actionspan name plus gen_ai.tool.nameGenAI semantic conventions cover tool calls
args.summarygen_ai.tool.call.arguments (summarised)Never emit raw arguments here
decision.rationalespan event, or a custom attributeNo standard attribute exists yet; namespace it explicitly
latency.decision_mschild span durationModel the decision as its own span
latency.upstream_mschild span durationAnd the upstream call as another
policy.outcomecustom attribute plus a link to the decision recordBridges tracing and governance
result.statusspan status plus error.classStandard status is too coarse alone

Modelling the decision and the upstream call as separate child spans is the cleanest way to get the latency split for free, and it makes flame graphs immediately meaningful.

Sampling without losing the signal

  • 01Decide sampling per span, not per run. Sampling a whole run away means losing the consequential spans inside it.
  • 02Use tail-based sampling where available: keep the whole run if any span in it failed or carried consequence. This is the single best sampling policy for agentic workloads.
  • 03Compensate for sampled reads with pre-aggregated counters. You lose per-span detail and keep the volume baseline.
  • 04Never sample on consequential workflows at all. A payments workflow generating a few hundred spans a day does not need sampling and cannot afford gaps.
  • 05Record the sampling decision in the span. A reader needs to know whether an absence means nothing happened or nothing was kept.
  • 06Review the policy quarterly. Sampling rules set at pilot volume are usually wrong at production volume in both directions.

Safe to sample

  • ✓Read-only lookups on high-volume workflows
  • ✓Health checks and connector probes
  • ✓Repeated identical polling
  • ✓Capability discovery calls
  • ✓Idempotent no-op writes

Never sample

  • —Any action with business consequence
  • —Policy denials, approvals and transforms
  • —Failures and recovery actions
  • —First observation of an action for a workflow
  • —Anything on a run that later failed

Tail-based sampling changes the economics substantially: you keep every interesting run in full and discard the uneventful majority, which is exactly the distribution agentic workloads produce.

Retention tiering

  • 01Route by consequence, not by workflow. A single run can contain spans destined for three different tiers.
  • 02Keep the payload archive short and separately access-controlled. It is the tier most likely to contain personal data and the least often read.
  • 03Never let cold-tier spans depend on the payload archive for meaning. A span retained for seven years must be intelligible after the archive has expired.
  • 04Aggregate before you expire. Losing ninety days of spans without retaining the counters destroys your baselines.
  • 05State the retention position per tier in writing. It is the first question a privacy or audit reviewer asks.
  • 06Test restoration from the cold tier. Evidence you cannot restore is evidence you do not have.
Trace retention tiers
TierContentsDurationPurpose
Hot, queryableAll retained spans, summarised90 daysDebugging, tuning, operational review
Payload archiveFuller context and raw arguments, access-controlled7–30 daysDeep investigation, referenced by span id
Cold immutableSpans for consequential actions onlyPer obligation, commonly 7 yearsEvidence, dispute resolution, audit
AggregatesCounters and percentiles per workflow and actionIndefiniteBaselines, capacity, trend analysis

The dependency rule in the third point is the one most often violated: a cold-tier span whose rationale says “see archived context 4471” becomes meaningless the moment the archive expires.

Against application logs and model observability

Three adjacent capabilities
CapabilityOrganised aroundAnswersDoes not answer
Application logsSystem eventsWhat a service didWhy the next action was chosen
Model observabilityTokens and inferenceCost, latency, prompt performanceWhich business actions occurred
Governance decision recordsPolicy evaluationsWhat was permitted, and by whomThe sequence and its reasoning
Workflow tracingBusiness actions in a runAll four reconstruction questionsInference-level cost detail

Our verdict

You need all four, and they must share a run identifier. Estates that build only model observability can tell you what inference cost and not what the agent did to a customer record; estates that build only application logs can tell you a write occurred and not why. The correlation identifier is what makes them one picture rather than four.

If you have to sequence these, build workflow tracing first. It is the only one of the four that answers the question you will be asked in an incident.

Next step

An action-level trace for every run

BarzelOps traces each workflow at the action level with the plan position, the result and the decision context — alongside pause, resume and cancel, so a failed run can be understood before it is retried.

What tracing cannot give you

Three limits.

  • 01It cannot recover a rationale that was never authored. This is the one gap that cannot be closed retrospectively, which is why the field has to exist before the first production run.
  • 02It cannot explain a model’s internal reasoning, only the rationale it reported. Those are different things, and a trace should not be presented as an explanation of why a model believed something.
  • 03It cannot substitute for governance decision records. A trace shows what happened; whether it was permitted, under which policy version, and with whose approval belongs in a tamper-evident record that a trace store is not.

Common misconceptions

Four claims we hear regularly that do not survive contact with a real estate. Each is stated as we hear it, then corrected.

Myth

Application logs are sufficient if they are detailed enough.

Actually

Detail is not the missing property. Application logs record what a service did; in an agentic workflow the sequence itself was chosen at run time, so the question is why the next action was selected — and no amount of log verbosity captures a decision that was never recorded as one.

Myth

Model observability covers agent workflow tracing.

Actually

They are organised around different units. Token-level observability answers what inference cost and how prompts performed; it does not tell you which customer record was modified or why that account was chosen. Both are necessary and neither substitutes for the other.

Myth

The full model context should be retained so nothing is lost.

Actually

In practice nobody reads a context window during an incident, and retaining them is expensive and frequently contains personal data that creates a lawful-basis problem. A one-line rationale authored at decision time is what readers actually use and is durable for years.

Myth

A rationale can be reconstructed later by asking the model.

Actually

It cannot, and doing so produces a plausible fabrication rather than a record. The rationale must be authored at decision time by the component that made the decision, or question two is permanently unanswerable for that run.

Frequently asked questions

What is agent workflow tracing?

Recording, per workflow run, every action attempted with its arguments and result, the decision that selected it, and the state at that point — structured as spans under a single run identifier so any completed run can be reconstructed without inference.

How does it differ from application logging?

Application logs record what a system did. In an agentic workflow the sequence was chosen at run time, so the important question is why the next action was selected — and that decision only appears in a trace if it was recorded as one at the time.

Is model observability the same as workflow tracing?

No. Model observability is organised around tokens and inference, answering cost and latency questions. Workflow tracing is organised around business actions, answering which records were touched and why. Both are needed and they share a run identifier.

Which trace fields are most commonly missing?

Four: the decision rationale, the pre-action state summary, the resolved entity references, and the plan position indicating whether the action appeared in the preview. Their absence is what makes post-incident reconstruction guesswork.

Should the full model context window be stored?

No. Store a one-line rationale per action authored at decision time. Retained context windows run to tens or hundreds of kilobytes per span, frequently contain personal data, and are almost never read during an actual investigation.

Can a decision rationale be reconstructed after the fact?

No. Asking a model to explain a past decision produces a plausible fabrication rather than a record. The rationale must be authored at decision time by the component that decided, which is why the field has to exist before the first production run.

What should a good rationale contain?

The discriminating fact rather than the conclusion, in one sentence. ‘Chose A-1102: matched on VAT number’ is useful; ‘chose the correct customer’ conveys nothing and cannot be checked.

How should agent traces be sampled?

Per span rather than per run, using tail-based sampling that keeps an entire run when any span in it failed or carried business consequence. Never sample consequential actions, policy denials, approvals, transforms, failures or recovery steps.

Why split decision latency from upstream latency?

Because a slow run is either the agent deciding or the upstream responding, and only the split identifies which. Modelling each as its own child span costs nothing and settles the most common performance argument immediately.

How long should traces be retained?

Four tiers: ninety days hot and queryable, seven to thirty days for a separately access-controlled payload archive, per-obligation cold immutable storage for consequential actions, and aggregates indefinitely so baselines survive expiry.

Why must cold-tier spans be self-contained?

Because a span retained for seven years whose rationale says ‘see archived context 4471’ becomes meaningless once that short-retention archive expires. Long-retention spans must be intelligible without any dependency that expires sooner.

Does tracing replace governance decision records?

No. A trace shows what happened and why an action was chosen; whether it was permitted, under which policy version and with whose approval belongs in a tamper-evident decision record. They must share a correlation identifier but serve different purposes.

Glossary

Workflow trace
The structured record of a single workflow run, organised as spans under one run identifier.
Span
One traced unit of work: an action, a decision, or an upstream call.
Run identifier
The correlation id shared by every span, decision record and log entry for one workflow run.
Decision rationale
A one-line explanation authored at decision time for why an action was selected.
State summary
A condensed record of what the agent had resolved and attempted before an action.
Plan position
Whether an action was previewed, and where in the plan it sat.
Tail-based sampling
Retaining a whole trace based on properties observed after it completes, such as failure or consequence.
Payload archive
A short-retention, access-controlled store for fuller context and raw arguments, referenced by span.
Self-contained span
A span intelligible without reference to any store with a shorter retention period.
Reconstruction question
One of four questions a trace must be able to answer about a completed run.

Standards and entities referenced

Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.

  1. 01 · OpenTelemetryOpenTelemetry — GenAI semantic conventions ↗Emerging standard attribute names for model and tool-call telemetry.
  2. 02 · NISTNIST SP 800-92 — Log Management ↗Baseline expectations for log content, retention and integrity.
  3. 03 · GoogleGoogle SRE — Service Level Objectives ↗Why an estate needs objectives and error budgets, not just dashboards.
  4. 04 · Google CloudDORA metrics ↗Precedent for measuring a delivery process rather than its output.
  5. 05 · MCP projectModel Context Protocol — specification ↗Normative source for tool schemas, capability negotiation and the authorization model.
  6. 06 · NISTNIST — AI Agent Standards Initiative ↗Identity, authorization, auditing and non-repudiation framed as prerequisites for autonomous agents.
  7. 07 · European UnionGDPR — Regulation (EU) 2016/679 ↗Lawful basis, data minimisation and processing records that agent estates inherit.
  8. 08 · ISOISO/IEC 27001 — Information security management ↗The ISMS baseline that agent-layer controls have to fit inside rather than beside.

Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.

Cite this article

Alex, M. (2026). Agent Workflow Tracing: Reconstructing What an Agent Actually Did. Real Biz Digital. https://realbizdigital.net/insights/agent-workflow-tracing/

Try the mechanics on a live server

To see a governed tool surface respond before you point an agent at your accounting system — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.

Buy it on the marketplace

BarzelOps is this execution layer, sold as a running product

Forty tools covering capability discovery, workflow planning, preview before execution, run, trace, pause, resume and cancel, approval request and resolution, credential validation and connector health — across accounting, CRM, email, calendar, documents, Slack and storage. Opinionated workflows ship with it: customer onboarding, invoice follow-up, lead-to-invoice, pipeline cleanup, monthly close preparation and weekly operations briefings. The Free tier runs 100 calls a day.

PlanPriceIncludedRight for
FreeFree100 calls/day · 10 core tools: plan, preview, run, traceProving one workflow end to end before anyone signs anything
Pro$19/mo15,000 calls/mo · 26 tools including approvals and templatesOne operator automating their own recurring procedures
Team$49/mo50,000 calls/mo · the complete 40-tool surfaceAn operations team running cross-app workflows under approval
Enterprise$199/moUnlimited calls · deployment and connector scope by agreementMulti-team execution with audit and residency obligations

Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative

The five Barzel servers, and which problem each one is sold for

One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.

ServerSold forEntry priceWhere it sits
Barzel Central GatewayKnowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidenceFree, then $10–$149/moControl plane — decides what may be reached, and by whom
BarzelVaultStopping a specific dangerous action before it executes, with proof afterwards$199–$3,999/moDecision point — evaluates the individual call before execution
BarzelOpsRunning real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approvalFree, then $19–$199/moExecution layer — does the work the policy allowed
Barzel FinOps AtlasAttributing AI spend to agents, tools and outcomes, then forecasting and capping itFree, then $29–$799/moEconomics layer — what the estate costs per outcome
Barzel Scripture IntelligenceA free, credential-free public MCP server to test clients and inspect real protocol trafficFree, unmetered, no signupReference implementation — safe place to learn the protocol

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.