Intelligent Operations · Measurement
Measuring Agentic Process Quality
Ninety-seven per cent accurate against what, over what, compared with whom? Most agent quality reporting cannot answer any of the three.
Key takeaways
- A model evaluation is an input to process quality, not a measure of it. The model is one of at least six ways an agentic process fails.
- Six metrics cover it: outcome correctness, straight-through rate, correction rate, unnoticed error rate, cost per completed outcome, and trigger-to-resolution time.
- Every ratio needs a denominator that includes abandoned runs. Leaving them out is the most common way agent reporting flatters itself.
- Without a stated human baseline, an accuracy figure is not interpretable. The comparison is not against perfection — it is against what the process did before.
- Report per workflow and per action type. An aggregate hides the one workflow that is quietly wrong.
The short answer
Measure the process against a system of record, not the model against an evaluation set. Six numbers carry it: outcome correctness on a verified sample, straight-through rate, correction rate, unnoticed error rate, cost per completed outcome including abandoned runs, and trigger-to-resolution time at the median and the 95th percentile. Report each per workflow, against a stated human baseline, and give the two that users actually feel an error budget with a consequence attached.
Why eval scores mislead
An evaluation set measures whether a model produces the right output for inputs someone has already curated. A process has to produce the right outcome on inputs nobody anticipated, using tools that fail, against data that has moved since the eval was written.
- 01The modelWrong extraction, wrong tool choice, wrong argument. The only surface an eval covers.
- 02The toolsTimeouts, rate limits, partial writes, schema drift on the far side of an integration.
- 03The dataStale records, duplicates, a vendor master that disagrees with the ledger.
- 04The policyLegitimate denials and step-ups that stall a run. Correct behaviour that still shows up as an incomplete outcome.
- 05The humanEscalations sitting in a queue. A process whose 95th-percentile resolution is three days has a quality problem located in an inbox.
- 06The loopRetries that repeat a failing step, and runs abandoned halfway with side effects already committed.
The reframing
Process quality is a property of the whole path from trigger to outcome. The model is one component in it, which is why every useful metric below ends at an outcome in a system of record.
The six metrics
Outcome correctness
Of the outcomes that reached the system of record, the share that is correct when verified against something authoritative. Measured on a sample, by a human or a deterministic check — never by the model that produced it.
Straight-through rate
Share of runs completed with no human intervention. This is the throughput number, and on its own it says nothing about quality.
Correction rate
Share of reviewed items a human changed, and what they changed. This is the calibration number: near zero means the gate is decorative, high means it is too tight.
Unnoticed error rate
Share of accepted outcomes later found to be wrong. Only audit sampling surfaces this, and it is the number that decides whether autonomy can rise.
Cost per completed outcome
Every cost attributable to the workflow — model calls, retries, tools, retrieval, human review time — over completed outcomes.
Trigger-to-resolution time
Median and 95th percentile from the event that started the run to the outcome. The 95th is where escalation queues appear; the median conceals them.
Four of the six are ratios, and all four are only as honest as their denominator.
Getting the denominator right
| Denominator | What it counts | What it hides |
|---|---|---|
| Model calls | Requests to a provider | Retries, abandoned runs, everything outside the model |
| Runs started | Every triggered attempt | Nothing — the honest base for rate metrics |
| Runs completed | Successful paths only | Every failure, which is the interesting part |
| Documents or tickets | Work items | Repeated attempts on the same item |
| Completed outcomes | Business results reached | Nothing, provided abandoned-run cost stays in the numerator |
Scroll the table horizontally on narrow screens.
- Rate metrics use runs started. Cost metrics use completed outcomes, with the cost of abandoned runs kept in the numerator.
- Count a run once. An item retried four times is one run with four attempts, or your automation rate improves every time it fails.
- Keep policy denials in a separate bucket from errors. A denied action is the system working as designed.
- Segment by action type. Aggregating a reversible update with an irreversible payment produces a number that describes neither.
A full worked calculation, including what abandoned runs do to a per-outcome figure, is in cost per outcome: a worked example.
Sampling and the human baseline
None of these numbers can be read without two things: a verified sample, and a baseline to compare it against.
- 01Sample deliberatelyDraw a fixed sample per workflow per period and stratify it: straight-through items, escalated items, and items the agent marked high-confidence. The straight-through stratum is where unnoticed errors live, and it is the one nobody looks at.
- 02Verify against something authoritativeA human check, or a deterministic reconciliation against a record. A model grading its own output measures agreement between two runs of the same system, not correctness.
- 03State the human baselineThe process had an error rate before the agent existed. If nobody measured it, say so — an unstated baseline silently changes the comparison to perfection, which is not the comparison anyone is actually making a decision about.
- 04Report a range, not a decimal placeA sample of sixty with two errors does not distinguish 3% from 10%. Quoting one decimal place from a small sample is the most common overstatement in agent reporting, and it is the one an auditor finds first.
Precision you have not earned
If your volume only supports a few dozen verified items a month, the honest output is a range and a direction of travel. That is not a weaker result — it is the same result, stated at the resolution the evidence supports.
Turning it into an error budget
Two of the six are felt directly by the people the process serves: unnoticed errors, and 95th-percentile resolution time. Those are the two that deserve a budget with teeth.
Workflow SLO — invoice intake
objective 99.0% of posted invoices correct on the monthly audit sample
95% of runs resolved within 8 business hours
window rolling 30 days
budget 1.0% incorrect -> about 6 items at current volume
consequence budget spent -> autonomy for this action drops one level
(execute-and-notify -> draft-and-hold)
2 clean windows -> eligible for promotion
excluded policy denials and step-ups (the system working)
owner AP controller (process) + platform team (pipeline)The consequence clause is the part that matters. A quality metric with no automatic consequence is a report; one that changes the autonomy level of the action is a control. The levels it moves between are in autonomous vs human-in-the-loop operations.
- One budget per workflow, not one for “the agent”.
- Denials and step-ups do not consume the budget; wrong outcomes do.
- A spent budget pauses promotion across that whole workflow, not only on the action that failed.
Metrics that flatter
- ›Model accuracy on a static evaluation set
- ›Share of tasks “handled by AI”
- ›Tokens or calls per task
- ›Hours saved, estimated from headcount
- ›Aggregate accuracy across all workflows
- ›Outcome correctness on a verified sample, per workflow
- ›Straight-through rate read together with correction rate
- ›Unnoticed error rate from audit sampling
- ›Cost per completed outcome, abandoned runs included
- ›95th-percentile trigger-to-resolution time
What unites the left column is that every one of those numbers can improve while the process gets worse. Automation share rises when more work is routed to an agent that is failing quietly. Tokens per task falls when the agent gives up earlier. Hours saved is an estimate multiplied by an assumption.
The instrumentation that makes the right column possible — attribution keys and a per-action execution record — is covered in AI agent cost tracking and how to audit AI agent actions.
The product
The execution record these numbers are computed from
BarzelOps runs the workflows and records every governed action with its approvals and its outcome. Barzel FinOps Atlas supplies the cost attribution and the reconciliation evidence the correctness checks run against.
Key terms
- Outcome correctness
- Whether the result recorded in the system of record is right, verified against an authoritative source.
- Straight-through rate
- Share of runs that complete with no human intervention.
- Correction rate
- Share of reviewed items a human changed.
- Unnoticed error rate
- Share of accepted outcomes later found to be wrong, discoverable only by audit sampling.
- Error budget
- The permitted quantity of failure in a window, with a defined consequence when it is spent.
- Trigger-to-resolution time
- Elapsed time from the event that started a run to the recorded outcome.
Frequently asked questions
How do you measure the quality of an agentic process?
Against outcomes in a system of record rather than model outputs: outcome correctness on a verified sample, straight-through rate, correction rate, unnoticed error rate, cost per completed outcome, and trigger-to-resolution time — each reported per workflow.
Are model evaluations useless for this?
No, they are necessary and insufficient. An eval catches regressions in one component. It cannot see tool failures, stale data, policy denials, escalation queues or abandoned runs.
What denominator should agent metrics use?
Runs started for rate metrics; completed outcomes for cost metrics, with the cost of abandoned runs kept in the numerator. Counting model calls, or completed runs only, flatters the result.
Why does the human baseline matter?
Because the alternative to the agent was never a perfect process. Without the error rate the process had before, an accuracy figure cannot be judged, and the comparison being implied is against perfection.
Can an agent grade its own output?
It can produce a triage signal, not a quality measurement. Self-grading measures agreement between two runs of the same system; correctness requires an authoritative check.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.
- 01 · Google Google SRE Book — Service Level Objectives ↗ Where the SLI/SLO/error-budget vocabulary used here comes from.
- 02 · Google Cloud DORA metrics / DevOps Research and Assessment ↗ Precedent for measuring a delivery process rather than its output.
- 03 · CNCF OpenTelemetry semantic conventions for GenAI ↗ Agreed attribute names for model, token and tool-call telemetry.
- 04 · NIST AI Risk Management Framework ↗ The govern / map / measure / manage structure these controls map onto.
- 05 · ISO ISO/IEC 42001 — AI management systems ↗ Certifiable management-system standard for AI governance programmes.
- 06 · FinOps Foundation FOCUS — FinOps Open Cost and Usage Specification ↗ Open schema for normalising cost and usage data across providers.
Last reviewed 19 August 2026. External links open in a new tab; we do not control their content.
Reference documentation
Want the specification rather than the argument?
Workflow traces, preview runs and the audit tools you measure process quality with.
Free tier · 100 calls/day · on the MCPize marketplace
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.