Why is the number of requests the wrong denominator?
Because a request is a unit of your architecture, not a unit of your business. Finance counts invoices processed and cases closed; the cost is incurred somewhere else entirely, in the fan-out between the outcome and the work it took to reach it.
A single business action decomposes roughly like this:
| Stage | Calls incurred |
|---|---|
| Context retrieval — records, documents, prior state | 1–n, depending on how much the agent decides it needs |
| Reasoning and planning steps | 1–n, unbounded if the agent loops |
| Tool calls against internal systems | 1 per tool, more when a result is ambiguous |
| Validation or self-check passes | 0–n |
| Retries after timeouts or unclear responses | 0–n |
| Audit-trail writes | several, on your side |
The spread between a well-built agent and a badly built one lives almost entirely in the last three rows, and none of it is visible in a count of documents. Hence the first measurement rule: count operations, count outcomes, and report the ratio. That ratio is independent of anyone's price list, so it moves only when your system changes — which makes it the cleanest quality signal you will get. The same argument, worked through for a regulated invoicing process, is set out in the Polish-market treatment of this problem.
What makes agent cost behave differently from ordinary cloud cost?
Three properties, each of which breaks an assumption conventional FinOps rests on.
Cost scales with input length, not with business value
A 200-line invoice costs many times what a one-line invoice costs to process, for identical business worth. Volume forecasts built on document counts are therefore forecasts of the wrong quantity, and a shift in input mix — a new document format, a longer retrieved context — moves spend with no change in throughput.
Failed calls cost the same as successful ones
A failed request against a conventional service costs approximately nothing; you retry and forget it. A failed model call costs what a successful one costs, because the input was processed regardless. An agent that interprets an ambiguous result as a failure and tries again pays in full for every attempt. This is usually the largest single source of variance in agent spend, and it is invisible unless retries are flagged as retries.
Behaviour changes without a deployment
When the underlying model is swapped, both output quality and unit cost can move with nothing changing on your side. Conventional capacity planning assumes a workload changes only when you change it. Here it does not.
| Property | Conventional service | Agent operation |
|---|---|---|
| Unit cost driven by | Requests, compute time, storage | Input and output length |
| Cost of a failed call | Near zero | The same as a successful one |
| Change without a deploy | Rare | Routine — model substitution |
| Cost across comparable units | Fairly tight | Long-tailed |
Why is the average cost per action the wrong number to report?
This is the most actionable point in the article. The distribution is the story, not the mean. Agent cost per action is long-tailed: a small share of actions routinely drives a large share of spend, and the mean sits well below that tail. A stable average is therefore no evidence of a stable process. The tail can thicken for a month while the average barely moves, and by the time the average moves you are reading a bill, not a signal.
Report p50, p95 and p99 alongside the share of spend taken by the most expensive decile. The pay-off is diagnostic. If every percentile rises together, something general changed — a price, a model, a prompt applied everywhere. If only the upper percentiles rise, the input mix changed or a retry path started firing on a subset of cases. Those need entirely different responses, and the average moves identically in both.
Why are governance and cost the same instrumentation problem?
The record needed to answer what did this agent do, and who approved it is the record needed to answer what did it cost: identity, operation, parameters, outcome, timestamp. Attach token counts and duration and the audit trail is also a cost ledger. Treated as separate projects, they get built twice — a compliance log nobody queries, and a finance dashboard that cannot attribute anything to a decision. More often, the argument over ownership means neither gets built at all.
The overlap is not only convenient, it is diagnostic. FINRA's 2026 Annual Regulatory Oversight Report names misaligned reward structures and autonomy without human validation among AI agent risks. Both have cost signatures. An agent optimising a proxy rather than the outcome shows up as operations per outcome climbing while outcomes stay flat. Autonomy without validation shows up in the tail — the most expensive actions are, reliably, the ones nobody looked at. That is also the argument for setting approval thresholds from the distribution rather than from intuition.
What does retention cost, and why does nobody budget for it?
Diagnostic logs rotate. Audit trails held for compliance do not. The retention period is set by regulation rather than by engineering: the California CPPA regulations require risk assessments to be retained for five years, and German § 14b UStG requires e-invoices to be kept for eight. Storage on that clock grows linearly and never plateaus for as long as the process runs.
Two consequences follow. Separate the retention classes explicitly, so verbose diagnostic output is not inherited into a compliance-clock archive for no evidential gain. And model the cost over the full period, not this year's volume: an eight-year clock means year one's records are still on the books in year eight, alongside seven further years of accumulation.
Which numbers belong in the monthly report?
| Metric | What it tells you |
|---|---|
| Operations per business outcome | Implementation quality. Price-independent, so it moves only when your system does. |
| Retry share of total calls, by cause | Whether spend is work or churn. Timeouts, ambiguity and validation failures need different fixes. |
| Cost per outcome, by process | Which processes are not worth automating. Aggregates hide this completely. |
| Agent cost per outcome, separate from infrastructure | The two have different dynamics; combined, each masks movement in the other. |
| p50 / p95 / p99 cost per action | The distribution. Never the mean alone. |
| Audit-trail storage growth | The line that compounds on a compliance clock rather than a business one. |
What should you instrument first?
- A correlation identifier binding every call to the business outcome it serves. Without it, nothing else on this list is computable.
- Token counts and duration on every call, recorded against the operation rather than rolled into a monthly total.
- An explicit retry flag with a cause. A retry indistinguishable from a first attempt makes the churn invisible exactly where it is largest.
- The process name on every operation, so cost attributes to something a business owner recognises and can decide about.
- Percentile reporting from day one. Distributions cannot be reconstructed from data already aggregated to monthly averages.
Frequently asked questions
What is the right unit of cost for an agent?
The operation. Count operations and outcomes separately, and report the ratio — it is the leading indicator.
Why not report an average cost per action?
Because the distribution is long-tailed and the mean sits below the tail. Report p50, p95 and p99 and the top decile's share of spend.
Why do failed calls matter so much?
They cost the same as successful ones. An agent that retries on ambiguity pays in full each time, and unflagged retries are invisible.
Should governance logging and cost measurement be one project?
Yes. The fields are the same; only the reader differs. Built separately, they are usually built twice or not at all.
How should retention be budgeted?
Over the full regulatory clock, with compliance archives kept separate from rotating diagnostic logs.
Related
- AI agent governance — the complete guide
- How to build an audit trail for AI agents
- Approving AI agent actions before execution
- Koszty automatyzacji KSeF — the Polish-market treatment of the same problem
FinOps Atlas measures cost at the level of the individual operation and binds it to the process and the business outcome it belongs to — reporting the distribution rather than the mean, and separating agent cost from infrastructure, where the two move for different reasons.
In practice
Know what every agent action costs before the invoice arrives.
Agent cost is decided per action, long before a monthly invoice makes it visible. FinOps Atlas allocates spend to workflows and outcomes as it happens, caps it, flags anomalies and produces the evidence month-end close and an audit need.
Barzel FinOps Atlas
Intelligent financial operations for AI agents and automation.
- Cost per action, workflow and business outcome, allocated as it happens.
- Spend limits and anomaly detection before the bill, not after.
- Financial evidence tracing and close-readiness for SOX, SOC 2 and external audit.
Free tier: 500 calls a monthPaid plans from $29 a monthLive on MCPize
BarzelVault
The AI action firewall: decide what an agent may do before it does it.
- Approval thresholds and policy checks enforced before execution; human approvals that expire and escalate.
- Cryptographically signed audit receipts: trigger, inputs, policy version, approver, outcome.
- Credential isolation, spend and action limits, and an emergency kill switch.
Free tier: 10,000 calls a monthPaid plans from $199 a monthLive on MCPize
Enterprise: written quote by email within two business days. No sales call.
Sources
- FINRA, 2026 Annual Regulatory Oversight Report — AI agent risks, including misaligned reward structures and autonomy without human validation.
- California Privacy Protection Agency, regulations requiring retention of risk assessments for five years.
- Umsatzsteuergesetz (Germany), § 14b — eight-year retention of invoices.
This article is for information and does not constitute legal, tax or financial advice.