5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

FinOps · Unit economics

Cost per Outcome: A Worked Example

Cost per token is easy to measure and answers no question anyone has. Here is the arithmetic for the number that does.

By Mark Alex, Founder Published 19 Aug 2026 10 min read
Part of the cluster FinOps for the AI Era →

Key takeaways

  • The denominator is the whole argument. Cost per run flatters; cost per completed outcome is decision-grade.
  • Retries and abandoned runs are not edge cases in non-deterministic systems — they are a standing percentage of volume.
  • Human review usually dominates. In the example it is 54% of total cost while being 18% of volume.
  • Compare against the human baseline honestly, including the cost of the reviewer you kept.
  • Once you have this number, optimisation priorities invert: reducing the approval rate beats switching models.

The short answer

Cost per outcome = all consumption attributable to a business result, divided by the number of results achieved — including retries, abandoned runs and human review time. In the worked example below, the naive per-token figure is $0.11 and the true cost per completed invoice is $1.30, because 22% of runs need a retry, 7% are abandoned, and 18% require four minutes of human review. Both numbers are arithmetically correct. Only one supports a decision.

The scenario

An accounts-payable agent processes supplier invoices: reads the document, extracts fields, matches to a purchase order, flags discrepancies, and either posts the invoice or routes it for approval. One completed outcome is one invoice correctly posted or correctly escalated.

The figures below are illustrative and internally consistent — they are here to demonstrate the method, not to represent a benchmark. Substitute your own; the structure is the point.

Observed over one month

completed outcomes           10,000 invoices

runs attempted               11,530
  └ first-attempt success     8,700   (75.5%)
  └ succeeded after retry     2,200   (19.1%)
  └ abandoned, no outcome       630    (5.5%)

human review                  1,800 invoices  (18%)
  └ average review time         4 min
  └ loaded reviewer cost      $52 / hour

The naive number

Take the model spend, divide by invoices. This is what most dashboards report and what most business cases are built on.

Per-token view

model spend for the month        $1,100
completed outcomes               10,000

cost per outcome                  $0.11

What it silently excludes

Retries are counted as successes rather than as the extra cost of one outcome. The 630 abandoned runs cost money and produced nothing, and are simply absent. Tool calls, retrieval and storage sit on other invoices. Human review is not considered a cost of the workflow at all. Every exclusion pushes the number down.

The real number

Same month, every attributable cost, same denominator.

Cost lineAmountShare
Model calls — successful first attempts$7405.7%
Model calls — retried runs$3102.4%
Model calls — abandoned runs$500.4%
Tool calls (OCR, PO lookup, posting)$1,1909.2%
Retrieval, embeddings and document storage$4803.7%
Gateway and policy evaluation$1801.4%
Human review — 1,800 × 4 min × $52/h$6,24048.1%
Escalation handling and exception work$1,78013.7%
Platform and on-call allocation$2,00015.4%
Total$12,970100%

Scroll the table horizontally on narrow screens.

Outcome view

total attributable cost         $12,970
completed outcomes               10,000

cost per outcome                  $1.30

naive figure                      $0.11
understated by                    11.8×

Why the gap is this large

Model tokens are 8.5% of the total. Everything else — tools, retrieval, policy, and above all human time — is the other 91.5%. Optimising the 8.5% is where teams instinctively start, and it is nearly the least useful place to work.

What the number tells you

With the breakdown in hand, the optimisation order is no longer a matter of taste.

  1. 01Reduce the approval rateHuman review is 48% of cost. Moving the review rate from 18% to 9% saves roughly $3,100 a month — more than eliminating model spend entirely. That is a threshold-design problem, covered in designing approval thresholds.
  2. 02Cut the retry rateRetries add cost and latency and are usually caused by a small number of recurring failure shapes. Fixing the top two failure modes typically moves this more than a model upgrade would.
  3. 03Attack tool-call costAt 9.2%, tool calls cost more than model tokens here. Caching PO lookups and batching OCR are unglamorous and effective.
  4. 04Only then consider the modelA cheaper model saves a fraction of 8.5% and risks raising the retry and escalation rates — both of which are more expensive. Routing decisions should be judged on this denominator, not on token price.

Comparing to the human baseline

The comparison that decides whether the workflow should exist. It has to be done on the same basis, which is where most business cases quietly cheat.

Dishonest comparison
  • ›Fully-loaded human cost vs model tokens only
  • ›Assumes 100% automation
  • ›Ignores the reviewer you still employ
  • ›Uses first-attempt success as the success rate
Honest comparison
  • ›Fully-loaded human cost vs fully-loaded agent cost
  • ›Uses the real automation rate (82% here)
  • ›Counts remaining human time as agent cost
  • ›Uses completed outcomes as the denominator

Same basis

manual baseline
  11 min/invoice × $52/h        =  $9.53 per invoice

agent workflow (all-in)          =  $1.30 per invoice

saving per invoice                  $8.23
at 10,000/month                    $82,300

automation rate                        82%
(18% still touched by a human)

That is a strong result and it remains true after scrutiny — which is the only kind of result worth presenting to a CFO. A case built on $0.11 collapses the first time somebody asks about the review queue, and it takes the credibility of the programme with it.

Turning this into an enforceable budget is the next step: see setting agent spend ceilings.

Diagram · free to reuse with attribution

The layers the arithmetic rolls up through

Each layer in the diagram is a row in the table above. The naive figure measures only the bottom band, and only part of it.

Original diagram by Real Biz Digital. Reuse it anywhere with a link back to this article.

The product

Unit economics you can put in front of a CFO

Barzel FinOps Atlas produces CFO summaries, controller memos and board finance packs with evidence citations — built on attributed cost rather than provider totals.

Key terms

Cost per outcome
Total attributable cost divided by the number of completed business results, including failed and retried attempts.
Automation rate
The share of outcomes completed without human involvement.
Loaded cost
An employee cost including salary, benefits, tax and overhead, used for honest labour comparisons.
Abandoned run
An execution that consumed resources and produced no outcome.
Unit economics
The cost and value of one unit of business activity, as opposed to aggregate spend.

Frequently asked questions

What is cost per outcome for AI agents?

Total attributable cost — model calls, tool calls, retrieval, policy evaluation, retries, abandoned runs, human review and platform allocation — divided by the number of completed business results.

Why is cost per token misleading?

It measures one input to one layer. In the worked example above model tokens are 8.5% of true cost, so the per-token figure understates cost per outcome by nearly twelve times and points optimisation at the smallest line.

Should human review time be included?

Yes, and it is usually the largest line. In the example it is 48% of total cost at 18% of volume, which makes reducing the approval rate a bigger lever than any model change.

Do abandoned runs count?

Yes. They consumed resources and produced no outcome. Excluding them flatters the result precisely where honesty matters, and it hides a failure rate somebody should be working on.

How do we compare fairly against a human baseline?

Compare fully-loaded to fully-loaded, use the real automation rate rather than assuming full automation, count the human review time that remains as a cost of the agent workflow, and use completed outcomes as the denominator on both sides.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.

  1. 01 · FinOps Foundation FinOps Framework ↗ Capability model and phases referenced throughout.
  2. 02 · FinOps Foundation FOCUS — FinOps Open Cost and Usage Specification ↗ Open schema for normalising cost and usage data across providers.
  3. 03 · FinOps Foundation FinOps for AI overview ↗ Where AI workloads change the practice.
  4. 04 · CNCF OpenTelemetry semantic conventions for GenAI ↗ Agreed attribute names for model, token and tool-call telemetry.
  5. 05 · Google Google SRE Book — Service Level Objectives ↗ Where the SLI/SLO/error-budget vocabulary used here comes from.

Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.

Reference documentation

Want the specification rather than the argument?

The finance workflows and reporting tools behind a cost-per-outcome number.

Free · 500 calls/mo · on the MCPize marketplace

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.