5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

FinOps · Attribution

AI Agent Cost Tracking

The invoice arrives as one number. Everything useful about it — who, why, whether it was worth it — has to be captured before the call is made.

By Mark Alex, Founder Published 19 Aug 2026 9 min read
Part of the cluster FinOps for the AI Era →

Key takeaways

  • Attribution is a write-time decision. No amount of later analysis recovers a key that was never captured.
  • Run ID is the load-bearing field. Retries, sub-agent calls and tool calls all belong to one outcome, and only that identifier ties them together.
  • Four costs get missed consistently: retries, abandoned runs, human review minutes, and retrieval or storage.
  • Tag with keys that map to accountable owners. Cost split by model name is a fact nobody can act on.
  • Use FOCUS field names where they exist, so agent spend joins your existing cloud cost data instead of forming a second silo.

The short answer

Attribution cannot be reconstructed after the fact; it has to be attached at call time. Instrument six keys on every model and tool call — run ID, agent identity, acting user or team, workflow, environment and outcome reference — and record token counts, tool-call costs, retries and human review time against them. Without the run ID linking retries and sub-calls to one outcome, you have a cost total and no unit economics.

Why the invoice cannot help

A provider bill is an accurate record of consumption and a useless record of causation. It knows tokens, models and API keys. It does not know your teams, workflows, customers or outcomes, and it never will.

  • One API key usually fronts many workflows, so key-level breakdown attributes everything to “the platform”.
  • Retries appear as independent calls, so a workflow that fails twice looks like three cheap successes.
  • Sub-agent and tool calls arrive as unrelated line items with nothing linking them to the run that caused them.
  • Human review time never appears at all, and in governed workflows it is frequently the largest single cost.

The reconstruction trap

Teams try to rebuild attribution from timestamps and log correlation after the first surprising bill. It produces an estimate that is confidently wrong and takes weeks. Instrument at call time or accept that you cannot answer the question.

The six attribution keys

Attach all six to every model call and every tool call. They are cheap to carry and impossible to add later.

KeyWhat it enables
Run IDJoins retries, sub-agents and tool calls into one outcome
Agent identityPer-agent cost, and therefore per-agent accountability
Acting user / teamChargeback and showback to an owner who can act
WorkflowUnit economics per business process, not per model
EnvironmentKeeps development experiments out of production trends
Outcome referenceLinks spend to the invoice, ticket or record it produced

Scroll the table horizontally on narrow screens.

The last one is what most instrumentation omits and what makes the data valuable. Cost per run is an engineering metric. Cost per resolved invoice is a business metric, and only an outcome reference bridges them.

One outcome, attributed

run_id:      run_8f2c41
agent:       agent.finance.ap
team:        accounts-payable
workflow:    invoice-to-pay
environment: production
outcome_ref: invoice/INV-20418

model calls:      14   (incl. 3 on retry)
tool calls:        9
retrieval:         2
human review:  4 min   (1 approval)

Four costs teams miss

Cost 01

Retries and reruns

Non-deterministic workloads retry as a matter of course. If retries are not attributed to the same outcome, cost per outcome is understated by whatever the retry rate is — often materially.

Cost 02

Abandoned runs

Work that started, consumed tokens and produced nothing still cost money. Excluding failures flatters the numbers in exactly the situations where you most need honesty.

Cost 03

Human review minutes

Every approval consumes a person’s time at a real hourly cost. In governed workflows this frequently exceeds the model spend, and it is the cost that scales worst with volume.

Cost 04

Retrieval, storage and egress

Embeddings, vector storage, document processing and data transfer sit on different invoices and get attributed to “infrastructure” rather than to the agent that caused them.

Together these routinely account for a large fraction of true cost. A cost-per-outcome figure that omits them is not conservative — it is wrong in the optimistic direction, which is the worse direction for a budget decision.

Making it stick

Instrumentation decays unless it is enforced somewhere structural. Four things that keep it alive.

Enforce at the gateway

If the layer that makes model and tool calls rejects calls without attribution keys, coverage stays at 100%. Voluntary tagging trends to zero.

Use FOCUS field names

The FinOps Open Cost and Usage Specification already names most of these concepts. Reusing its vocabulary lets agent spend join your cloud cost reporting.

Emit with OpenTelemetry conventions

The GenAI semantic conventions cover model, token and tool-call attributes. Custom names guarantee a second, incompatible pipeline.

Report to owners, monthly

Cost data nobody receives changes no behaviour. Per-team, per-workflow reports with a named recipient are what turn measurement into savings.

Once attribution exists, the two things it unlocks are unit economics — see cost per outcome — and enforceable limits, covered in setting agent spend ceilings.

Diagram · free to reuse with attribution

The four layers cost has to roll up through

Consumption is what the invoice shows. Outcome is what a budget owner decides with. The layers between are where attribution either exists or does not.

Original diagram by Real Biz Digital. Reuse it anywhere with a link back to this article.

The product

Attribution, optimisation and cost governance

Barzel FinOps Atlas carries 30 tools, 21 resources and 16 prompts for spend attribution, evidence lineage and cost governance — with a free sandbox at 100 calls a month.

Key terms

Attribution key
A tag attached at call time that links consumption to an accountable owner, workflow or outcome.
Run ID
An identifier joining every model call, tool call and retry belonging to a single agent execution.
Showback
Reporting costs to the team that caused them without transferring budget.
Chargeback
Transferring the cost of consumption to the causing team’s budget.
FOCUS
The FinOps Open Cost and Usage Specification, an open schema for normalising cost and usage data.

Frequently asked questions

Why can I not just use the provider invoice?

An invoice records consumption accurately and causation not at all. It knows tokens, models and API keys; it does not know your teams, workflows, customers or outcomes, so it cannot answer who caused the spend or whether it was worthwhile.

What should be tagged on every AI call?

Six keys: run ID, agent identity, acting user or team, workflow, environment and outcome reference. The run ID is the most important, because it is what joins retries and sub-calls into one attributable outcome.

Which AI costs are most often missed?

Retries, abandoned runs, human review minutes, and retrieval or storage and egress. Together they are frequently a large share of true cost, and omitting them understates cost per outcome.

Should human review time be counted as AI cost?

Yes. It is a direct cost of operating the workflow, it is often larger than model spend in governed processes, and it scales worse with volume — which makes it the most decision-relevant number you can capture.

How do we keep instrumentation complete?

Enforce it at the layer that makes the calls: reject calls lacking attribution keys. Voluntary tagging degrades quickly, and partial coverage produces numbers that are worse than none because they look authoritative.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.

  1. 01 · FinOps Foundation FinOps Framework ↗ Capability model and phases referenced throughout.
  2. 02 · FinOps Foundation FOCUS — FinOps Open Cost and Usage Specification ↗ Open schema for normalising cost and usage data across providers.
  3. 03 · FinOps Foundation FinOps for AI overview ↗ Where AI workloads change the practice.
  4. 04 · CNCF OpenTelemetry semantic conventions for GenAI ↗ Agreed attribute names for model, token and tool-call telemetry.
  5. 05 · CNCF OpenTelemetry ↗ Standard for the traces and spans an agent execution record should emit.

Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.

Reference documentation

Want the specification rather than the argument?

Cash, variance and reporting tools, and the evidence chain behind every figure.

Free · 500 calls/mo · on the MCPize marketplace

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.