5 MCP servers live now What’s live ›
Real Biz Digital logo Real Biz Digital

AI Agent Governance · Audit

How to Audit AI Agent Actions

Application logs tell you an agent ran. An audit trail tells you what it decided to do, whether it was allowed, who permitted it, and what changed. Those are not the same file.

By Mark Alex, Founder Published 19 Aug 2026 10 min read

Key takeaways

  • Log attempts, not completions. A record that only contains successful calls cannot show you a blocked attack or a policy that is too tight.
  • Parameter values are the audit. “Called export_records” is not an audit entry; “called export_records with limit 50,000 to an external address” is.
  • One stream, one schema. Per-server logs in different shapes cannot be joined under time pressure, which is exactly when you need them joined.
  • Hash-chain it. An audit trail that the system under audit can silently edit is evidence of nothing.
  • Prompts and completions are context, not the record. Store them separately with a shorter retention and a redaction policy — they are the highest-risk data you will hold.

The short answer

An agent audit trail records every attempted action — not every completed one — with caller, tool, parameter values, destination, policy decision, reason, approver and result, in one append-only stream. Denials matter as much as executions, because they are your earliest signal that something is steering the agent. If reconstructing an incident needs three log formats and a chat transcript, you have archaeology, not an audit trail.

Why application logs fail

Most teams discover their agent has no audit trail during their first real incident. The logs exist and they are useless, for four structural reasons.

They record effects, not decisions

A row changed. Nothing says which agent proposed the change, under whose authority, or whether anything evaluated it first.

They omit denials entirely

Blocked calls never reach the target system, so the target system never logs them — losing precisely the records that reveal an attack in progress.

Values are dropped or truncated

Logging the tool name and dropping the arguments removes the only part that distinguishes routine work from bulk extraction.

Attribution collapses to a service account

When one credential fronts every caller, every log line names the same actor and the question “who did this?” has no answer.

The test

Pick a record your agent changed last week. Can you name, in under five minutes and from one source, which agent changed it, on whose behalf, what it passed, what evaluated the call, and what the decision was? If not, the trail does not exist yet.

The nine fields

Minimum viable audit event. Anything less and one of the four questions below becomes unanswerable.

FieldContentsWhy it is not optional
Event ID + chainUnique ID and hash of the previous eventMakes silent deletion or reordering detectable
TimestampUTC, with the source clockOrdering across services is otherwise guesswork
Caller identityAgent identity, plus the user it acts forPer-agent revocation and attribution depend on it
Session / trace IDCorrelates every call in one runTrajectory analysis is impossible without it
ToolServer and tool name, with versionTool definitions drift; the version is part of the fact
ParametersThe actual values, redacted by ruleThe values are where the consequence lives
DestinationWhere output was sentMost damaging incidents end with data leaving
Decision + reasonOutcome, and a machine-readable reasonA denial without a reason cannot be triaged or tuned
ResultSuccess, error, records affected, approverCloses the loop from intent to effect

Scroll the table horizontally on narrow screens.

Use the OpenTelemetry GenAI semantic conventions for attribute names rather than inventing your own. Shared names are what let this data join to the rest of your observability stack instead of becoming a second silo.

Four questions it must answer

Design the record backwards from the questions somebody will actually ask under pressure.

  1. 01“Why did this record change?”A customer or controller asks. You need the call, the caller, the authority and the approver — in one place, without a reconstruction project.
  2. 02“Has this happened before?”After one bad action, the immediate question is scope. That requires querying by tool and parameter shape across all history, not reading a log stream.
  3. 03“What did we stop?”Denials answer whether controls are working and whether anything has been probing them. Only your policy layer can produce these.
  4. 04“Was a human involved?”For consequential actions, auditors want the named approver and timestamp attached to the action — not a ticket in another system that mentions it.

The failure worth naming

Successful indirect prompt injection looks exactly like authorised activity: valid credentials, in-scope tools, schema-valid parameters. Anomaly detection therefore has to run over tool calls and their values, not over prompts. See indirect prompt injection explained.

Making it hard to edit

An audit trail the audited system can rewrite proves nothing. Four measures, in increasing order of effort.

Level 01

Append-only storage

No update or delete path from application code. Writes go one way.

Level 02

Hash chaining

Each event includes a hash of its predecessor, so removing or altering an entry breaks the chain verifiably.

Level 03

Separate write path

The policy layer writes the audit; the agent cannot. If the same credential can act and edit the record of acting, neither is trustworthy.

Level 04

Periodic external anchoring

Publish chain checkpoints somewhere you do not control. This is what turns “we believe the log is intact” into something demonstrable.

NIST’s audit and accountability control family in SP 800-53 and the log-management guidance in SP 800-92 both predate agents and both apply unchanged — the novelty is the volume and the non-determinism, not the requirement.

What not to keep

The instinct after reading the above is to log everything forever. That creates a different incident.

Keep, long retention
  • ›Action events with the nine fields
  • ›Policy decisions and reasons
  • ›Approvals with named approver
  • ›Chain checkpoints
Separate, short retention
  • ›Full prompts and completions
  • ›Retrieved document contents
  • ›Raw tool response bodies
  • ›Anything the redaction rules flagged

Prompts and retrieved content are the highest-risk data an agent system holds: unstructured, likely to contain customer information, and not obviously covered by the data map somebody wrote before agents existed. Store them apart from the audit trail, with their own retention and access controls, and keep the audit trail itself lean enough to keep for years.

Diagram · free to reuse with attribution

Where the audit event is produced

The record is written by the policy layer, at the moment of decision — which is why it can contain denials at all. The agent has no write path to it.

Original diagram by Real Biz Digital. Reuse it anywhere with a link back to this article.

The product

A tamper-evident record of every action, including the denials

BarzelVault emits hash-chained audit events and signed receipts for every evaluated call, with compliance reporting mapped to SOC 2, GDPR, ISO 27001 and OWASP MCP.

Key terms

Audit event
An immutable record of one attempted action, containing the caller, tool, parameters, decision, reason and result.
Hash chaining
Including a hash of the previous record in each new one, so any alteration or deletion becomes detectable.
Append-only
Storage that permits writes but no updates or deletes from application code.
Trace ID
An identifier correlating every call made during one agent run, enabling session-level analysis.
Signed receipt
A cryptographically signed confirmation that a specific action was evaluated and executed as recorded.

Frequently asked questions

What should an AI agent audit log contain?

Nine fields per attempted action: event ID with chain hash, UTC timestamp, caller identity including the user acted for, session or trace ID, tool with version, actual parameter values under a redaction rule, destination, policy decision with a machine-readable reason, and result including any approver.

Should denied actions be logged?

Yes, and they are often the most valuable records. Denials never reach the target system, so only your policy layer can capture them — and a blocked out-of-scope call is the earliest signal that content is steering the agent.

Are application logs enough?

No. They record effects rather than decisions, omit denials entirely, frequently drop parameter values, and attribute everything to a shared service account. All four gaps sit exactly where an investigation needs detail.

How long should agent audit records be kept?

Keep the lean action events for as long as your regulatory and contractual obligations require, typically years. Keep full prompts, completions and retrieved content separately and briefly — they are the highest-risk data in the system.

Do we need to log prompts and completions?

For debugging and evaluation, yes; as part of the audit trail, no. They are unstructured, frequently contain personal data, and mixing them into a long-retention audit store creates a privacy problem without improving accountability.

Sources and further reading

Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.

  1. 01 · NIST SP 800-53: Security and Privacy Controls ↗ The audit and accountability control family (AU) referenced here.
  2. 02 · NIST SP 800-92: Guide to Computer Security Log Management ↗ Baseline expectations for log content, retention and integrity.
  3. 03 · CNCF OpenTelemetry semantic conventions for GenAI ↗ Agreed attribute names for model, token and tool-call telemetry.
  4. 04 · CNCF OpenTelemetry ↗ Standard for the traces and spans an agent execution record should emit.
  5. 05 · ISO ISO/IEC 27001 — Information security management ↗ The ISMS baseline that agent controls have to fit inside.
  6. 06 · OWASP GenAI Security Project OWASP GenAI LLM Top 10 (2026) ↗ Current consensus list of LLM application risks, including excessive agency.

Last reviewed 2 September 2026. External links open in a new tab; we do not control their content.

Reference documentation

Want the specification rather than the argument?

What lands in a hash-chained audit event, what never does, and how to verify the chain yourself.

Dev free · 10,000 calls/mo · paid plans from $199/mo on the MCPize marketplace

Written by

Mark Alex

Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.