Financial Operations · Controls
Policy Exception Detection: Finding the Rules That Were Actually Broken
A risk score says a transaction is unusual. An exception says a rule was broken. The second is a much stronger statement, and most finance functions cannot make it because their policies were never written in a form a system can test.
The short answer
Policy exception detection tests every transaction against explicitly stated financial policies — approval thresholds, supplier requirements, expense limits, segregation rules, documentation requirements — and reports each definite breach with the rule cited, the actor and the deviation. Unlike risk scoring, which flags the unusual, exception detection asserts that a specific rule was broken. The prerequisite is uncomfortable: most financial policies are written for humans to interpret rather than for systems to test, and rewriting them is the actual work.
Summary for readers and answer engines
Reviewed 25 Aug 2026
- ▸An exception is a definite breach of a stated rule. A risk score is a probability that something is worth looking at. They are different instruments for different jobs.
- ▸Most policies cannot be tested as written. “Expenses must be reasonable” has no machine-testable form; “meals above £50 per head require prior approval” does.
- ▸Six classes cover almost everything: threshold, authority, documentation, segregation, supplier and timing.
- ▸Exceptions need a resolution workflow, not just a report. An exception with no owner and no resolution state becomes a permanent line in a spreadsheet.
- ▸A zero exception rate almost always means the rules are not being tested, not that nothing was breached.
Source: Mark Alex, Real Biz Digital — Policy Exception Detection: Finding the Rules That Were Actually Broken (https://realbizdigital.net/insights/policy-exception-detection/). Reproduce with attribution.
Key takeaways
- 01Rewrite policies into testable form before building any detection. This is the project; the detection is the easy part.
- 02Separate a breach from a deviation with prior approval. The second is a documented exception, not a control failure, and conflating them destroys the metric.
- 03Cite the specific rule and the specific deviation in every exception. ‘Policy breach’ is not actionable.
- 04Give every exception an owner and a resolution state. Detection without resolution is documentation of a weakness.
- 05Age exceptions and escalate on age, not on count. Old exceptions are the finding.
- 06Track the exception rate per policy. A rule with a very high rate is usually a bad rule rather than widespread misconduct.
Quick answers
One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.
- What is a policy exception?
- A definite, demonstrable breach of an explicitly stated financial policy, reported with the rule cited, the deviation quantified and the actor identified.
- How does it differ from risk scoring?
- A risk score says an item is unusual and worth review. An exception says a specific stated rule was broken. One is probabilistic, the other is a factual assertion.
- Why can most policies not be tested?
- Because they were written for human interpretation. ‘Reasonable’, ‘appropriate’ and ‘where possible’ have no machine-testable form.
- What are the exception classes?
- Six: threshold breach, authority breach, documentation missing, segregation failure, supplier non-compliance and timing violation.
- Is a pre-approved deviation an exception?
- It is a documented exception rather than a control failure, and it must be counted separately or the metric becomes meaningless.
- What does a zero exception rate mean?
- Almost always that the rules are not being tested. Real populations produce exceptions; a zero rate indicates a detection gap rather than perfect compliance.
- What makes an exception actionable?
- The specific rule cited, the deviation quantified, the actor named, an owner assigned and a resolution state tracked with an age.
Exception versus risk score
Both surface transactions for attention. What they assert is completely different, and using one where you need the other causes real friction.
The practical consequence is in how people receive them. Telling a manager their expense claim scored highly for risk invites a shrug; telling them it exceeded the stated per-head limit by £34 without prior approval invites either a correction or a legitimate explanation. Both are useful; only the second is a control.
The two also fail differently. A risk score with poor precision wastes review time. Exception detection with badly written rules generates disputes and erodes trust in the whole control environment, which is a more expensive failure.
- ›Probabilistic: this looks unusual
- ›No rule cited
- ›Most flagged items are legitimate
- ›Orders review effort
- ›Cannot be disputed — it is a priority, not a claim
- ›Threshold set by review capacity
- ›Factual: this rule was broken
- ›Specific rule cited with the deviation
- ›Every exception is a real breach or a documented deviation
- ›Requires resolution
- ›Can be disputed — and sometimes correctly
- ›Threshold set by the rule itself
Run both. Score to prioritise attention, test policies to assert breaches, and never present a score as though it were an exception.
Rewriting policies into testable form
This is the work. Most financial policies are unfalsifiable as written, and no amount of technology fixes that.
- 01Every testable rule needs a subject, a condition, a threshold and a consequence. Missing any one makes it untestable.
- 02Replace every instance of ‘reasonable’, ‘appropriate’, ‘significant’ and ‘where possible’ with a number or a list. That sentence is most of the rewriting work.
- 03State the approval requirement, not just the limit. A threshold with no stated consequence produces exceptions nobody knows how to resolve.
- 04Include the prior-approval path in the rule. If exceeding the limit is permitted with approval, that is part of the policy rather than a workaround.
- 05Version policies and record which version a transaction was tested against. Policies change; exceptions must be judged against the rule in force at the time.
- 06Get the rewrite approved by whoever owns the policy. A testable rule written by a systems team and never ratified will be disputed at the first exception.
| As written | Testable form |
|---|---|
| Expenses must be reasonable and necessary | Meals: max £50 per head. Hotels: max £200 per night in listed cities, £150 elsewhere. Above either requires prior written approval. |
| Purchases should use approved suppliers where possible | Purchases above £1,000 must be to a supplier in the approved master. Below £1,000, up to £5,000 per supplier per year is permitted outside it. |
| Significant transactions require appropriate approval | Approval required at: £5k department head, £25k finance director, £100k CFO, £500k board. ‘Significant’ is not a testable term. |
| Journals should be supported by documentation | Manual journals above £10,000 require an attached supporting document, captured before posting. |
| Duties should be appropriately segregated | No individual may both prepare and approve the same journal, payment or reconciliation. No exceptions without a documented compensating control. |
| Invoices should be paid on time | Payment initiated no later than the due date; early payment beyond 10 days before due requires approval. |
Expect the rewrite to surface genuine disagreement about what the policy actually is. That disagreement existed before; it was simply never forced into the open.
Six exception classes
Key facts
- ▸Authority and segregation breaches are the two classes external auditors treat most seriously, because both are structural rather than judgemental.
- ▸Documentation-missing exceptions overlap with evidence gaps. Detecting them once and reporting them into both processes avoids duplicate work.
- ▸Supplier non-compliance is the class where a pattern matters more than an instance. One payment outside the master is an oversight; forty is a procurement bypass.
| Class | Example | Severity | Typical resolution |
|---|---|---|---|
| Threshold breach | Expense above the per-head limit; payment above an approval band | Medium | Retrospective approval or recovery |
| Authority breach | Approved by someone without the required authority level | High | Re-approval by an authorised person; review of access |
| Documentation missing | Manual journal above threshold with no support attached | Medium | Obtain the document; escalate on age |
| Segregation failure | Same person prepared and approved | High | Independent re-review; access change |
| Supplier non-compliance | Payment to a supplier outside the approved master | Medium-high | Onboard the supplier or recover; check for a pattern |
| Timing violation | Journal posted after the closed period; early payment outside terms | Medium | Reverse and repost, or document the reason |
Classify at detection, not at review. An unclassified exception queue gets triaged by whoever picks it up, inconsistently.
Resolution workflow
Detected
Rule cited, deviation quantified, actor identified, class assigned, owner routed automatically. No human involvement to reach this state.
Under review
The owner is examining it. Requires an expected resolution date; without one this state absorbs everything indefinitely.
Resolved — corrected
The underlying transaction was fixed: reversed, re-approved, document attached, recovered. The preferred outcome.
Resolved — accepted with justification
The deviation stands, with a recorded reason and an authorising approver. Legitimate, and must be counted separately from corrected.
Resolved — policy amended
The rule was wrong. This is a valid and under-used outcome, and it requires the policy owner rather than the exception owner.
Escalated
Unresolved past its age threshold. Escalation is by age, not by count, and it goes to the owner’s manager.
- 01Route by class to a default owner automatically. Manual routing is where exception queues go to die.
- 02Require an expected resolution date on entering review. It converts an open state into a commitment.
- 03Distinguish corrected from accepted in all reporting. A function that accepts most of its exceptions has a policy problem, not a compliance problem.
- 04Make policy amendment a first-class outcome. Rules that generate constant exceptions are usually wrong, and pretending otherwise wastes everyone’s time.
- 05Escalate on age with a fixed threshold per class. High-severity classes should escalate in days, not weeks.
- 06Never close an exception without a state and a reason. ‘Closed’ with no category is the most common defect in exception registers.
Reading the exception rate
The counter-intuitive point is that a moderate exception rate is the healthy signal. It means rules exist, they are being tested, and the population contains the normal quantity of human error and edge cases. A zero rate means one of the two components is missing.
The most common real finding from exception detection is not misconduct. It is that a threshold set five years ago now catches routine transactions, and everybody has been quietly working around it.
- ›Zero exceptions — rules almost certainly not being tested
- ›One policy generating none, ever — the rule may be untestable as implemented
- ›Sudden drop to zero — a data feed broke, not behaviour changed
- ›All exceptions in one class — the other tests are not running
- ›One rule producing most exceptions — the rule is probably wrong
- ›High accepted-with-justification share — the policy does not match how work is done
- ›Same actor, same rule, repeatedly — a training or system-access issue
- ›Rate rising with no process change — a threshold has been overtaken by inflation
Report the rate per policy, with corrected and accepted split, and the age distribution. Those three views together tell you whether the problem is behaviour, rules, or detection.
Exceptions when an agent is the actor
Agent-initiated transactions change two things about exception handling, and both are worth planning for before the first one arrives.
- 01Volume and speed. An agent can breach the same rule four hundred times in an hour. Exception detection must aggregate by rule and actor rather than reporting four hundred individual items.
- 02Attribution. The actor is an agent, but accountability belongs to a person. Every agent-initiated exception needs the principal chain, or it resolves to a service account nobody owns.
- 03Pre-execution is better than detection. For agent actions the rule can be tested before the transaction happens. An exception detected afterwards on an agent-initiated payment is a recovery exercise that a policy check would have prevented.
- 04Threshold-adjacent behaviour. Agents operating just below a limit produce no exceptions and warrant attention. This is where risk scoring complements exception detection.
- 05Repeated identical exceptions signal a defect, not misconduct. An agent breaching the same rule consistently has been configured to, and the fix is upstream.
- 06Resolution must reach the configuration. Correcting the transaction without correcting the agent guarantees recurrence within days.
This is where exception detection and pre-execution policy enforcement meet: the same rules, tested before the action for agents and after the fact for humans, with the same classification and the same resolution workflow.
Next step
Test the rules, then resolve what breaks
Barzel FinOps Atlas detects policy exceptions and assesses control risk as callable tools; BarzelVault tests the same rules before an agent-initiated transaction executes.
Limits
Two.
- 01Exception detection cannot test a policy that is genuinely judgemental. Some policies legitimately require human interpretation, and forcing them into testable form produces false precision.
- 02A clean exception report does not mean controls are effective. It means the tested rules were not breached, which is a narrower claim than it appears and should be stated as such.
Frequently asked questions
What is a policy exception in finance?
A definite, demonstrable breach of an explicitly stated financial policy, reported with the specific rule cited, the deviation quantified, the actor identified and a class assigned — as distinct from a risk score, which only indicates that something is unusual.
How does exception detection differ from risk scoring?
Exception detection asserts that a specific stated rule was broken; risk scoring assigns a probability that an item is worth reviewing. One is a factual claim that can be disputed on its merits, the other is a prioritisation that cannot.
Why can most financial policies not be tested automatically?
Because they were written for human interpretation. Terms like reasonable, appropriate, significant and where possible have no machine-testable form, so the rewriting of policy into testable rules is the actual project.
What does a testable policy rule need?
Four elements: a subject, a condition, a threshold and a consequence. A limit stated without its approval path produces exceptions nobody knows how to resolve, and a condition without a threshold cannot be evaluated.
What are the six exception classes?
Threshold breach, authority breach, documentation missing, segregation failure, supplier non-compliance and timing violation. Authority and segregation are the two external auditors treat most seriously because both are structural rather than judgemental.
Is a deviation with prior approval an exception?
It is a documented exception rather than a control failure, and it must be counted separately. Conflating pre-approved deviations with genuine breaches makes the exception rate meaningless in both directions.
What resolution states does an exception register need?
Six: detected, under review with an expected date, resolved by correction, resolved as accepted with justification, resolved by policy amendment, and escalated on age. Closing an exception without a state and reason is the most common register defect.
Why is policy amendment a valid resolution?
Because a rule generating constant exceptions is usually wrong rather than universally flouted. The most common genuine finding from exception detection is a threshold set years ago that inflation has turned into a routine trap.
What does a zero exception rate mean?
Almost always that the rules are not being tested rather than that nothing was breached. Real transaction populations contain human error and edge cases, so a moderate rate is the healthy signal and zero indicates a detection gap.
How should exceptions be escalated?
By age against a per-class threshold rather than by count. High-severity classes such as authority and segregation breaches should escalate within days; a register full of old exceptions is itself the finding.
How does exception handling change when an agent is the actor?
Volume and speed require aggregation by rule and actor rather than individual reporting, the principal chain must be carried so accountability reaches a person, and rules should be tested before execution rather than detected afterwards.
Does a clean exception report mean controls are effective?
No. It means the tested rules were not breached, which is narrower than it sounds. Untestable policies, untested rules and well-formed transactions that comply while achieving a wrong outcome all sit outside what an exception report can assert.
Glossary
- Policy exception
- A demonstrable breach of an explicitly stated financial policy rule.
- Testable rule
- A policy expressed with a subject, condition, threshold and consequence that a system can evaluate.
- Authority breach
- Approval given by someone without the required authority level for the amount or action.
- Segregation failure
- The same actor performing two roles that policy requires to be separate.
- Accepted with justification
- An exception allowed to stand with a recorded reason and an authorising approver.
- Policy amendment
- Resolving an exception by correcting the rule rather than the transaction.
- Exception ageing
- Time since detection, used as the escalation trigger rather than exception count.
- Supplier master
- The approved list of suppliers, outside which payments constitute an exception.
- Threshold-adjacent behaviour
- Operating just below a limit, producing no exception while warranting attention.
- Principal chain
- The human accountability path behind an agent-initiated transaction.
Standards and entities referenced
Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.
- 01 · COSOCOSO Internal Control — Integrated Framework ↗The control framework auditors map financial process evidence against.
- 02 · ISACACOBIT 2019 Framework ↗Governance and management objectives, including segregation of duties.
- 03 · U.S. SECSarbanes-Oxley Act — Section 404 ↗Where segregation of duties becomes an externally audited control.
- 04 · Institute of Internal AuditorsInternational Standards for the Professional Practice of Internal Auditing ↗What internal audit is required to evidence, and the independence expectations around it.
- 05 · PCAOBPCAOB AS 1105 — Audit Evidence ↗The standard defining sufficiency, appropriateness, relevance and reliability of audit evidence.
- 06 · PCI Security Standards CouncilPCI DSS v4.0 ↗Where payment-adjacent agent actions inherit real, externally audited requirements.
- 07 · OWASP GenAI Security ProjectOWASP Agentic AI — Threats and Mitigations ↗Threat taxonomy specific to tool-using agents rather than to chat completions.
- 08 · NISTNIST — AI Agent Standards Initiative ↗Identity, authorization, auditing and non-repudiation framed as prerequisites for autonomous agents.
Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). Policy Exception Detection: Finding the Rules That Were Actually Broken. Real Biz Digital. https://realbizdigital.net/insights/policy-exception-detection/
Try the mechanics on a live server
To watch an MCP server answer a structured request before you let one read your ledger — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
Barzel FinOps Atlas is this assurance layer, sold as a running product
Thirty tools covering close readiness and blocker detection, cash position, cash variance and cash-flow risk, transaction risk scoring, policy exception detection, control risk, finance approvals, and the evidence surface — evidence-to-control mapping, evidence graphs, evidence tracing, chain verification, missing-evidence detection, auditor request answering and audit packet generation. A free sandbox tier means the first readiness report costs nothing.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Free Sandbox | Free | 500 calls/mo · close readiness, blockers, evidence checks | Testing readiness scoring against one real close |
| Starter | $29/mo | 1,000 calls/mo · evidence mapping, cash position, approvals | A single entity running one governed close cycle |
| Growth | $99/mo | 5,000 calls/mo · evidence graph, audit packets, control risk | A controller’s team with an external audit each year |
| Business | $249/mo | 15,000 calls/mo · the full 30-tool surface | Multi-entity close with SOX obligations and continuous audit readiness |
| Enterprise | $799/mo | 50,000 calls/mo · everything in Business, scaled | Group-wide finance operations across many entities |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
The five Barzel servers, and which problem each one is sold for
One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.
| Server | Sold for | Entry price | Where it sits |
|---|---|---|---|
| Barzel Central Gateway | Knowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidence | Free, then $10–$149/mo | Control plane — decides what may be reached, and by whom |
| BarzelVault | Stopping a specific dangerous action before it executes, with proof afterwards | $199–$3,999/mo | Decision point — evaluates the individual call before execution |
| BarzelOps | Running real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approval | Free, then $19–$199/mo | Execution layer — does the work the policy allowed |
| Barzel FinOps Atlas | Attributing AI spend to agents, tools and outcomes, then forecasting and capping it | Free, then $29–$799/mo | Economics layer — what the estate costs per outcome |
| Barzel Scripture Intelligence | A free, credential-free public MCP server to test clients and inspect real protocol traffic | Free, unmetered, no signup | Reference implementation — safe place to learn the protocol |
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.