Financial Operations · Risk
Transaction Risk Scoring: Finding the Entries That Need a Human
Reviewing every journal entry is impossible and reviewing a random sample is nearly useless. Risk scoring is how you review the two percent that matters.
The short answer
Transaction risk scoring assigns each financial transaction a score from explicit weighted factors — amount relative to history, timing, preparer, account combination, round numbers, reversals, manual origin, and others — so that limited review capacity is directed at the highest-scoring few percent rather than at a random sample. Use transparent weighted rules rather than an opaque model, because the first question an auditor asks is why this transaction scored highly.
Summary for readers and answer engines
Reviewed 25 Aug 2026
- ▸Random sampling finds problems at the base rate. Risk scoring concentrates review on the few percent where the base rate is many times higher.
- ▸Twelve factors cover most of what matters, and they are all computable from ledger data you already have.
- ▸Use transparent weighted rules, not a black-box model. Explainability is not a nice-to-have when an auditor asks why an item was or was not reviewed.
- ▸Calibrate the threshold to your review capacity, not to a statistical cutoff. A score that flags more than your team can review is a score nobody uses.
- ▸Measure precision at the threshold. If reviewing the top two percent finds nothing, the weights are wrong.
Source: Mark Alex, Real Biz Digital — Transaction Risk Scoring: Finding the Entries That Need a Human (https://realbizdigital.net/insights/transaction-risk-scoring/). Reproduce with attribution.
Key takeaways
- 01Score everything, review the top slice. Scoring is cheap; review is the scarce resource.
- 02Weight manual origin heavily. Manual journal entries carry far more risk per item than system-generated ones.
- 03Combine factors multiplicatively for genuine compounds: a large manual entry posted after cut-off by an unusual preparer is not three separate concerns.
- 04Recalibrate quarterly against what review actually found. A score never revised against outcomes is a guess with a decimal point.
- 05Keep the score explainable per transaction. ‘Scored 84: unusual account pair, round amount, posted day 3 after cut-off’ is defensible; a number alone is not.
- 06Track precision and the finding rate. Those two numbers tell you whether the exercise is worth continuing.
Quick answers
One-line answers to the questions this page is most often asked. Each is expanded further down, and each is written to be quoted on its own.
- What is transaction risk scoring?
- Assigning each financial transaction a score from explicit weighted factors so that scarce review capacity is directed at the highest-risk items rather than a random sample.
- Why not review a random sample?
- Because a random sample finds problems at the population base rate. Scoring concentrates review where the base rate is many times higher, for the same effort.
- What factors matter?
- Amount relative to account history, manual versus system origin, timing relative to cut-off, preparer history, account pair unusualness, round numbers, reversals, weekend posting, and several others.
- Should this use a machine learning model?
- Generally not. Transparent weighted rules are explainable, auditable and adjustable, and the first question asked about any flagged or unflagged item is why.
- How is the threshold set?
- By review capacity. If your team can review forty items a month, the threshold is wherever the fortieth-highest score sits — not at a statistical cutoff.
- How do you know the score works?
- Precision at the threshold: what proportion of reviewed items had a genuine issue. If reviewing the top slice finds nothing, the weights need recalibration.
- Does this apply to payments as well as journals?
- Yes, with different factors. Payments add beneficiary novelty, bank detail changes and duplicate-signature checks.
Why sampling fails and scoring works
The arithmetic here is not subtle, and it is the entire justification.
- 01Sampling is designed for statistical inference about a population, not for finding specific problems. Auditors sample to conclude; controllers need to detect.
- 02Issues are not randomly distributed. They cluster in manual entries, in period-end postings, in unusual account combinations — all of which are computable.
- 03Review capacity is fixed and small. Any method that produces more candidates than capacity is not a method.
- 04Scoring everything is nearly free once the factors are defined. It is a query over data you already hold.
- 05The score orders work; it does not judge. A high score means look, not wrong.
- 06This is complementary to statutory sampling, not a replacement. Auditors will still sample; you should still score.
Detection arithmetic
population: 40,000 journal entries issues present: 60 (0.15% base rate) review capacity: 400 entries random sample: 400 × 0.15% ≈ 0.6 issues found risk-scored: 400 entries at ~8% precision ≈ 32 issues found
Same effort, roughly fifty times the yield. The precision figure is what the whole exercise turns on, and it is measurable.
The fifty-fold figure varies with base rate and precision, but the direction is robust: any concentration of review on higher-risk items beats uniform random selection.
Twelve factors with indicative weights
Key facts
- ▸Manual origin and preparer-equals-approver are the two factors that most reliably identify genuine issues, in our experience of calibrating these scores.
- ▸Round numbers are a weaker signal than folklore suggests in isolation, and a strong one in combination with manual origin and late posting.
- ▸Missing evidence is worth including even though it is not a transaction property, because it makes the score dual-purpose: risk and audit readiness.
| Factor | Weight | Rationale |
|---|---|---|
| Manual origin (not system-generated) | High | Manual entries carry far more risk per item; the single strongest factor |
| Amount vs account history (z-score) | High | An entry far outside an account’s normal range warrants attention |
| Posted after period cut-off | High | Late entries are disproportionately adjustments and errors |
| Unusual account pair | High | Debit-credit combinations rarely or never seen before in this entity |
| Preparer unusual for this account | Medium-high | Someone posting to an account they do not normally touch |
| Round number | Medium | Estimates and plugs are round; genuine transactions frequently are not |
| Reversal or reversing pair | Medium | Reversals are legitimate and cluster around error correction |
| Weekend or out-of-hours posting | Medium | Not inherently suspicious; correlates with rushed work |
| No supporting evidence attached | Medium | Missing support is both a risk factor and an evidence gap |
| Preparer equals approver | High | A segregation-of-duties failure and an independent finding |
| Description missing or generic | Low-medium | Weak descriptions correlate with weak substantiation |
| Related-party or intercompany account | Medium | Higher inherent risk and higher audit interest |
Weights should be indicative to start with and calibrated against outcomes within two quarters. Anyone who tells you the correct weights without seeing your data is guessing.
Why rules beat models here
The decisive consideration is not accuracy. It is that both a flagged item and an unflagged item will be questioned, and a controller must be able to explain the score in both directions. “The model gave it 0.34” is not an answer in a control environment.
There is a legitimate place for pattern detection: finding factor combinations you had not thought of, and surfacing them as candidate rules for a human to adopt. That keeps the discovery benefit and the explainability.
- ›Every score explainable per factor
- ›Auditable: why this scored 84
- ›Adjustable by a controller, not a data team
- ›Stable: same input, same score
- ›No training data required
- ›Defensible when an item was not flagged
- ›Score without attribution
- ›Hard to explain to an auditor
- ›Requires labelled training data
- ›Drifts with retraining
- ›Better at genuinely unknown patterns
- ›Difficult to defend a miss
A practical hybrid: rules produce the score that drives review; unsupervised analysis runs alongside to propose new rules quarterly. The score stays explainable and the factor set keeps improving.
Payments need different factors
- 01Bank detail changes deserve their own control, not just a score contribution. A change within thirty days of a payment should require independent verification, full stop.
- 02Threshold-adjacent amounts matter more in agent estates. An agent that learns a ceiling will operate just below it, and per-call limits cannot see the pattern.
- 03Duplicate detection needs a window and a reference, not just amount matching. Two legitimate identical payments do happen.
- 04New beneficiary plus urgency plus threshold-adjacent amount is the classic combination. Score it multiplicatively rather than additively.
- 05Check against the supplier master, not against payment history. History includes any fraud that already succeeded.
- 06Payments scored above the threshold should require verification before release, not review afterwards. Unlike journals, the money is gone.
| Factor | Weight | Why |
|---|---|---|
| New beneficiary (first payment) | High | The single most exploited vector in payment fraud |
| Beneficiary bank details changed recently | Highest | Classic mandate fraud signature; days since change matters |
| Amount just below an approval threshold | High | Deliberate threshold avoidance, whether by a person or an agent |
| Duplicate: same beneficiary, amount and window | High | Duplicate payment, frequently a retry artefact |
| Payment to a jurisdiction unusual for this supplier | Medium-high | Redirection or a legitimate change nobody recorded |
| Urgency flag or out-of-cycle request | Medium-high | Urgency is the standard social-engineering lever |
| Beneficiary not in the approved supplier master | High | Payments outside the master file bypass supplier controls |
The asymmetry matters: a journal entry can be reviewed after posting and corrected. A payment reviewed after release is a recovery exercise.
Calibrating the threshold
Key facts
- ▸Step five is the discipline that keeps scoring honest. Without a random control group you cannot know what the score misses.
- ▸Precision at threshold of five to fifteen percent is a working score in our experience. Below two percent, the weights need work.
- ▸Recalibrating continuously rather than quarterly produces a score that chases noise and cannot be compared period to period.
Establish review capacity honestly
How many transactions can be genuinely reviewed per period? Not how many can be ticked — genuinely reviewed. This is usually a much smaller number than people expect.
Set the threshold at that capacity
If capacity is forty items, the threshold is the fortieth-highest score. Everything else is a queue nobody works.
Review, record outcomes, measure precision
For each reviewed item, record whether a genuine issue was found. Precision at threshold is the number that matters.
Recalibrate weights against outcomes
Factors present in items that turned out fine should lose weight; factors present in genuine findings should gain it. Quarterly, not continuously.
Add a small random sample
Five to ten percent of review capacity spent on randomly selected items, to detect what the score is systematically missing. This is the check on the score itself.
The random sample is also what allows you to state a defensible position to an auditor: here is what the score found, and here is what an independent sample found, and the difference is the score’s contribution.
Measuring whether it works
- 01Precision at threshold. Share of reviewed items with a genuine issue. The headline number.
- 02Random-sample finding rate. Issues found per item in the random control group. The comparison that proves the score adds value.
- 03Lift. Precision at threshold divided by the random finding rate. Below three, the score is barely earning its keep.
- 04Factor contribution to findings. Which factors were present in items that turned out to be genuine issues. Drives recalibration.
- 05Review completion rate. Share of flagged items actually reviewed. Below eighty percent means the threshold is set above capacity.
- 06Time from posting to review. For payments this must be pre-release; for journals, within the period.
Report lift rather than precision alone. Precision without a comparison is uninterpretable, and lift is the number that answers whether the exercise is worth the effort.
Next step
Score the population, review the top slice
Barzel FinOps Atlas scores transaction risk, detects policy exceptions and assesses control risk as callable tools — explainable per factor, with a free sandbox tier.
Limits
Three.
- 01A high score is not an accusation. Most high-scoring transactions are legitimate, and treating the score as a verdict damages relationships with preparers quickly.
- 02The score cannot detect a well-formed fraudulent transaction that resembles normal activity. That is what segregation of duties and reconciliation exist for.
- 03Weights calibrated on one entity may not transfer to another with different volume, seasonality and account structure. Calibrate per entity where volumes allow.
Frequently asked questions
What is transaction risk scoring?
Assigning each financial transaction a score computed from explicit weighted factors — manual origin, amount relative to account history, timing, preparer, account pair unusualness and others — so that limited review capacity is directed at the highest-scoring items.
Why is risk scoring better than random sampling?
Because issues are not randomly distributed. A random sample finds problems at the population base rate, while scoring concentrates the same review effort where that rate is many times higher — in one worked example, roughly fifty times the yield for identical capacity.
Should transaction risk scoring use machine learning?
Generally not for the score that drives review. Transparent weighted rules are explainable per factor, auditable, adjustable by a controller, and defensible both when an item was flagged and when it was not — which is the requirement in a control environment.
Which risk factors matter most for journal entries?
Manual origin is the single strongest factor, followed by preparer equals approver, amount far outside an account’s normal range, posting after period cut-off, and unusual debit-credit account pairs.
Are round-number amounts a strong risk signal?
Weaker in isolation than folklore suggests, and genuinely strong in combination with manual origin and late posting. This is why factors should compound multiplicatively for real combinations rather than simply adding.
What additional factors apply to payments?
Recent beneficiary bank detail changes, first-time beneficiaries, amounts just below an approval threshold, duplicates within a window, unusual jurisdictions, urgency flags, and beneficiaries absent from the approved supplier master.
Why do bank detail changes deserve special treatment?
Because they are the classic mandate fraud signature. A payment to a beneficiary whose bank details changed within thirty days should require independent verification as a control, not merely contribute to a score.
How should the review threshold be set?
At genuine review capacity rather than a statistical cutoff. If a team can properly review forty items a month, the threshold is the fortieth-highest score; anything beyond that is a queue nobody works.
Why keep a random control sample?
Because it is the only way to know what the score systematically misses. Spending five to ten percent of review capacity on randomly selected items gives the comparison that proves the score adds value and reveals blind spots.
What is lift and what is a good value?
Precision at the review threshold divided by the finding rate in the random control sample. Below three the score is barely earning its keep; five to fifteen percent precision at threshold against a base rate of a fraction of a percent is a working score.
How often should weights be recalibrated?
Quarterly, against recorded review outcomes. Continuous recalibration chases noise and destroys period-to-period comparability, while never recalibrating leaves you with a guess that has a decimal point.
Is a high risk score an accusation?
No. Most high-scoring transactions are entirely legitimate — the score orders review effort rather than reaching a conclusion. Treating it as a verdict damages relationships with preparers quickly and makes the programme harder to sustain.
Glossary
- Transaction risk score
- A computed figure ordering transactions by the likelihood that review will find an issue.
- Base rate
- The proportion of a population containing genuine issues.
- Precision at threshold
- The share of reviewed items above the threshold that contained a genuine issue.
- Lift
- Precision at threshold divided by the random-sample finding rate.
- Manual origin
- A transaction entered by a person rather than generated by a system process.
- Account pair
- The debit-credit combination of an entry, scored by historical unusualness.
- Threshold-adjacent amount
- A value just below an approval limit, indicating possible avoidance.
- Mandate fraud
- Redirection of payment by altering a legitimate beneficiary’s bank details.
- Random control sample
- Randomly selected items reviewed alongside scored ones to detect what the score misses.
- Factor contribution
- How often a given factor was present in items that proved to be genuine issues.
Standards and entities referenced
Every named framework on this page resolves to a public definition. If you are checking our claims, start here rather than with us.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own operating judgement rather than something a standard states, the text says so.
- 01 · COSOCOSO Internal Control — Integrated Framework ↗The control framework auditors map financial process evidence against.
- 02 · PCAOBPCAOB AS 1105 — Audit Evidence ↗The standard defining sufficiency, appropriateness, relevance and reliability of audit evidence.
- 03 · U.S. SECSarbanes-Oxley Act — Section 404 ↗Where segregation of duties becomes an externally audited control.
- 04 · Institute of Internal AuditorsInternational Standards for the Professional Practice of Internal Auditing ↗What internal audit is required to evidence, and the independence expectations around it.
- 05 · PCI Security Standards CouncilPCI DSS v4.0 ↗Where payment-adjacent agent actions inherit real, externally audited requirements.
- 06 · MITREMITRE ATLAS ↗Adversary technique knowledge base for AI systems, useful for naming what a risk score is scoring.
- 07 · ISACACOBIT 2019 Framework ↗Governance and management objectives, including segregation of duties.
- 08 · TrintechTrintech — AI agents for financial close ↗Market reference: variance and flux agents with reviewer signoff and traceable evidence.
Last reviewed 2 September 2026 by Mark Alex. External links open in a new tab; we do not control their content.
Cite this article
Alex, M. (2026). Transaction Risk Scoring: Finding the Entries That Need a Human. Real Biz Digital. https://realbizdigital.net/insights/transaction-risk-scoring/
Try the mechanics on a live server
To watch an MCP server answer a structured request before you let one read your ledger — Barzel Scripture Intelligence is free and public at scripture-intelligence-server.mcpize.run: no signup, no key, 54 tools. Setup is in the reference.
Buy it on the marketplace
Barzel FinOps Atlas is this assurance layer, sold as a running product
Thirty tools covering close readiness and blocker detection, cash position, cash variance and cash-flow risk, transaction risk scoring, policy exception detection, control risk, finance approvals, and the evidence surface — evidence-to-control mapping, evidence graphs, evidence tracing, chain verification, missing-evidence detection, auditor request answering and audit packet generation. A free sandbox tier means the first readiness report costs nothing.
| Plan | Price | Included | Right for |
|---|---|---|---|
| Free Sandbox | Free | 500 calls/mo · close readiness, blockers, evidence checks | Testing readiness scoring against one real close |
| Starter | $29/mo | 1,000 calls/mo · evidence mapping, cash position, approvals | A single entity running one governed close cycle |
| Growth | $99/mo | 5,000 calls/mo · evidence graph, audit packets, control risk | A controller’s team with an external audit each year |
| Business | $249/mo | 15,000 calls/mo · the full 30-tool surface | Multi-entity close with SOX obligations and continuous audit readiness |
| Enterprise | $799/mo | 50,000 calls/mo · everything in Business, scaled | Group-wide finance operations across many entities |
Sold on the MCPize marketplace · prices as listed 2 Sep 2026 · the listing is authoritative
The five Barzel servers, and which problem each one is sold for
One estate rarely needs all five. This is the honest mapping, so you buy the layer your problem actually lives in.
| Server | Sold for | Entry price | Where it sits |
|---|---|---|---|
| Barzel Central Gateway | Knowing and governing the estate: inventory, registry, routing, risk scoring, approvals, evidence | Free, then $10–$149/mo | Control plane — decides what may be reached, and by whom |
| BarzelVault | Stopping a specific dangerous action before it executes, with proof afterwards | $199–$3,999/mo | Decision point — evaluates the individual call before execution |
| BarzelOps | Running real business workflows across HubSpot, Xero, Gmail, Drive and Slack under approval | Free, then $19–$199/mo | Execution layer — does the work the policy allowed |
| Barzel FinOps Atlas | Attributing AI spend to agents, tools and outcomes, then forecasting and capping it | Free, then $29–$799/mo | Economics layer — what the estate costs per outcome |
| Barzel Scripture Intelligence | A free, credential-free public MCP server to test clients and inspect real protocol traffic | Free, unmetered, no signup | Reference implementation — safe place to learn the protocol |
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.