When financial institutions discuss AI in their transaction monitoring programmes, two architecturally distinct things tend to be described in the same words. The first is detection improvement: the model finds suspicious activity that the existing rule set would not have flagged. The second is alert management: the rule set is unchanged, and the model scores, routes, hibernates, and escalates the alerts the rules generate.
Both can produce measurable value. Both are deployed in production. But they sit in different parts of the workflow, need different evidence to validate, require different governance, and should be measured against different benchmarks. Conflating them is routine – in vendor presentations, in industry discussion, in internal proposals. It produces misjudged procurement, the wrong governance frame, and metrics that measure the wrong thing.
This post draws the distinction precisely, examines what the published evidence actually shows for each, and explains what the difference means for teams evaluating or deploying AI in transaction monitoring.
What detection improvement means
In a rule-based transaction monitoring system, the rules define the detection logic. A rule fires when a specific pattern, threshold, or sequence of conditions is met. The total population of alerts is determined by the rules: if a suspicious pattern is not described by any rule, it generates no alert and reaches no investigator.
Detection improvement means the model finds activity that the rules would not have flagged. It might surface behavioural anomalies that are statistically unusual for a given customer but match no explicit rule, identify relationships across transactions and entities that the rule language cannot efficiently express, or pick up patterns that were not known when the rules were written.
This is the stronger claim, and it needs the stronger evidence base. A serious detection-improvement claim should be able to demonstrate one of:
- more confirmed cases of money laundering or terrorist financing identified at the same or lower false-positive rate;
- equivalent or higher true-positive retention at a substantially lower false-positive rate;
- comparison against a well-tuned baseline – not against an untuned rule engine that was never properly segmented or calibrated;
- and, ideally, identification of cases or patterns the rule set genuinely did not pre-specify.
Producing that evidence publicly is hard. Institutions cannot disclose case-level data, and externally auditable detection evidence needs validated ground truth that most AML environments do not have. That is why the public evidence base for production detection-improvement claims is thinner than vendor decks suggest.
When a vendor presents a large false-positive reduction as evidence of detection improvement, the right first question is whether the model changed what gets detected, or whether it changed what happens to alerts after the rules generated them. The answer determines what the number actually means.
What alert management means
Alert management means the rule set is unchanged and the model changes what happens to the alerts the rules produce. The same rules generate the same alert population. The model sits above that population and performs a different function: it scores each alert for the probability that it represents genuine suspicious activity, and uses those scores to change how alerts move through the workflow.
Three routing paths are typical:
- Low-scored alerts are hibernated or deprioritised. They remain accessible but are not surfaced to investigators in the standard queue.
- Medium-scored alerts are routed through standard Level 1 review.
- High-scored alerts are escalated directly to Level 2, or flagged for priority attention.
The value here is prioritisation. The highest-risk alerts reach the most experienced investigators first. Cases that would eventually result in SARs are resolved faster. Analyst capacity is concentrated where it is most likely to be needed. None of the cases the rules would have caught are missed; they are simply ordered.
That is alert management. It is not detection improvement. The model has not found any case the rules would have missed; it has made the existing case population more navigable.
The distinction matters for governance. If a low-scored alert is hibernated and the underlying activity later turns out to be part of a financial-crime pattern, the institution needs to be able to show that hibernation was a defensible decision, that the hibernated pool is monitored, and that a mandatory-review safeguard exists. The governance model for alert management – oversight of the hibernation logic, the audit trail, the escalation thresholds – is different from the governance model for a detection layer, and institutions deploying alert-management AI need to design for that explicitly.
Alert-management architecture: bolt-on ML scoring above an unchanged rule engine, three-tier routing – the model optimises how existing alerts are processed; it is not detection improvement.
What the published evidence shows
The public evidence base for AI in transaction monitoring is smaller than the volume of vendor claims would suggest. Most production deployments are not described in peer-reviewed literature, and the literature that does exist is more often technical than operational.
One study that makes the architecture explicit is Halford et al., Journal of Money Laundering Control, 2025 (CC-BY 4.0). The paper describes a bolt-on machine-learning scoring model deployed above an existing rule-based transaction monitoring system at one anonymised bank. The country and the institution are both anonymised in the paper. The rule engine was not modified; the ML model scored the existing alert population and routed alerts into three tiers based on those scores. The detection layer was unchanged. The pilot and rollout covered 11 months of continuous production data (April 2023 – February 2024), 117,494 alerts.
The reported results: 13.4 % of alerts hibernated, 15.3 % auto-escalated, about 57 % of SARs concentrated in the top 15 % of alerts, no missed SARs in the hibernated pool during the six-month pilot window, about 61 % alert-to-SAR turnaround reduction, and roughly 45.6 FTE saved in February 2024 alone. The cumulative figure of about 280 over the 11 months is the total across the ramp-up period, not an ongoing headcount reduction. All institution-specific and not benchmarks – but a reproducible blueprint for the alert-management pattern.
These are clearly alert-management results. The paper does not claim the model found cases the rules would have missed; it claims the model made the existing alert population substantially more efficient to process. The metrics – turnaround time, SAR concentration in high-scored alerts, behaviour of the hibernated pool – are the right benchmark for this class of AI deployment. Applying detection metrics to an alert-management system measures the wrong thing and produces misleading assessments of whether it is working.
Methodology counterparts on test data exist: Bakry et al., Journal of Supercomputing, 2024 reports ~73 % false-positive reduction at ~92 % recall on a 1,926-customer test set, and Jensen & Iosifidis, Expert Systems with Applications, 2023 reports >33 % false-positive reduction at ~98.8 % true-positive retention at Spar Nord. Both are also institution-specific results, and neither is a benchmark. Together with Halford, they describe the current published shape of the field: small, predominantly single-institution, mostly bolt-on alert-management — not detection improvement at scale.
Two operational consequences follow. First, a zero missed-SAR rate during a pilot period is meaningful but is not the steady state – institutions deploying alert-management AI need ongoing monitoring of the hibernated population, not just the pilot result, because typologies evolve. Second, alert-management evidence does not transfer cleanly across institutions: a 61 % turnaround improvement at one anonymised bank does not predict the same improvement at a payment institution with a different customer mix, product portfolio, and rule configuration.
The governance implications
The detection-vs-alert-management distinction is not only a measurement question. It determines the governance structure the AI deployment requires.
Detection-layer changes – modifying which rules fire, adding model-based detection above or alongside rules, or suppressing rules on the basis of model output – affect the institution’s compliance obligations directly. If the model causes the institution to stop detecting a category of activity it was previously detecting, that is a material change to the AML programme that needs internal validation, documented justification, and, depending on jurisdiction and materiality, regulatory consideration. Institutions should consult their compliance and legal teams before changing detection logic that may affect their reporting obligations.
Alert-management changes are different in character. They do not change what is detected; they change how detected activity is processed. The obligation to review and resolve alerts remains; the model assists with prioritisation and routing inside that review process. That said, alert-management governance is still real governance, not a light-touch operational question. The questions to design for, in roughly the order they will be asked of you, are:
- Who owns the model? Who validates it? Who monitors it in production?
- What is the hibernation policy – how long can a low-scored alert remain unreviewed before a mandatory escalation triggers?
- What is the false-negative monitoring mechanism – how does the institution know if the hibernated pool begins to accumulate genuine cases?
- What documentation does the model produce for audit purposes – scores, routing decisions, hibernation log, review timestamps?
- How does the model interact with the institution’s obligation to report suspicious activity within prescribed timeframes?
The EU AI Act (Regulation (EU) 2024/1689) brings obligations for high-risk AI systems into application from 2 August 2026. The Articles most relevant to AI in transaction monitoring are 9 (risk management), 10 (data and data governance), 11 (technical documentation), 13 (transparency), 14 (human oversight) including 14(5) (no autonomous decision on specified decisions), 15 (accuracy, robustness, cybersecurity), 26 (deployer obligations – the institution side), and 86 (right to an explanation of individual decisions). Whether a given AML AI system is high-risk under Annex III turns on Commission classification guidelines under Article 6(4), signalled to publish during 2026; until those land, the responsible posture is to scope readiness against the conservative reading.
FATF Recommendation 15 (New Technologies) sets the standard that obliged entities identify and assess the money-laundering and terrorist-financing risks that arise from new technologies – including AI and automated systems – before launch, and take appropriate measures to manage them. The FATF Methodology distinguishes technical compliance on paper from effectiveness in practice (the eleven Immediate Outcomes). Technology does not reduce the institution’s responsibility for the effectiveness of its controls – that is the test obliged entities are assessed against. A sophisticated scoring model does not satisfy the obligation to have effective transaction monitoring if the alert-management logic systematically deprioritises genuine suspicious activity; human oversight above the model’s outputs is not optional.
Governance requirements by AI type – detection-layer AI vs alert-management AI, same institution, different obligations.
Five questions before any transaction-monitoring AI deployment
Whether the AI under evaluation is detection-layer or alert-management, five questions help ensure the decision is grounded in operational reality rather than the vendor’s presentation.
- Did the rule set change, or did the model sit above an unchanged rule engine? This is the most direct way to know which lane you are in. If the rules are unchanged, the AI is doing alert management; evaluate accordingly.
- What was the baseline configuration? A well-tuned, segmented, suppression-aware rule engine is a very different baseline from an untuned default. If the vendor cannot describe the baseline precisely, the headline improvement number is not yet useful.
- What is the hibernation safeguard? For alert management specifically: how long can a low-scored alert remain unreviewed; what triggers a mandatory escalation; how is the hibernated pool monitored for accumulating genuine cases? This is the governance point most often left vague in vendor materials.
- What metric proves value? Match the metric to the workflow stage. Alert-to-SAR turnaround, SAR concentration in the top-scored band, and missed-SAR rate in the hibernated pool for alert management. True-positive rate at a constant false-positive rate, against a strong tuned baseline, for detection. Detection metrics do not evaluate alert management, and vice versa.
- Where does human review remain? AI in transaction monitoring should make human investigators more effective, not remove them from the process. The points at which human review is required – and the documentation that review generates – are governance obligations where the system is in scope of AI Act Article 14, not optional enhancements.

Five questions before any transaction-monitoring AI deployment – match the evidence to the workflow stage before committing.
The distinction between detection improvement and alert management is not a technicality. It determines the appropriate benchmark, the governance model, the regulatory framing, and the operational promise an institution can reasonably make about its AML programme. The value AI brings to transaction monitoring is real; the discipline is in naming it precisely. An alert-management system that sharply reduces SAR turnaround and concentrates SARs in the highest-risk alerts is a strong result – and it deserves to be evaluated for what it is, not as though it were doing something else.
Read related blog articles:
AI in AML: Three European Events That Shape the Landscape
AML AI Software: The Questions to Ask Before You Buy
Sources & notes
- Halford et al., Journal of Money Laundering Control, 2025 – DOI 10.1108/JMLC-09-2024-0152 (CC-BY 4.0). One anonymised bank; country and institution both anonymised in the paper. Reported figures are institution-specific; they are not benchmarks.
- Bakry et al., Journal of Supercomputing, 2024 – DOI 10.1007/s11227-023-05708-z. Test-environment results; institution-specific.
- Jensen & Iosifidis, Expert Systems with Applications, 2023 – DOI 10.1016/j.eswa.2022.119037. Single-institution (Spar Nord) results; institution-specific.
- EU AI Act — Regulation (EU) 2024/1689 – EUR-Lex. High-risk obligations apply from 2 August 2026; Annex III classification is conditional on Commission guidelines under Article 6(4), signalled to publish during 2026.
- FATF Recommendation 15 (New Technologies) – fatf-gafi.org.
- FATF Methodology for Assessing Technical Compliance with the FATF Recommendations and the Effectiveness of AML/CFT Systems – fatf-gafi.org. The eleven Immediate Outcomes are the effectiveness lens referenced above.
AMLYZE builds transaction monitoring, customer screening, and AML case management software for regulated financial institutions: configurable rule catalogue, no-code rule design, real-time alert processing. Get a demo here.







