Most AI-in-AML software vendor conversations begin with a familiar claim: our AI significantly reduces false positives compared with legacy systems. Sometimes that is true. Sometimes it is true but incomplete. Often it is an accurate number drawn from a misleading comparison.
The problem is not usually dishonesty. False-positive reduction is an output metric without a defined baseline. Without that baseline, the number says little about what the AI is doing, or whether it will do the same work inside your institution. Two questions, asked early, remove much of that ambiguity. They belong before metrics, before demos, and well before procurement.
The first question: compared with what?
The most important thing to establish about any AI performance claim in AML is the baseline: was the AI measured against an untuned legacy rule engine, or against a well-configured, segmented, suppression-aware one?
The distinction matters. A rule-based AML system running on default thresholds is a weak baseline. It generates large numbers of low-quality alerts because it has not been calibrated to the institution’s customer population, product mix, or transaction behaviour. Comparing AI with that baseline will almost always look favourable. Some of the improvement is the AI doing real work. Some is the AI receiving credit for work the rule engine could have done if anyone had configured it.
A well-built rule-based system is not a static threshold engine. Without any machine learning, it can already do the following:
- compare a customer’s activity against their own historical behaviour;
- compare activity against peer-cohort benchmarks segmented by customer type, product, geography, or risk band;
- apply different alert thresholds per segment rather than one universal threshold;
- detect sequences of activity over time, not individual transactions in isolation;
- suppress repeat alerts for activity that has been reviewed and dispositioned without a SAR;
- consolidate duplicate alerts across rule families, periods, and entities;
- apply sophisticated screening logic with alias handling, multi-script matching, date-of-birth variants, and configurable name-match similarity.
These are configuration decisions, not machine learning. If an AI product’s headline improvement was measured against a system that was not doing this work, some of that improvement was available without AI. The AI’s actual contribution can only be read against a baseline that already captured the rule-engine work.
AML software vendors that cannot describe their baseline precisely are not necessarily being deceptive. But the number they are presenting cannot be used as a reliable prediction of what their product will do for your institution, because your baseline is different from the one in their case study. Ask for the baseline description before evaluating the result.
What a well-tuned rule engine already does, without AI — and the single row where AI adds what rules cannot.
The second question: which value proposition?
After the baseline question comes the second: which lane is the AI in? AI in AML divides into four distinct value propositions, each sitting in a different part of the workflow, each with its own evidence requirements and metrics:
Detection improvement. The model changes what gets detected. Rule logic is amended or supplemented; suspicious activity that no rule would have flagged is identified by the AI. This is the strongest claim and needs the strongest evidence: true-positive rate at a constant false-positive rate, or false-positive rate at a constant true-positive rate, measured against a well-tuned baseline.
Alert management. The rule set is unchanged and continues to generate the same alert population. The AI scores each alert above that rule set, hibernates low-risk alerts, escalates high-risk alerts, and routes the middle differently. The detection layer is the same; the value comes from prioritisation, faster escalation, and freed analyst capacity.
The peer-reviewed reference for this pattern is Halford et al., Journal of Money Laundering Control, 2025 – an anonymised bank, 11 months of continuous production data (April 2023 – February 2024), 117,494 alerts. The reported results: 13.4 % of alerts hibernated, 15.3 % auto-escalated, about 57 % of SARs concentrated in the top 15 % of alerts, no missed SARs in the hibernated pool over the six-month pilot window, about 61 % reduction in alert-to-SAR turnaround, and roughly 45.6 FTE saved in February 2024 alone. The cumulative figure of about 280 over the 11 months is the total across the ramp-up period, not an ongoing headcount reduction. These are institution-specific results, not benchmarks – but a reproducible blueprint.
Alert management is a legitimate value proposition. It should, however, be evaluated as alert management, with the metrics that follow from that role: alert-to-SAR turnaround, SAR concentration in the top-scored band, missed-SAR rate in the hibernated pool, and analyst capacity released. Evaluating an alert-management product with detection-layer metrics either rejects products that would have improved operations or accepts products that will underperform against the wrong benchmark.
Investigation reasoning. AI assistance inside the investigation workflow – case summarisation, narrative drafting, document analysis, counterparty mapping. This operates after alerts have been escalated, not at the detection or triage layer. Its value is in investigation throughput, time per case, and documentation quality.
Knowledge structuring. Converting unstructured AML material – typology reports, SAR narratives, regulatory guidance, internal case notes – into structured, consistent, machine-readable outputs. This underpins detection training and rule engineering but is not itself a detection product. Its value lies in the consistency and portability of the labels it produces.
Each lane is valuable in the right context. None substitutes for the others. An AML software vendor presenting an alert-management result as evidence of detection improvement is measuring the wrong lane.
The four AI value propositions in the AML workflow – different positions, different evidence requirements, different metrics.
What that means for procurement
The practical consequence of not asking these two questions is that AI procurement decisions get made on numbers that do not reliably predict performance in the buyer’s environment. A reported false-positive reduction at a large retail bank, measured against an untuned legacy stack, does not translate predictably to a payment institution with a different customer mix, product portfolio, and rule configuration. The number describes a past result in a different context. It does not describe a future result in yours. Equally, strong alert-management results do not predict detection performance: a product that efficiently routes and hibernates alerts does nothing about whether the right alerts are being generated.
Good AI procurement in AML is not sceptical of AI. It is precise about what the AI is being asked to do, what evidence exists that it does that thing, and whether the evidence was generated in a context similar enough to the buyer’s to be predictive. The questions to ask an AML software vendor before signing should include:
- the baseline configuration against which improvement was measured;
- whether the model changed the detection logic, or operated above an unchanged rule engine;
- which workflow stage the model is in, and whether the metric matches that stage;
- the safeguards around hibernated or low-scored alerts;
- how human oversight is structured above the model and where override paths sit;
- the model-governance package the AML software vendor provides – technical documentation, audit trail, explainability;
- who owns model risk inside the institution once the product is deployed.
The last two are becoming regulatory questions, not just operational ones. The EU AI Act (Regulation (EU) 2024/1689) brings obligations for high-risk AI systems into application from 2 August 2026. The Articles most relevant to procurement readiness are 9 (risk management), 10 (data and data governance), 11 (technical documentation), 13 (transparency), 14 (human oversight), 14(5) (no autonomous decision on specified decisions), 15 (accuracy, robustness, cybersecurity), 26 (deployer obligations on the institution side), and 86 (right to an explanation of individual decisions). Whether a given AML AI system is high-risk under Annex III turns on Commission classification guidelines under Article 6(4), signalled to publish during 2026. Until those land, the responsible posture is to scope readiness against the conservative reading. AML software vendors whose architecture and documentation are being built with Arts. 9–15, 26 and 86 in mind are easier to integrate; AML software vendors whose materials trail those Articles will be harder to defend to a supervisor.
AI in AML has demonstrated value across all four propositions. The point of these questions is to buy the right product for the right problem, against a benchmark that tells you whether it will work.
Define the problem first. Find the technology that addresses it. Not the other way around.
A procurement checklist in three stages – before evaluating the headline number, before signing, after deployment.
What durable AI adoption looks like
These two questions help at procurement time. But the institutions that have deployed AI in financial crime prevention most durably are distinguished less by what they bought than by what they did first: they defined the problem before selecting the tool, built their baseline before measuring improvement, and treated model governance as an ongoing operational responsibility rather than a deployment milestone.
A well-configured rule-based system can already deliver substantial false-positive reduction, faster investigations, and better alert prioritisation through tuning, segmentation, and suppression. Institutions that have not done that work before introducing AI may find that the marginal value of the AI is smaller than expected, because some of the available gain was never captured by the existing system and the AI is receiving credit for it.
The strongest AI results in AML tend to come from institutions whose rule-based foundations were already strong. They introduce AI to do what rules cannot –detect previously unknown patterns, score complex multivariate risk, handle linguistic and cross-cultural name variation, synthesise case narratives from large document sets – and they measure it against the strong baseline they already built. The risk-based approach embedded in FATF Recommendation 1, and operationalised through the FATF Methodology’s Immediate Outcomes, distinguishes effectiveness in practice from technical compliance on paper. AI is not exempt from that test.
The two questions in this article – compared with what, and which value proposition – are a starting point for closing the gap between AML software vendor claims and operational reality. They are not difficult to ask. They are simply not asked often enough.
Sources & notes
- Halford et al., Journal of Money Laundering Control, 2025 – DOI 10.1108/JMLC-09-2024-0152 (CC-BY 4.0). Reported figures are from an anonymised bank’s 11-month continuous pilot/rollout and are institution-specific; they are not benchmarks.
- EU AI Act – Regulation (EU) 2024/1689 – EUR-Lex. High-risk obligations apply from 2 August 2026; Annex III classification is conditional on Commission guidelines under Article 6(4), signalled to publish during 2026.
- FATF Methodology for Assessing Technical Compliance with the FATF Recommendations and the Effectiveness of AML/CFT Systems – fatf-gafi.org. The eleven Immediate Outcomes are the effectiveness lens referenced above.
AMLYZE builds transaction monitoring, customer screening, and AML case management software for regulated financial institutions: configurable rule catalogue, no-code rule design. Get a demo here.








