Crypto AML Rule Tuning
Crypto AML rule tuning covering the ATL/BTL threshold-testing lifecycle, a hedged sub-1 percent SAR-conversion retirement signal, the back-test gate before a threshold ships and what MiCA vs the incoming AMLR actually require of CASP transaction monitoring.
Key takeaways: crypto AML rule tuning 5
The tuning mechanics, retirement signal and regulatory framing that matter most when cutting false positives out of a crypto CASP's transaction-monitoring rules without dropping true positives.
- ATL/BTL is the standard tuning technique Raise a threshold and confirm the alerts you would lose were false positives; lower it and check whether new alerts would have been escalated. Run both against 6-12 months of historical data.
- The sub-1% retirement signal is our heuristic, not a standard Real conversion data runs around 2.8%, typical ranges are 1-5%, a healthy target is often cited as 5-15%, and a low ratio is not inherently bad, per Jim Richards. A rule below 1% is a review candidate, not an automatic kill.
- The back-test gate proves a window, not eternity A back-test shows no SAR-worthy alert was suppressed in the sampled 6-12 month look-back, not literal zero false negatives across all future traffic.
- Crypto typologies need crypto-native rules Structuring, rapid in-out within 24 hours, dormant-then-active, mixer proximity and weak-KYC counterparties are FATF-anchored; peel chain is a vendor term feeding the structuring rule, not a standalone FATF typology.
- MiCA licenses, AMLR mandates monitoring MiCA Article 62 covers internal controls at authorisation; the binding ongoing-monitoring mandate is AMLR plus AMLD6, applying from 10 July 2027, with AMLD5 (applied from January 2020) as today's actual crypto entry point.
Crypto AML rule tuning is the piece of transaction monitoring almost nobody writes about honestly. Vendors sell a rule engine and ship it with default thresholds copied from fiat-bank playbooks, and most public guidance on tuning those thresholds is written for a bank, not a CASP. This guide is the tuning half we could not find written for crypto: how to cut false positives without dropping true positives, when a rule has earned retirement and what a defensible back-test actually proves. It draws on the standard above-the-line and below-the-line (ATL/BTL) tuning technique used across traditional AML programs, adapted to on-chain typologies, plus the EU rules that define what CASP monitoring has to demonstrate.
In short: ATL/BTL threshold testing, raising a threshold to see what you would lose and lowering it to see what you would gain, is the standard, defensible way to tune a rule, and it works the same way on a crypto typology as it does on a wire-transfer rule. A rule whose alert-to-SAR yield sits persistently below roughly 1 percent is a candidate for retirement, not an automatic kill, because a low conversion rate is not inherently a bad sign on its own. A back-test gate proves no SAR-worthy alert was suppressed in a sampled historical window, not literal zero false negatives forever. And MiCA licensing is not where the ongoing-monitoring mandate lives, that sits in the incoming EU AML package.
Crypto AML Rule Tuning Starts With One Question
Every rule that goes into a monitoring engine should be able to answer one question before it ships: what typology does this rule catch. A threshold with no articulated typology behind it cannot be tuned, because tuning means deciding whether raising or lowering it improves detection of something specific, and cannot be defended to an examiner, because "we lowered it because alerts felt high" is not a documented rationale.
In practice that means naming the typology (structuring, rapid in-out, dormant-then-active, mixer exposure, sanctioned-address hops), the parameters that express it (an amount, a count, a time window, a counterparty class, an on-chain feature) and the risk the rule is meant to surface, before a single threshold number gets picked. Everything that follows, the lifecycle, the ATL/BTL mechanics, the retirement signal and the back-test gate, assumes that starting point is already in place.
Why Fiat-Bank AML Rules Misfire on Crypto Transaction Monitoring
A rule engine tuned on wire-transfer patterns carries assumptions that do not hold on-chain. A fiat structuring rule watches account-level deposits against a reporting threshold; a crypto equivalent has to watch wallet-level transfers, and the Financial Action Task Force explicitly analogizes the two, describing structuring VA transactions "in small amounts, or in amounts under record-keeping or reporting thresholds" as a direct crypto counterpart to cash structuring. FATF organizes its red flags into six categories: transactions, transaction patterns, anonymity, senders or recipients, source of funds or wealth and geographical risk, and several of those categories have no clean fiat analogue at all.
Anonymity tooling is the clearest gap. A fiat rule engine has nothing like a mixer, a tumbler or a privacy coin to watch for, and FATF names exactly that anonymity category, funds moved through multiple wallets or addresses specifically to break the traceable relation between them, as a standalone red-flag class. Rapid in-out is a fiat pattern too, but FATF's crypto phrasing is more precise about the window that matters: "multiple high-value transactions... in short succession, such as within a 24-hour period." A rule copied wholesale from a fiat rail and pointed at a wallet address will alert on none of this and drown in false positives on the parts that do overlap, because it has no concept of on-chain-specific signal at all.
One on-chain pattern worth naming precisely, because it gets miscited often: the peel chain, funds moved through a series of wallets with small amounts peeled off at each hop, is a blockchain-analytics vendor term (Chainalysis and Merkle Science both use it), not a FATF-defined typology. FATF does not use the phrase. It is a real, useful pattern, and it feeds directly into a structuring rule's on-chain signal, but a rule engine's documentation should attribute it to the vendors that coined it rather than to a regulator that never defined it.
The 6-Step Rule-Tuning Lifecycle
What follows is our own synthesis of the standard ATL/BTL tuning technique used across AML programs, adapted for crypto-native rules and the typologies above. Treat it as a repeatable loop, not a one-time project.
- Define the scenario. Pin the typology the rule targets, its parameters and the risk it is meant to surface. No typology, no defensible rule.
- Baseline the current alert population. Measure live output over a fixed look-back window, commonly 6 to 12 months: alert count, case count, SAR count, alert-to-case and alert-to-SAR yield and analyst hours spent. This baseline is the reference point every later change gets measured against.
- Tune thresholds via ATL/BTL testing. Raise the threshold in a sandbox and confirm the alerts that would disappear were genuine false positives; lower it and check whether any newly-surfaced alerts would have been escalated. Iterate toward the range that maximizes yield without losing a true positive.
- Back-test the candidate threshold. This is the gate a change has to clear before it ships, and it gets its own section below because the framing matters.
- Deploy with documentation. Ship the change with a written rationale, before-and-after metrics, the ATL/BTL evidence and the back-test sample, so an examiner can reconstruct why the threshold sits where it does. Where the rule change touches how or when a verdict reaches the orchestration layer that acts on it, webhook delivery, polling cadence, cache invalidation, that is an integration-architecture decision rather than a tuning one, and we cover it separately in our crypto transaction monitoring integration guide rather than re-covering it here.
- Monitor and retire. Track post-deployment yield against the baseline. A rule whose yield stays persistently near the floor becomes a retirement candidate, covered in its own section below, but retirement is a decision that itself has to survive a back-test, not a threshold crossed automatically.
When you don't have the history yet
Step 2 above assumes 6 to 12 months of live alert-and-SAR history to baseline against, and step 4's back-test assumes the same kind of window to sample from. A newly licensed CASP, or a brand-new rule bolted onto an established program, has neither. Our practice for that cold-start case: seed the initial threshold from a proxy baseline, peer or typology priors rather than the program's own traffic, set it provisionally on an analyst's expert judgment rather than a measured yield, run a shorter initial look-back window than the standard 6 to 12 months and tighten the threshold in stages as real alert-and-SAR data accrues. Treat every threshold set this way as provisional until it has earned enough of its own history to clear the same ATL/BTL and back-test gates described below, not as a permanent value.
Above-the-Line and Below-the-Line Threshold Testing Explained
ATL/BTL testing is the mechanic behind step 3, and it is worth explaining on its own because the two halves test opposite failure modes. Above the line means raising the threshold, illustratively a 10 percent move from a 5,000 baseline to 5,500, so the rule fires less often. You then examine specifically the alerts that would disappear under the new threshold and confirm each one really was a false positive, not a productive alert you would be losing. The purpose is finding how high the threshold can go before it starts dropping genuine risk.
Below the line means lowering the threshold, illustratively 5,000 to 4,500, so the rule fires more often. You sample the new alerts that appear just under the current line and investigate whether any of them would have been escalated as suspicious. The purpose here is the opposite: proving the current threshold is not silently hiding productive alerts that a slightly wider net would catch.
Both halves run in a sandbox against historical data, typically the same 6 to 12 month window used for baselining, never against live production traffic. And both halves have to be periodic, not a one-off exercise, with every tuning decision documented well enough that an examiner could reconstruct the reasoning later. A threshold that has not been revisited in over a year is itself a finding worth flagging, independent of what the number happens to be.
SAR-Conversion Yield as a Tuning and Retirement Signal
Alert-to-SAR conversion, the share of alerts that end up as a filed suspicious activity report, is a useful signal for both tuning and retirement, but it needs careful framing because the honest numbers around it are easy to misuse. Real survey data from the Mid-Size Bank Coalition of America, cited by former Bank of America BSA officer Jim Richards, put one institution's actual conversion at 2.8 percent (9,648,000 transactions produced 3,908 alerts, 348 cases and 108 SARs in a single month). Industry practitioners generally describe typical alert-to-SAR conversion as somewhere in a 1 to 5 percent range, with a commonly cited "healthy" target sitting closer to 5 to 15 percent, the inverse of an 85 to 95 percent false-positive rate that industry estimates (tracing back to a widely-cited PwC figure, not a regulator) put as the norm. There is no regulatory benchmark for an acceptable false-positive rate or a mandated SAR-conversion floor; every number in this paragraph is industry practice, not law.
Against that backdrop, we propose a specific practitioner heuristic, not an industry or regulatory standard: a rule whose alert-to-SAR yield sits persistently below roughly 1 percent, measured over the same baseline window used in step 2, is a retirement candidate. Persistently below the floor even the low end of typical ranges is a fair signal that a rule is mostly generating analyst workload rather than catching anything a differently-tuned rule, or a different rule entirely, would not also catch.
The load-bearing caveat here is Richards' own argument, and it cuts directly against treating this heuristic as a kill switch: a low alert-to-SAR ratio is not inherently bad, and optimizing toward raw SAR counts is, in his words, "like a car manufacturer tracking how many cars it builds but not how many it sells." He cites a Treasury OIG review that found an 11 percent hit rate on law-enforcement information requests and argues a 10 to 15 percent hit rate could be optimal for that kind of check, precisely because volume alone says nothing about whether the SARs filed have any tactical or strategic value to investigators. A rule below the retirement floor is a candidate for review, not a rule to switch off on sight, and the review itself is a back-test, covered next, proving the rule is not the sole net catching a live typology before it gets retired.
Retirement candidacy is also where teams most often get tempted to build a bespoke tuning-and-yield-tracking tool instead of buying or extending what a vendor already offers. That is a build-vs-buy economics question in its own right, and we work through the real cost trade-offs, vendor pricing against a custom stack's launch economics, separately in our crypto AML screening cost guide rather than re-covering it here.
The Back-Test Gate: A Design Goal, Not a Literal Zero
Every threshold change we described under ATL/BTL testing, and every rule retirement decision, has to clear the same gate before it ships: a back-test proving the change does not suppress a SAR-worthy alert. We frame this as a zero-false-negative gate deliberately, because it is the standard the change has to meet, but it is worth being precise about what that phrase can and cannot prove.
The mechanics: run the proposed threshold over historical data, the same 6 to 12 month window used elsewhere, in a sandbox, then sample and investigate the below-the-line population, the alerts the change would suppress, almost as if they were live alerts. The change passes the gate only if that sampled population contains no alert that would have been escalated to a SAR. Document the sample size and the findings alongside the deployment record from step 5.
What that gate proves is narrower than "this rule will never miss a SAR." It proves no SAR-worthy alert was suppressed in the specific historical window and sample that got tested. It is bounded by how far back the look-back window reaches and how much of the below-the-line population actually got sampled rather than exhaustively reviewed. A rule can pass a rigorous back-test on a 12-month window and still miss something a typology shift six months later would have caught, because a back-test is a statement about the past traffic it was tested against, not a guarantee about traffic that has not happened yet. Treat "zero false negative" as the design goal and the gate criterion the change has to clear, never as a literal mathematical claim about all future transactions.
Crypto Detection Scenarios and Tunable Thresholds
The table below lays out the crypto-native scenarios a tuning program has to cover, alongside the parameters typically tuned for each and what ATL/BTL testing is actually checking. The thresholds shown are illustrative examples for discussion, not regulatory values, and should never be copied into a production rule engine without running them through the lifecycle above first.
| Scenario | Crypto signal | FATF anchor | Example tunable parameters | ATL/BTL note |
|---|---|---|---|---|
| On-chain structuring | Many sub-threshold transfers under a reporting limit | FATF structuring (verbatim) | Amount ceiling, transfer count, rolling window (for example, N transfers under X within 24 hours) | BTL sampling tests whether lowering X surfaces real structuring rather than noise |
| Peel chain | Funds moved through a chain of wallets, small amounts peeled off at each hop | Structuring plus multi-wallet movement concept; term itself is vendor-attributed (Chainalysis, Merkle Science) | Hop count, peel ratio, chain depth | Feeds the structuring rule's on-chain signal; do not present as a standalone FATF typology |
| Mixer or tumbler proximity | Inflow or outflow within a set number of hops of a known mixer or privacy protocol | FATF anonymity category | Hop distance, exposure percentage, direct vs indirect | Vendor exposure scores drive the threshold; attribute the scoring source in documentation |
| Sanctioned-address exposure | Direct or N-hop exposure to a sanctioned wallet | Screening obligation under the EU AML package; analytics via blockchain-analytics vendors | Hop depth, direct vs indirect exposure percentage | Near zero-tolerance on direct exposure; ATL/BTL tuning applies mainly to the indirect hop-depth cutoff |
| Rapid in-out | High-value deposit followed by a near-full withdrawal within 24 hours | FATF multiple high-value transactions in short succession (verbatim) | In/out ratio, time window, value floor | ATL testing checks whether raising the value floor loses genuine layering cases |
| Dormant-then-active | A long-idle account transacts suddenly, then goes quiet again | FATF staggered pattern with a long gap afterward (verbatim) | Dormancy period, reactivation value, post-burst silence window | Low base rate by nature; tune conservatively to avoid noise from ordinary reactivation |
What Regulators Require of CASP Transaction Monitoring
Getting the regulatory anchor right matters more here than in most compliance writing, because the instrument that actually mandates ongoing monitoring is not the one most articles reach for. MiCA (Regulation (EU) 2023/1114) is a licensing and prudential regime. Article 62 requires a CASP to demonstrate robust internal controls to manage risks including money laundering and terrorist financing as part of its authorisation, but it does not itself specify an ongoing transaction-monitoring obligation. Citing MiCA as the source of the monitoring mandate is a common miscite worth avoiding.
The binding ongoing-monitoring mandate sits in the EU's AML package instead: the AML Regulation (Regulation (EU) 2024/1624, AMLR) and the sixth AML Directive (Directive (EU) 2024/1640, AMLD6), which name CASPs as obliged entities required to run customer due diligence, ongoing transaction monitoring, five-year record-keeping and suspicious-transaction reporting without delay. AMLR and AMLD6 apply from 10 July 2027, an incoming obligation rather than a currently-enforced one, and the new supervisory body created to enforce them, AMLA, has been operational since 1 July 2025 and is issuing the technical standards that will fill in the detail through 2026. The original entry point that brought crypto exchanges and custodian wallet providers into the obliged-entity perimeter at all was the fifth AML Directive (Directive (EU) 2018/843, AMLD5), applied from January 2020 following entry into force in July 2018, and national AMLD5-derived transposition, alongside MiCA licensing and the Transfer of Funds Regulation, is the regime actually binding on CASPs today, ahead of AMLR taking effect.
Two further layers matter for a tuning program specifically. The EBA's ML/TF Risk Factors Guidelines (EBA/GL/2021/02) were amended in January 2024 to bring CASPs into scope, issued under Articles 17 and 18(4) of Directive (EU) 2015/849; these are supervisory guidelines on a comply-or-explain basis, not a regulation, and they reference blockchain-analytics tools as an expected part of risk mitigation without mandating a specific rule engine or threshold. And FATF's red-flag indicators, the source for the typologies covered earlier, are standards rather than directly-binding law; they carry force in the EU once transposed, for the Travel Rule specifically via the Transfer of Funds Regulation (EU) 2023/1113, and a rule-tuning program should treat them as authoritative guidance to build against, not as a regulation to cite by number.
Tune Your Crypto AML Rules With a Team That Has Shipped This
Rule tuning is not a one-time configuration exercise, it is an ongoing discipline that needs the lifecycle, the ATL/BTL mechanics and a defensible back-test gate running continuously against real crypto typologies, not a fiat rulebook pointed at a wallet address. If your CASP is carrying a false-positive load a rule engine's default thresholds never accounted for, or is preparing its monitoring program for AMLR and AMLD6 ahead of the 2027 application date, our MiCA compliance software development team can help scope the tuning program, the back-test evidence trail and the integration work around it.
Sources: FATF, "Virtual Assets - Red Flag Indicators of Money Laundering and Terrorist Financing" (September 2020). Abrigo, "Optimizing Your AML Program with Above-the-Line, Below-the-Line Testing." Jim Richards, RegTech Consulting, "Rules-Based Monitoring, Alert-to-SAR Ratios and False-Positive Rates: Are We Having the Right Conversations?" Chainalysis and Merkle Science glossaries for peel-chain and on-chain vendor terminology. Jones Day, "Crypto-Assets, CASPs and AML/CFT Compliance: the New European Regulatory Landscape Under MiCA and AMLR" (July 2025). EBA press releases and the FIAU Malta final report on the amended ML/TF Risk Factors Guidelines. Regulation (EU) 2023/1114 (MiCA), Regulation (EU) 2024/1624 (AMLR), Directive (EU) 2024/1640 (AMLD6), Directive (EU) 2018/843 (AMLD5) and Regulation (EU) 2023/1113 (Transfer of Funds Regulation). The 6-step tuning lifecycle here is our own synthesis grounded in the ATL/BTL sources above, not an attributed ACAMS methodology. This article was last reviewed 22 July 2026.
FAQ
Quick answers to common questions about custom software development, pricing, process and technology.
Type to filter questions and answers. Use Topic to narrow the list.
Showing all 6
No matches
Try a different keyword, change the topic or clear filters
-
Rule tuning is the ongoing process of adjusting a transaction-monitoring rule's thresholds so it catches genuine risk without drowning analysts in false positives. It runs on the ATL/BTL (above-the-line, below-the-line) technique used across AML programs generally: raise a threshold and confirm the alerts you would lose were false positives, lower it and check whether new alerts would have been escalated.
For crypto specifically it has to be tuned against on-chain typologies, structuring, mixer exposure, rapid in-out, dormant-then-active, that a fiat-bank rule engine was never built to watch for.
-
Above-the-line (ATL) testing raises a rule's threshold in a sandbox and examines the specific alerts that would disappear, confirming each was a genuine false positive rather than a productive alert you would be losing. Below-the-line (BTL) testing lowers the threshold and samples the new alerts that appear, checking whether any would have been escalated to a suspicious activity report.
Run together against 6 to 12 months of historical data, the two halves find the range that maximizes yield without dropping true positives.
-
We treat a rule whose alert-to-SAR yield sits persistently below roughly 1 percent as a retirement candidate, not an automatic kill. That is a practitioner heuristic we are proposing, not a regulatory standard: real survey data shows conversion around 2.8 percent, typical ranges cited across the industry run 1 to 5 percent and a commonly cited healthy target is 5 to 15 percent.
Former BSA officer Jim Richards explicitly warns that a low conversion rate is not inherently bad, so a rule under the floor gets a back-test review before retirement, never an automatic switch-off.
-
Not literally. A back-test gate proves that no SAR-worthy alert was suppressed within the specific historical window and sample tested, typically 6 to 12 months, not that the rule will never miss a suspicious transaction in future traffic.
Treat "zero false negative" as the design goal and the gate criterion a threshold change has to clear before shipping, bounded by the look-back window and the sample size, never as a literal mathematical guarantee.
-
Not directly. MiCA (Regulation (EU) 2023/1114) is a licensing regime; Article 62 requires a CASP to show robust internal controls against money-laundering risk as part of its authorisation, but the ongoing transaction-monitoring mandate itself comes from the EU's AML package, the AML Regulation (2024/1624) and the sixth AML Directive (2024/1640), applying from 10 July 2027.
The original crypto entry point was the fifth AML Directive (2018/843), applied from January 2020, and that AMLD5-derived regime alongside MiCA licensing is what currently binds CASPs ahead of AMLR taking effect.
-
No. Peel chain is a blockchain-analytics vendor term, used by Chainalysis and Merkle Science, for funds moved through a series of wallets with small amounts peeled off at each hop. FATF does not use the phrase; the underlying pattern is covered conceptually under FATF's structuring and multi-wallet movement red flags.
A rule engine should attribute the peel-chain pattern to the vendors that coined it and feed it into a structuring rule rather than citing it as a standalone FATF category.
I work with startup founders who need a dedicated software development team but don’t want to gamble on hiring, random outsourcing, or opaque delivery.
Most founders face the same problem sooner or later.
Early technical and team decisions lock the product into tech debt, slow delivery, missed milestones and constant re-hiring. By the time this becomes visible, fixing it is already expensive.As a CTO and software architect, I help founders design, build and run dedicated development teams that work as a true extension of the startup. Not as a black-box vendor.
My focus is on complex products where mistakes are costly:
- Web3 and blockchain platforms
- FinTech and regulated products
- High-load startup systems
- MVP → scale transitions
We don’t do body-shopping.
We don’t sell generic outsourcing.Instead, we help founders:
- build the right team structure from day one
- keep technical ownership and transparency
- scale delivery without losing control
- avoid vendor lock-in and hidden risks
Teams are aligned with the product roadmap, business goals and long-term architecture. Not just short-term velocity.