Oscilar Team

AML Model Validation: What Changed on April 17, 2026

Posted

Posted

Oscilar Team
Contents

Share this article

Last updated: September 2026

TL;DR

AML model validation is independent review of whether a transaction monitoring model is sound in concept, built on fit data, performing as intended, and still doing so. The supervisory guidance most AML programmes cite for it was replaced on 17 April 2026: SR 11-7 and the 2021 BSA/AML interagency statement are both superseded, the definition of a model has narrowed, and generative and agentic AI are explicitly outside the new guidance while examiners still assess safety and soundness. Validation was never an attestation you could obtain against a document. It is an exercise that produces evidence, and the evidence it has to produce has changed.

What AML model validation is

AML model validation is independent review of a model used in anti-money laundering compliance, establishing that it is conceptually sound, built on data fit for the purpose, performing as intended, and continuing to perform as conditions change.

"Independent" is the load-bearing word. The review has to be performed by a party other than the one that built or owns the model. That single constraint drives most of the cost and nearly all of the friction, because it means the people who understand the model best cannot be the ones who sign off on it.

It is worth being direct about what the output is, because this is where programmes most often mislead themselves. A validation is an exercise, not an attestation. The deliverable is a summary of work that was actually performed — tests run, data examined, weaknesses found and what was done about them. There is no certificate that can be issued against a document, and no framework you can adopt that produces one.

That distinction is why the April 2026 changes matter in practice rather than only on paper. If validation were an attestation against a standard, a new standard would mean new paperwork. Because it is an exercise, a new standard changes what the exercise has to demonstrate.

What changed on 17 April 2026

On 17 April 2026 the Federal Reserve, OCC and FDIC issued revised model risk management guidance — SR 26-2 on the Federal Reserve side, OCC Bulletin 2026-13 on the OCC's. It supersedes and replaces SR 11-7, "Guidance on Model Risk Management" of 4 April 2011, and SR 21-8, the "Interagency Statement on Model Risk Management for Bank Secrecy Act/Anti-Money Laundering Compliance" of 9 April 2021.

Both of those are still cited constantly — in programme documentation, in vendor material, and in most of what has been written about AML model validation. If your model risk policy names either one as its governing authority, that reference is now stale.

The table below is worth checking your own documentation against.

Document

Status

SR 11-7, Guidance on Model Risk Management (April 2011)

Superseded by SR 26-2

SR 21-8, Interagency Statement on Model Risk Management for BSA/AML Compliance (April 2021)

Superseded by SR 26-2

OCC Bulletin 2011-12, Sound Practices for Model Risk Management

Rescinded

OCC Bulletin 2021-19, BSA/AML Model Risk Management

Rescinded

OCC Bulletin 1997-24, Credit Scoring Models

Rescinded

Model Risk Management booklet, Comptroller's Handbook

Rescinded

One scope point deserves emphasis, because the institutions most likely to over-apply the guidance are the ones it was least written for. SR 26-2 states that the letter "is expected to be most relevant to banking organizations with over $30 billion in total assets regulated by the Federal Reserve." A community bank treating the letter as a set of binding requirements for its own programme is reading it more strictly than it asks to be read.

It is also worth remembering what the 2021 statement did and did not do while it was live. The Federal Reserve's press release announcing it recorded that it "does not alter existing BSA/anti-money laundering (AML) legal or regulatory requirements" and that no specific model risk management framework is required. Neither the old guidance nor the new one prescribes a framework you must adopt.

Is your monitoring system a model?

This is the most practically useful question on the subject, and its official answer is one of the things that changed.

The revised guidance defines a model as "a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates." It explicitly excludes simple arithmetic calculations and deterministic rule-based processes that lack underlying statistical or economic theories.

Read against a real AML stack, that definition draws a line somewhere inside most AML transaction monitoring platforms rather than around them. A threshold rule that flags every cash transaction over a fixed amount is a deterministic rule-based process with no statistical theory underneath it. A scenario score derived from a statistical model of expected customer behaviour applies exactly the kind of theory the definition describes. The same platform can contain both.

That is a way of reading the definition, and reading it is your programme's job rather than this article's. What can be said usefully is where the analysis actually bites: the question is not what a system is called or who sold it, but whether a statistical, economic or financial theory sits underneath the number it produces.

How programmes made this call before April 2026

Under the superseded 2021 statement, the customary approach was a three-component test — an information input component, a processing component converting inputs into estimates, and a reporting component translating estimates into useful output. Systems that merely aggregated cash transactions for CTR reporting, or that flagged on a single threshold, were generally characterised as probably not models.

That test is worth knowing because it is embedded in a great deal of existing programme documentation, and because it is how the industry read the previous regime rather than a verbatim quotation of the agencies' text. Treat it as history. The current definition is the one to apply.

Conceptual soundness: the evidence that the model makes sense

The revised guidance organises effective model development and use around model testing, model validation and monitoring — including validating conceptual soundness and outcomes analyses — and governance and controls. Conceptual soundness is the first thing a validator will probe, and the hardest to produce after the fact.

The question is whether the approach is appropriate for the typology it is meant to detect, and whether the design choices behind it can be justified. A scenario built to catch structuring should be recognisable as a scenario built to catch structuring, with a stated rationale for its logic and its parameters. "It was inherited from the previous system" is the answer that fails.

Data fitness is the other half, and it is where most validations find their real findings. Three properties carry it:

  • Coverage — whether every account and product that should be monitored is actually in scope.

  • Lineage — where each input comes from and what happens to it in transit.

  • Representativeness — whether the data the model was developed on resembles the population it now scores.

A conceptually sound model fed unrepresentative data is not a sound model in operation.

Both are documentation problems as much as analytical ones. The evidence that establishes conceptual soundness and data fitness is written down at design time or reconstructed expensively later.

Ongoing monitoring, drift and outcomes analysis

Validation is not an event, and treating it as one is the most common structural mistake. The guidance pairs validation with monitoring, which puts the real question elsewhere: what the model does in the long gaps between reviews.

Drift shows up in three forms, all of which matter in any modelling context:

  • Data drift — the inputs change distribution.

  • Feature drift — the relationship between an input and the derived feature shifts.

  • Concept drift — the relationship between features and the thing being predicted changes underneath the model.

The AML case has a property that makes it distinctive, and it is worth stating plainly: in AML detection the typology changes because someone is adapting to you. Concept drift in a credit model is mostly the economy moving. Concept drift in a monitoring model is often an adversary who has worked out where the thresholds sit. A model that quietly stops catching a typology looks identical, in aggregate, to a model catching a typology that has stopped occurring.

This is why outcome KPIs have to be measured by segment rather than in aggregate. An overall false-positive rate can improve while performance on one customer segment, product, or corridor degrades badly — and the aggregate improvement will read as success. Segment-level measurement is what makes the degradation visible while there is still time to act on it.

At platform level, that means auto model monitoring covering data drift, feature drift and concept drift, with outcome KPIs measured by segment. What a platform can contribute here is the evidence — a continuous record of how a model behaved between reviews, which is precisely what a validator asks for and what most programmes have to reconstruct.

Tuning, thresholds and change approval

Threshold tuning as a model change is the principle to start from. Treating tuning as configuration rather than as a change requiring approval is the single most common finding in this area, and it is easy to fall into because the interface makes it feel like a setting.

A defensible change record answers five questions:

  • Why the change was made.

  • What testing preceded it.

  • Who approved it.

  • When it took effect.

  • How the model behaved before and after.

A programme that can produce those five things for every threshold change in the last two years has a straightforward conversation with a validator. One that cannot has an expensive one.

Governance, policies and controls are the third of the guidance's organising features, and this is where they land concretely. Who is permitted to change what, what evidence is required first, and whether the change is recorded in a form someone else can audit.

Two platform properties do real work here: human-in-the-loop by design, so that a change has an accountable approver rather than an automated path, and a full audit trail per decision with the reasoning stored alongside it. Bias and drift monitoring sit in the same evidence file. None of that constitutes a validation — it is the material a validation consumes.

What a validator will ask you for

The practical shape of a validation is less about analysis than most programmes expect.

A risk lead at a credit union described an external validation where the friction was not the findings but the logistics: the validators needed data extracts pulled for them in order to run their analysis, and producing those extracts became the project. That is the pattern worth internalising. Most of the cost of a validation is the cost of producing evidence on demand, and a programme that can export its own decision history cheaply has a cheaper validation than one that cannot.

Cadence varies more than the literature suggests. A compliance lead at a community bank described going through validation every two years, with the previous one in 2024 and the next already scheduled. That is one institution's practice rather than a standard interval — the right cadence depends on model materiality, change frequency, and what your own policy commits you to.

There is also a two-layer problem that comes up whenever a platform is involved. A risk lead at a firm preparing for a New York licence asked whether there was a validation for the platform as a whole, and separately what happens to the things they build on top of it. Both halves are real questions.

A vendor's own testing of its platform does not validate your configuration of it, and your validation of your configuration does not extend to the vendor's underlying components. The evidence file has to be explicit about which layer each piece of evidence covers.

Where AI-assisted monitoring sits now

The revised guidance is unambiguous on this: "generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance."

Out of scope is not out of supervision, and reading it otherwise would be a serious mistake. Examiners continue to assess safety and soundness, and they will still ask how an institution tests an AI-assisted control, how it monitors it, and where human oversight sits. The absence of a standard removes the reference point, not the scrutiny.

The agencies have also said they plan to issue a request for information addressing model risk management generally and considering, in particular, banks' use of AI. The intention is joint, from the OCC, Federal Reserve Board and FDIC. It has not been published, and there is no comment period to respond to yet.

The practical consequence for anyone deploying AI-assisted monitoring today is straightforward. You cannot point at the guidance for a standard, so you have to be able to describe your own — what the system does, how you tested it, what it would take for you to notice it had stopped working, and who is accountable for the decisions it influences. Institutions that can articulate that will be in a reasonable position whenever the guidance does arrive. The ones treating the scope carve-out as permission not to have an answer will not.

The wider question of how AI fits into an AML programme is its own subject, and the governance answer above is the part that matters for a validation.

If your monitoring runs on risk agents rather than on fixed rules, the governance question is the one above, and it is answerable with the same evidence a validation already wants: what the system decided, on what basis, and who reviewed it.

The pattern worth taking away

The reason this article exists is that the most-cited authority on AML model validation stopped being current in April 2026, and most of the material a practitioner will find has not caught up. If your programme documentation names SR 11-7 or the 2021 interagency statement, that is the first thing to fix.

What has not changed is the underlying discipline. Validation still asks whether the model makes sense, whether the data fits, whether it still works, and whether you can prove it. The programmes that handle a new regime well are the ones already producing that evidence continuously rather than assembling it when a validator arrives — which is also, incidentally, how you find out that a model has stopped working before someone else tells you.

If the immediate question is the monitoring layer itself rather than its governance, transaction monitoring is the closer starting point, and AML for banks covers how Oscilar approaches the programme around it.

Frequently asked questions

Is SR 11-7 still current?

No. SR 11-7 was superseded on 17 April 2026 by SR 26-2, the revised model risk management guidance issued by the Federal Reserve, OCC and FDIC. SR 21-8, the 2021 interagency statement on model risk management for BSA/AML compliance, was superseded at the same time.

What is AML model validation?

AML model validation is independent review establishing that a model used in anti-money laundering compliance is conceptually sound, built on data fit for its purpose, performing as intended, and continuing to perform. Independent means performed by a party other than the one that built or owns the model. The output is a record of work actually done, not a certificate.

Do AML rules count as models?

Not necessarily. The revised guidance defines a model as a complex quantitative method, system or approach applying statistical, economic or financial theories to process inputs into quantitative estimates, and it explicitly excludes deterministic rule-based processes that lack such theories. A fixed-threshold rule and a statistically derived scenario score can sit on opposite sides of that line within the same platform.

Who has to perform model validation?

Someone other than the model's owner or developer — that independence is definitional rather than a preference. It can be an internal function with genuine separation from model development, or an external firm. In practice the choice tends to be driven by the model's materiality and by whether the institution has the internal capacity to be credibly independent.

What evidence establishes conceptual soundness and data fitness?

For conceptual soundness: a stated rationale for why the approach suits the typology it targets, and justification for the design choices and parameters. For data fitness: coverage of all in-scope accounts and products, documented lineage for each input, and evidence that the development data represents the population now being scored. Both are far cheaper to record at design time than to reconstruct.

How should drift and threshold changes be monitored?

Track data, feature and concept drift continuously, and measure outcome KPIs by segment rather than in aggregate, because an aggregate improvement can conceal a segment degrading. Treat every threshold change as a model change with a record of reason, testing, approver, date, and pre- and post-change behaviour. In AML specifically, assume some drift is adversarial rather than incidental.

Does model risk guidance cover generative and agentic AI?

No. The revised guidance states that generative AI and agentic AI models are novel and rapidly evolving and are not within its scope. Examiners still assess safety and soundness, so the supervisory interest remains — what is missing is a reference standard. The agencies plan to issue a request for information covering model risk management generally and banks' use of AI in particular.

Does the revised guidance apply to smaller institutions?

SR 26-2 states it is expected to be most relevant to banking organizations with over $30 billion in total assets regulated by the Federal Reserve. Smaller institutions are not required to apply it as though they were in that population, and over-applying it is a real and common cost. Existing BSA/AML obligations are unaffected either way.

Oscilar Team

The Oscilar Team is comprised of experts from many domains of risk operations. These articles express viewpoints and knowledge from a variety of sources and contributors across the organization.