Oscilar Team

Best Deepfake Detection Tools: 8 Criteria to Consider

Posted

Posted

Oscilar Team
Contents

Share this article

Last updated: September 2026

The best deepfake detection tools for a financial institution are the ones that catch the attacks it has already seen, without failing the genuine customers it wants to keep. That means catching injected video as well as fakes held up to a camera, checking documents and voices as well as faces, and explaining each result in terms an analyst can report. The reliable way to tell tools apart is to test them on your own confirmed attacks and your own good customers before signing.

This guide sets out eight criteria for choosing deepfake detection, each with the questions to ask a provider and what a weak answer looks like. It closes with a test you can run on your own cases and a short table of all eight criteria.

TL;DR

  • Test deepfake detection on attacks your institution has already confirmed, and on sessions from genuine customers, before you compare features.

  • A tool has to catch injection attacks, where fake video bypasses the camera, as well as presentation attacks, where something fake is held up to it.

  • Liveness checks, document checks and voice checks are separate problems, and a provider strong on one may be weak on another.

  • A detection result has to be explainable and reportable: FinCEN's November 2024 deepfake alert asks institutions to use the key term FIN-2024-DEEPFAKEFRAUD in suspicious activity reports, and it is still current as of September 2026.

  • Detection works best as one signal in the decision, alongside device, behavior and application history, with uncertain cases sent to a person.

The short answer: how a financial institution should choose deepfake detection

To choose deepfake detection, a financial institution should run its own confirmed attacks and a sample of genuine customer sessions through each candidate tool, then compare what each one caught, what it wrongly flagged, and whether its explanation would hold up in a suspicious activity report. The attacks worth testing include generated selfies, injected video and generated or edited documents as well as photos held up to a camera. A tool that performs well on a provider's demonstration set can still miss the attacks your institution actually receives.

No provider has closed the problem. A VP of operations at a financial services company described deepfakes this way: "That is a problem I know I need to solve but I don't know if the industry has fully solved it yet." A useful evaluation therefore asks how well a tool performs on your cases today and how it will keep up, rather than whether it stops deepfakes.

Oscilar wrote this list, and Oscilar does not detect deepfakes, run liveness checks or verify documents itself. Oscilar orchestrates the detection providers that do and makes the decision over their results, combined with other signals. The criteria below name no providers or products, and a later section applies the same criteria to Oscilar.

For the wider defense, from blocking mass registrations to catching account takeover, see deepfakes and synthetic identities in finance. This guide covers one part of that defense: choosing the detection tool.

The deepfake attacks financial institutions are actually seeing

Financial institutions are seeing deepfakes in six forms: generated selfies and face swaps, photos or screens held up to the camera, video injected past the camera, generated or edited documents, cloned voices on contact center calls, and genuine people coached through verification by someone else. Most of these arrive during onboarding or at a step-up check, where the institution asks for more proof. A tool that covers only the first form covers only part of the risk.

FinCEN, the US Treasury's financial crimes unit, issued Alert FIN-2024-Alert004 on 13 November 2024 on fraud schemes that use deepfake media to target financial institutions. The alert lists red flags, including a customer photo that does not match other identifying information, identity documents that are inconsistent with each other, and the use of a third-party webcam plugin during live verification. As of September 2026, FinCEN's list of advisory key terms still includes the alert.

Attack

Where it shows up

What detection has to check

What else in the decision helps

Generated selfie or face swap

Onboarding selfie, step-up check

Whether the face was generated or altered

Device, application history, reuse of the same face

Photo or screen held up to the camera

Liveness check

Whether a live person is in front of the camera

Device and session behavior

Injected video or virtual camera

Live video verification

Whether the video came from a real camera

Device signals, camera and software used in the session

Generated or edited document

ID scan, bank statement, business paperwork

Whether the document was generated or altered

Consistency with other application data

Cloned voice

Contact center call

Whether the voice is synthetic

Device, phone number, account history

Coached genuine person

Video verification

Whether someone else is directing the applicant

Behavior across sessions, links between applications

The table maps each attack to the check it needs and to the signals outside detection that help catch it.

Exposure is uneven across institutions. A compliance operations manager at a payments-infrastructure fintech said, "We've had a little bit of AI selfies that have come through." A partnerships manager at a business-banking fintech said AI selfies had been noticeable over the past year, were hard to detect, and were being pushed into manual review.

Some teams see documents first. A senior AML compliance officer at an e-money fintech said identity documents had not been a problem so far, but that forged "documents, agreements, invoices" had already come through in a couple of cases. Others have seen little of any kind: a senior manager of fraud strategy at a money transfer company said deepfakes were "not on our radar" because the team had not yet faced them, as far as it knew. That spread is the reason to test on your own attacks rather than on a general list.

The eight criteria, and how to use them

Use the eight criteria below as a test plan, not a feature checklist: for each one, check a provider's answer against attacks your institution has already seen and against genuine customers it wants to approve. Each criterion has the same three parts: what it is, what to ask a provider, and what a weak answer looks like. A tool can be strong on some criteria and weak on others, so score them separately.

Claims on a provider's website are where an evaluation starts, not where it ends. A senior director of fraud risk management at a payments company put it plainly: "And then you test them and you're like, well, what do they actually do?" The criteria are written so that each question can be answered with a test result.

1. It catches injected video, not only fakes held up to the camera

A deepfake detection tool has to catch injection attacks, where fake video is fed into the verification session without passing through a real camera, as well as presentation attacks, where something fake is shown to the camera. The two need different defenses. A tool built only to judge what the camera sees can pass an injected video that looks perfect, because the camera never saw it.

What it is. A presentation attack holds a printed photo, a screen or a mask up to a real camera. An injection attack replaces the camera feed itself, often through a virtual camera or a browser plugin, so the verification step receives prepared video. FinCEN's 2024 alert lists the use of a third-party webcam plugin during live verification as a red flag, along with attempts to switch the communication method by claiming technical glitches.

What to ask a provider. Ask how the tool tells a real camera feed from an injected one, and which parts of that check run on the device and which on the uploaded image. Ask which injection methods the provider tests against and how often that list changes. Ask for results on presentation attacks and injection attacks reported separately, not blended into one figure.

What a weak answer looks like. A weak answer treats liveness as the whole defense. A risk policy manager at a digital lender described applicants whose face matched the national identity record passing a liveness check by holding a photograph at an angle to the camera. A check that passes a held-up photo detects neither kind of attack.

Passing liveness also does not prove who someone is. A risk lead at an online marketplace, which is a platform business and not a financial institution, described people who pass identity verification, a live interview and liveness checks, yet are not who they say they are. Liveness answers whether a live person is present, and a separate part of the decision has to answer whether that person is the applicant.

2. Its liveness check is one your genuine customers can get through

A deepfake detection tool's liveness check should stop fake faces while letting genuine customers through on the first attempt, and it should be used where the risk justifies it, not on everyone. Liveness that fails good customers moves the cost from fraud losses to abandoned applications. The right question is which liveness method runs, for whom, and at what cost to the genuine customer.

What it is. Passive liveness judges a single capture, such as a selfie or a short video, without asking the customer to do anything extra. Active liveness asks the customer to respond, for example by turning their head or following a prompt on screen. Passive checks add less friction, while active checks give the tool more to analyze, and many tools use both depending on the risk.

What to ask a provider. Ask which method the tool uses by default, whether you can switch methods by risk level, and how often genuine customers fail on the first attempt in your own test. Ask how the check performs on older phones, poor lighting and customers with accessibility needs. Ask what happens after a failure: a retry, a different check, or a person.

What a weak answer looks like. A weak answer is a single liveness step applied to every applicant. A risk lead at a point-of-sale lender moving its business online said its only anti-fraud control had been a liveness check, and that for an online flow where friction matters, biometrics alone did not look like the right path. A tool that only works as a gate for everyone leaves the institution choosing between fraud and friction.

Stepping up by risk works in both directions. A senior director of fraud analytics at a consumer lender said, "I think if we could identify very low risk applications that we could reduce the friction on from authentication perspective, I think that would help us." A head of card fraud risk at a consumer lender accepted the other side of that trade: "there will be chances that we inevitably will have to step up genuine customers."

So the step-up has to be easy to complete. A senior director of fraud risk management at a payments company described the aim as a check that is seamless rather than frictionless, meaning easy for the customer to get through even though it adds a step. Sending every applicant through every check has a cost too: at one lending marketplace, applicants who missed the fast path were sent through every verification page even when the data to decide was already on hand, and many dropped off.

3. It covers documents, not only faces

A deepfake detection tool for a financial institution has to check documents as well as faces, because generated identity documents, edited bank statements and mocked-up business paperwork now arrive as often as fake selfies in some lines of business. Documents usually come in at a step-up, when the institution asks for more proof. If the step-up document is itself fake, the step-up is where the attack succeeds.

What it is. Document deepfakes include generated driver's licenses and passports, edited bank statements and pay stubs, and invented business records such as articles of incorporation. A risk policy manager at a digital lender said, "You can't differentiate between the fake and the original." A VP of risk mitigation and quality control at a mortgage lender named bank statement fraud as one of the team's biggest problems: "They're getting better at it with AI."

What to ask a provider. Ask which document types the tool checks, and whether that includes supporting documents such as bank statements and business paperwork, or only identity documents. Ask how it detects a generated or edited document beyond matching it to a known template. Ask whether it checks the document against the rest of the application, such as the name, address and date of birth already given.

What a weak answer looks like. A weak answer relies on template matching alone. A fraud manager at a telecommunications company, outside financial services but facing the same problem, said the business still relies on submitted documents and "I cannot continue to rely on templates." Generated documents are built to match the template.

Human review catches some of these and misses others. An executive at an auto refinance lender described an AI-generated license submitted with an application: "our human eyes caught it, but our technology needs to catch that for them." An SVP of financial crimes strategy at a large regional bank asked the question every buyer in this category should put to a provider about business paperwork: "do you have any AI that would look at it and go, this was mocked up?"

The step-up itself is part of the exposure. A VP of operations at a financial services company described the design many institutions share: "If we aren't sure we just step you up into an id scan." When the ID scan is the fallback for every uncertain case, a tool that cannot spot a generated ID weakens the whole flow.

4. It covers voice, if your contact center is a way in

A deepfake detection tool needs to cover voice if customers can reset credentials, move money or change account details over the phone, because a cloned voice is a way into the account that bypasses every onboarding check. Voice is a harder detection problem than images. An institution whose contact center is a real channel should test voice detection separately and should not assume a face-detection provider covers it.

What it is. A voice deepfake is synthetic or cloned speech used to impersonate a customer on a live call or in a recorded message. An industry fraud advisor noted how much work has gone into detecting deepfake pictures and added, "Voice biometrics is even tougher in my opinion." Some institutions are building programs for it now: a VP of enterprise authentication and fraud at a credit union said, "We are in the middle of launching a contact center program right now for fraud for voice biometrics".

What to ask a provider. Ask whether the tool analyzes live calls or only recordings, and how quickly it returns a result during a call. Ask what the voice result is paired with, such as the device, the phone number and the account's history. A business architect at an asset manager asked a version of this question directly: whether voice biometrics could work together with device recognition.

What a weak answer looks like. A weak answer treats the voice match as proof of identity. A voice that sounds right is the thing a clone is built to produce, so the result needs other signals behind it. If a provider cannot say how its voice check performs on a call from a new device or number, it has not been tested against the attack that matters.

5. It looks across applications, not one at a time

A deepfake detection tool should compare each application with earlier ones, because some of the clearest deepfake signals exist only across sessions: the same face on different identity documents, a repeat offender returning with a new identity, or a genuine person being coached through verification. A check that judges each selfie on its own cannot see any of these. The comparison has to cover the institution's own history, not just the provider's reference data.

What it is. Cross-application checks look for the same face, document image, device or voice appearing under different identities. A senior director of fraud analytics at a consumer lender described seeing the "same picture on different driver's licenses" in store, and repeat offenders online as well. A senior fraud analyst at a crypto platform described reviewing video submissions to "make sure there's not a third party talking to them guiding them through the process", which can indicate a scam in which the applicant is real but not acting alone.

What to ask a provider. Ask whether the tool matches new faces and documents against your earlier applications, including those you declined, and how long that history is kept. Ask whether it can link applications through shared devices, contact details or document images as well as faces. Ask how it handles a genuine applicant who is being directed by someone off camera.

What a weak answer looks like. A weak answer scores every session in isolation and leaves linkage to a separate system. If the tool cannot tell you that a face it just passed was declined last month under another name, the institution has to find that out some other way. Cross-application checks overlap with synthetic identity detection, and evaluating synthetic identity fraud detection software is a separate choice with its own criteria.

6. It explains what it found in terms you can report

A deepfake detection tool should say what it found and why, in terms an analyst can act on and a compliance team can put in a suspicious activity report (SAR). A score with no reasons forces the analyst to redo the work or to trust the number blindly. The explanation is also what makes human review consistent from one analyst to the next.

What it is. An explainable result names the evidence: which part of the image or document looked altered, which signal suggested injected video, or which earlier application shared the same face. A senior AML compliance officer at an e-money fintech said analysts can often see that a document was generated by AI, but what they need when they report it is an official, justified conclusion that it was. FinCEN's alert asks institutions filing a SAR about deepfake media to include the key term FIN-2024-DEEPFAKEFRAUD in SAR field 2 and in the narrative.

What to ask a provider. Ask to see the analyst view for a flagged case, and whether it shows reasons or only a score. Ask whether the result can be stored with the case so the reasons are still there when a SAR is written or an examiner asks. Ask whether the device, session and account history sit next to the detection result, because an alert that says only that something looked anomalous is not actionable without that context.

What a weak answer looks like. A weak answer is a confidence score with no supporting detail. Without reasons, the outcome depends on who reviews the case. The same compliance officer said, "it depends on the analyst and their experience", and noted that for less experienced colleagues a fake can take longer to spot and is sometimes released as genuine. A VP of risk transformation at a merchant acquirer made the same point more bluntly: "The humans are inconsistent in the application of the policy".

7. It feeds the decision rather than acting as the gate

A deepfake detection tool works best as one signal in the onboarding or authentication decision, combined with device, behavior and account history, rather than as a single pass-or-fail gate. Every detection method misses some attacks. Layered controls mean the miss is caught by something else, and the uncertain cases go to a person.

What it is. The FFIEC's 2021 guidance on authentication and access describes layered security as incorporating "multiple preventative, detective, and corrective controls" designed to compensate for potential weaknesses in any one control. Several of FinCEN's deepfake red flags are not image analysis at all: geographic or device data that does not match the identity documents, a refusal to use multi-factor authentication, and a new account that moves money quickly. For how behavior-based signals work alongside identity checks, see behavioral authentication.

What to ask a provider. Ask whether the tool returns a result your decision platform can combine with other signals, or only a final verdict. Ask whether you can set thresholds by product and channel, and send borderline results to review instead of straight to a decline. Ask what the tool returns when it cannot decide, and whether that state is distinct from a pass.

What a weak answer looks like. A weak answer positions the tool as the whole defense. A VP of operations at a financial services company described the selfie as a safety net used "in far too many places", and said that within a few years "we're going to need to have a different way to feel confident that someone is who they say they are without seeing their face in a picture or a video". A risk policy manager at a digital lender described the need for a place where "a human being has to intervene and actually have a proper check" when something looks wrong.

8. It keeps up as the attacks change

A deepfake detection tool should show how it keeps up as generation methods improve, because a tool that performed well in last year's test can miss this year's fakes. Attackers use the same generative tools that defenders do. The Federal Reserve Bank of Boston has described how generative AI can learn from failure: if a batch of fabricated identities starts to get caught, the pattern is identified and the next batch improves.

What it is. Keeping up means retraining detection models on new attack types, telling customers when detection has changed, and giving the institution a way to confirm performance has not drifted. It also means the provider can show how it handles attack types it has not seen before. For the wider picture of how generative models are used on both sides, see generative AI in fraud detection.

What to ask a provider. Ask how often models are retrained, what triggers an update, and how you will be told when detection changes. Ask how you will know if detection has drifted on your own traffic, and whether you can rerun your test set after each update. Include these questions in provider due diligence: the June 2023 interagency guidance on third-party relationships is in force and was proposed for replacement on 11 September 2026, and as of 29 September 2026 the replacement is still a proposal.

What a weak answer looks like. A weak answer points to a one-time certification or test result. Detection that is not retested decays quietly, and the first sign is often a rise in confirmed fraud that passed the check. A provider that cannot describe its update cycle, or cannot support a retest on your own cases, leaves the institution to discover drift after the losses.

How to test deepfake detection on attacks you have already seen

The most reliable way to compare deepfake detection tools is to send each candidate the same set of your own confirmed attacks and genuine customer sessions, then measure what each one catches and what it wrongly flags. A test built this way reflects your customers, your channels and the attacks that reach you. It takes six steps.

  1. Collect your confirmed attacks by type. Gather generated selfies, held-up photos, sessions you believe were injected, generated or edited documents and, if your contact center is a channel, recorded calls. Label each by attack type so results can be reported per criterion.

  2. Collect genuine sessions and documents from good customers. Include the hard cases: older phones, poor lighting, worn documents and customers who needed several attempts. A tool that flags these is a friction problem.

  3. Send the same set to every candidate. A director of operational risk at a small-business lender described sending the same batch of known-fraudulent documents to several document-fraud tools and comparing what each caught; the tool the lender already used did not perform well.

  4. Measure the genuine set as well as the attacks. A senior director of fraud risk management at a payments company found that one document-fraud tool in a test flagged far more documents as fraudulent than actually were, which would have added friction for most of the genuine customers it flagged.

  5. Score against your own success criteria. A VP of payment risk, compliance and operations at a payments company set three: "how effective is it in detecting bad documents that are humanized?", how much faster it would be, and "how much more efficient will our manual reviews be with this tool". Add whether each result carried an explanation an analyst could report.

  6. Run the leader live on a slice before committing. Route a share of real traffic through the leading tool, alongside your current process, and compare outcomes on the cases that reach review. A test set shows what a tool can do, and live traffic shows what it does with your applicants.

Keep the test set after the purchase. Rerun it after each provider model update, and add new confirmed attacks as they arrive, so the test becomes the way you check for drift under criterion 8.

Where Oscilar fits, judged on the same criteria

Oscilar does not detect deepfakes, run liveness checks or verify documents itself. Oscilar is the decision and orchestration layer around the detection providers an institution chooses: it routes applications to those providers, combines their results with device, behavior and account history, and records why each decision was made. Judged on the eight criteria, that means Oscilar answers some of them directly and relies on a detection provider for others.

  • Criteria 1 to 4: injection, liveness, documents and voice. These depend on the detection provider. Oscilar's consumer onboarding product orchestrates KYC checks and enhanced identity checks from third-party sources, and its device and behavioral product is built "to ensure security and prevent spoofing and reverse engineering". Oscilar's pages do not describe voice deepfake detection, so an institution that needs it should test a voice provider on its own calls.

  • Criterion 2: liveness customers can get through. Onboarding workflows adapt to each applicant's risk profile, and step-up verifications trigger additional checks when they are needed, so a selfie or document check can be reserved for the applicants who warrant it.

  • Criterion 5: across applications. Oscilar analyzes cross-account linkages to uncover fraud rings, which lets a detection result be read against earlier applications that share devices or other details.

  • Criterion 6: explanation. Case management provides prioritized queues, a holistic view of each user and natural-language explanations of why a case was created. Oscilar keeps a full audit trail per decision with stored reasoning, and human-in-the-loop by design.

  • Criterion 7: one signal among several. This is the core of what Oscilar does: detection results sit beside device intelligence, behavioral biometrics and data enrichments in one decision, with decisions specified at under 100 milliseconds.

  • Criterion 8: keeping up. Oscilar's auto model monitoring covers data drift, feature drift and concept drift, plus outcome KPIs measured by segment. Backtests on historical data and A/B tests let a team compare a change before it goes live.

On third-party recognition, Oscilar was named to Chartis's 2026 FCC50, its ranking of financial crime and compliance technology vendors, with category wins for Low-Code/No-Code Customization and Agentic AI Innovation. Oscilar is also a Nacha Preferred Partner for Account Validation, Fraud Monitoring, and Risk and Fraud Prevention. Neither is a ranking of deepfake detection.

If deepfake detection is one part of a wider onboarding decision, choosing onboarding software as a whole is a broader decision than this list covers. For how the pieces fit together in practice, see account opening fraud prevention on Oscilar.

The eight criteria at a glance

The table below summarizes the eight criteria, the question that tests each one, the answer that should worry you, and how to check it on your own cases.

Criterion

What to ask

Weak answer

How to test it

1. Catches injected video, not only held-up fakes

How do you tell a real camera feed from an injected one?

Liveness is the whole defense

Include injected sessions and held-up photos in the attack set, scored separately

2. Liveness genuine customers can get through

Which liveness method runs, for whom, and what happens after a failure?

One liveness step for every applicant

Measure first-attempt failures on the genuine set

3. Covers documents, not only faces

Which document types, and how beyond template matching?

Template matching only

Include generated IDs, edited statements and business paperwork

4. Covers voice, if the contact center is a way in

Live calls or recordings, and paired with what?

The voice match is treated as proof

Test on recorded calls, with device and number data

5. Looks across applications

Do you match against our earlier and declined applications?

Each session scored alone

Include repeat faces and documents from past cases

6. Explains what it found

What does the analyst see, and is it stored with the case?

A score with no reasons

Have analysts draft a SAR narrative from the output

7. Feeds the decision

Can the result be combined with other signals and routed to review?

A single pass-or-fail gate

Check what the tool returns when it cannot decide

8. Keeps up as attacks change

How often are models updated, and how will we know?

A one-time certification

Rerun the test set after an update

Each row is one criterion from this guide, with the question, the warning sign and the test that checks it.

Frequently asked questions

What does a deepfake detection tool check that ordinary identity verification does not?

Ordinary identity verification checks that an identity exists and that the person matches the document they submitted. A deepfake detection tool checks whether the face, video, document or voice itself was generated, altered or injected. The two answer different questions, which is why a synthetic face can pass a match against a synthetic document.

What is the difference between passive and active liveness detection?

Passive liveness detection judges a single selfie or short video without asking the customer to do anything extra. Active liveness detection asks the customer to respond, for example by turning their head or following an on-screen prompt. Passive checks add less friction, and active checks give the tool more to analyze, so many institutions use each depending on risk.

What is an injection attack, and why does it get past a selfie check?

An injection attack feeds prepared video into a verification session through a virtual camera or plugin instead of a real camera. It gets past a selfie check that only judges the image, because the image can look like a real face even though no camera captured it. FinCEN's 2024 deepfake alert lists the use of a third-party webcam plugin during live verification as a red flag.

Can deepfake detection work on a call to the contact center?

Voice deepfake detection exists, and it is generally considered harder than detecting generated images. It works best when the voice result is combined with other signals such as the device, the phone number and the account's history. An institution should test a voice tool on its own recorded calls before relying on it.

What should a bank put in a suspicious activity report about a deepfake?

FinCEN asks institutions filing a SAR about fraud involving deepfake media to include the key term FIN-2024-DEEPFAKEFRAUD in SAR field 2 and in the narrative, under its November 2024 alert, which is still current as of September 2026. The narrative should describe what was found and why it points to manipulated media, such as the red flags the alert lists. A detection tool that explains its result makes that narrative easier to write and to defend.

How does Oscilar measure up against these criteria?

Oscilar does not detect deepfakes, run liveness checks or verify documents itself, so criteria 1 to 4 depend on the detection provider an institution connects. Oscilar covers the decision around detection: combining results with device, behavior and cross-account linkage, routing uncertain cases to review with explanations, and keeping a full audit trail per decision with stored reasoning, and human-in-the-loop by design. Test it on the same cases as any other part of the stack.

The best deepfake detection tool for your institution is the one that catches your confirmed attacks, passes your genuine customers, explains its results well enough to report, and keeps doing so after the next update. Build the test from your own cases, and treat detection as one signal in a decision that also weighs device, behavior and history. Oscilar's approach to account opening fraud is one way to put that decision around the detection provider you choose.

Oscilar Team

The Oscilar Team is comprised of experts from many domains of risk operations. These articles express viewpoints and knowledge from a variety of sources and contributors across the organization.

DISCLAIMER

The content on this website is provided for informational purposes only and does not constitute legal, tax, financial, investment, or other professional advice. Any views or opinions expressed by quoted individuals, contributors, or third parties are solely their own and do not necessarily reflect the views of our organization.

Nothing herein should be construed as an endorsement, recommendation, or approval of any particular strategy, product, service, or viewpoint. Readers should consult their own qualified advisors before making any financial or investment decisions.

Oscilar makes no representations or warranties as to the accuracy, completeness, or timeliness of the information provided and disclaims any liability for any loss or damage arising from reliance on this content. This website may contain links to third-party websites, which Oscilar does not control or endorse.