Last updated: September 2026
An AI proof of concept is only worth running if it can fail. Before an agent touches a single alert, four things belong in writing: the threshold each measurement has to clear, the data the exercise will run against, the people who will sign that the thresholds were met, and what happens if they are not. Settle those and a short exercise produces a decision. Leave them open and you get a few weeks of interesting output followed by another meeting.
This page is the list to settle before week one, written from the buying side of the table. Most published guidance on running an AI POC comes from firms that would like to run it for you, which is why so little of it tells you how to set a number that could come out wrong.
TL;DR
Acceptance criteria written after the results arrive are not criteria. The exercise stops being a test and becomes a story about what happened.
Speed is the easiest thing to demonstrate and the least informative. An agent that is faster and wrong is worse than the queue you have.
Three things belong in writing before the start date: a fixed timeline, success criteria both sides signed, and an executive who has already said that meeting them means proceeding.
The data choice decides what you can prove. Without your analysts' recorded dispositions there is no agreement number to be had, however well the exercise goes.
Set thresholds on agreement, unexplainable disagreement, population completeness, evidence sufficiency, review time, behavior on missing data, and auditability. Set the threshold for missed true positives separately, and set it low.
Those dimensions sit inside four families: decision quality, operational fit, the commercial case, and a written decision rule. Change the numbers, keep the four families.
Nobody can tell you which numbers to pick. There is no published benchmark for agent acceptance thresholds, so measure your own baseline first, including how often your analysts agree with each other. The worked example on this page is one team's set, chosen before they saw any results, and not a number to copy.
Write the failure branch down. A proof of concept with no agreed consequence for failing cannot be failed.
The proof of concept that passes and changes nothing
Buyers rarely ask for a proof of concept out of curiosity. They ask because the last purchase went badly. A stakeholder who owned contracts and regulatory exposure at one institution put the motive plainly: the current solution is not what the team wanted, and nobody wants to repeat getting into implementation and discovering there were things they simply could not do. An engineering lead at the same institution said the executive sponsor wanted the same thing proven, because the sponsor did not want to be in the position the team was in already.
That motive is worth holding onto, because it explains why the criteria matter more than the demo. The exercise is insurance against a repeat, so it has to be able to return a negative.
The failure it needs to guard against sits at the other end. Both teams spend weeks on setup, the agreed measurements come out fine, and the executive sponsor decides to stay with the incumbent anyway. We describe that as racing towards a red light, and when we put it to an engineering lead on a live deal the answer came back immediately: that was fair, and they had seen it many times.
There are two mechanisms behind it, and they are the same mechanism twice. Criteria chosen after the results appear cannot be missed. Criteria that only measure speed cannot be missed either, because an agent will always be faster than a person reading an alert. In both cases the exercise cannot fail, which means it cannot inform a decision.
What an AI risk agent proof of concept is, and what it is not
An AI proof of concept is a time-boxed exercise that tests whether an agent can do one specific job on your data, to a standard you agreed before the work started. For a risk agent that job is usually a judgment rather than a task: disposition an alert, assemble the evidence behind a case, or decide what a reviewer needs to see next.
The neighboring terms get used interchangeably and they license different conclusions.
Exercise | The question it answers | What it may conclude |
|---|---|---|
Prototype | Can something like this be built? | That a mechanism exists. Nothing about your data. |
Proof of concept | Can it do this job, on our data, to the standard we set first? | That the job is achievable at the standard you wrote down. |
Minimum viable product | Will anyone use the smallest shippable version? | That the workflow survives contact with real users. |
Pilot, sometimes called an AI pilot program | Does it hold up in one live corner of the operation? | That it works in production at limited scope. |
Proof of value | Is the benefit worth the cost? | That the economics work, given that it already functions. |
The distinction that decides everything downstream is the last one. A proof of concept asks whether it can work. A proof of value asks whether it is worth it. Most disappointing exercises were sold as the first and judged as the second, and the gap surfaces at the end, when the criteria have been met and the decision still does not arrive.
A machine learning proof of concept has an easier target than an agent one. There you are asking whether a model predicts well enough against a labeled outcome. With a risk agent you are asking whether a judgment is good enough for an analyst to act on, and that question has no answer until you have something to compare the judgment against. Everything in the data section below follows from that.
One boundary before the criteria. The question of which dimensions an agent should be judged on at all is a separate exercise, covered in how to evaluate an AI risk agent before production. This page starts one step later, at the pass conditions those dimensions need.
When it is worth running one, and when it is not
Run one when three conditions hold. You have historical decisions to compare against. Your queue carries enough volume to produce a representative sample. And a named person will own the answer.
Do not run one when the criteria would be generic. A growth lead on one buying team described their own gap accurately when asked whether they had a business requirements document for the functionality under evaluation: they had a very generic one, which is how they identified that they needed a person to own it, because it had to get more specific about risk tolerance and about which controls they wanted in place. That person had not been hired yet, so the stakeholders on their side were still unknown and the outcome was undefined before anything started. They had spotted the problem and were fixing it in the right order, which is more than most teams manage.
Do not run one when nobody has the authority to act on a pass. That is the red light, and the right response is to delay rather than to proceed carefully.
Do not run one yet if your queue is generating noise the agent will faithfully process. An agent pointed at a duplicate-heavy queue proves that the duplicates are real. Basic rule hygiene comes first.
The diagnostic question is short. If the agent performs exactly as you hope, what decision happens next week, and who makes it? If that has no answer, the criteria are not ready.
Three things to settle before week one
Before any proof of concept starts, three things have to exist.
A fixed timeline. A start date, an end date, and a named event at the end. An intention to wrap up in a few weeks is not a timeline.
Mutually agreed success criteria, with named signatories on both sides. The word doing the work is mutual. Criteria one side wrote alone are a wish list or a sales script, depending on which side wrote them. Name the people who will say that the criteria were met, and get their names down before the first alert is loaded.
Executive alignment in advance. The sponsor states, before the work begins, that meeting these criteria means proceeding. Most buyers treat this as a follow-up step. It is the one prerequisite that prevents the red light, because it converts a technical result into a decision that someone has already committed to.
Two additions belong on the same page.
Settle pricing before the exercise runs rather than after. A proof of concept that succeeds technically and then dies on a budget surprise has failed, and the inputs a quote needs are knowable up front: expected volumes by line of business and expected onboarding events. Waiting weeks into an exercise to get financial information is how a good result becomes a stalled one.
Then agree which internal workstreams run in parallel, because the tension there does not resolve itself. An engineering lead on one deal was direct about it: a month or two of vendor onboarding and background work is reasonable, but they were not going to start that process or ask management to sign before the team could prove it could use the solution.
The vendor's interest runs the other way, since finishing the exercise and then discovering that contracting takes another six weeks adds a quarter to the calendar. Neither side is wrong. Decide explicitly which reviews start now and which wait.
Calibrated or blind: the data choice that decides what you can prove
Start by fixing the scope, because fixing the scope and choosing the data are the same decision. Criteria written against a fuzzy scope get argued about later, and the argument arrives in week three, when it is expensive. This page carries one worked example all the way through, an agent on AML level-one alert triage, and in that example the scope is settled first.
Scope element | Fixed at kickoff, in this example |
|---|---|
Workflow under test | AML level-one alert triage: context assembly and a recommended disposition. Not level-two investigation, not SAR filing. |
Case set | 1,200 historical alerts from the previous six months, all carrying the disposition an analyst recorded at the time |
Source systems the agent reads | Transaction monitoring alerts, customer records from the system of record, the KYC onboarding file, prior case history |
Explicitly out of scope | Device and behavioral signals, which an existing tool already covers. Sanctions screening. Any automated action on an alert. |
Duration | Four weeks, with the end date fixed at kickoff and a checkpoint in week two |
Named decision-maker | One person with the authority to say no. In this example, the BSA officer. |
The row teams most often skip is the third one. If you already run a device intelligence tool, say plainly that the exercise is not evaluating device signals; otherwise a fortnight disappears into comparing the agent against a capability you already own. An exclusion written at kickoff is far easier to defend than one argued for in week three.
Then the data choice. It is the most consequential decision on the page and the one most often made by accident.
A calibrated exercise runs against your historical alerts together with the disposition rationale your analysts recorded at the time. Because the past decisions are present, agreement becomes measurable: for each alert you can ask whether the agent reached the same disposition your team did, and where it did not, why.
A blind exercise runs without those dispositions. You still see the agent's full analysis on every item, and you cannot produce an agreement number no matter how well it goes. The comparison has nothing to compare against.
The trap is the combination. A buyer who has not made this choice consciously has usually agreed to a blind exercise while expecting a calibrated result. What happens then is predictable: the team judges the agent on whether its reasoning looked sensible. That is a demo with a longer runtime.
A third option is legitimate and often correct. If your dispositions live in unstructured notes, you can wait until they are captured in structured form as a matter of course and run the calibrated exercise later against a cleaner baseline. Choosing that deliberately is a better outcome than running a blind exercise and calling it calibrated.
Synthetic data has a real place, and it is a narrower place than it looks. Formatted in your own schema, it removes the integration work entirely, so an exercise can start without engineering time on your side and the screens look like what you would see once live. Be exact about what it proves: that the workflow handles your data shapes, not that the judgment matches your analysts.
The 1,200-case set in the example above is what a full calibrated run looks like by the end. The inputs needed to start one are modest:
Roughly 25 historical alerts to start, or 25 to 50 with the matching rationale if that rationale sits in documents rather than fields.
60 to 90 days of transaction history for the entities on those alerts, per your own procedure.
Basic entity and identity data.
A session of about 45 minutes watching how your analysts work alerts today.
That last one is the step most often skipped and the one the whole measurement rests on. You cannot set a threshold for agreement with your analysts without first observing what your analysts do.
The acceptance criteria, by dimension
Eight dimensions are worth setting a threshold on, and they sit inside four families: decision quality, what the agent has to get right; operational fit, whether anyone can actually use it; the commercial case, whether it is worth buying; and a written decision rule for the review.
Change the numbers, keep the four families. A set missing one of them has a predictable failure mode: no decision-quality criteria and you argue about whether the output was good, no operational criteria and you get an exercise that works and never converts, no commercial criteria and the purchase stalls in procurement, no decision rule and you get an extension.
Gate on fewer criteria than you track. Three or four that would genuinely change your answer beat a dozen that produce a scorecard nobody acts on. If missing a criterion would not change the decision, it is a metric: measure it, do not gate on it.
Circulate the criteria before kickoff. The value is in the disagreement it surfaces while disagreement is still cheap. If the risk lead and the operations lead write different targets into the first row, that is a conversation worth having in week zero rather than an argument in week five.
Here is the deliverable. Set a threshold on each dimension, in writing, before the exercise starts. The third column is where the work is, because none of these numbers can be handed to you.
Dimension | What to measure | How to set your threshold | What a failure means |
|---|---|---|---|
Agreement with your analysts | Share of sampled alerts where the agent reaches the same disposition your team recorded | Anchor to your own baseline, including how often two of your analysts agree on the same alert | The judgment does not transfer to your risk appetite, whatever it does elsewhere |
Missed true positives | Confirmed-true cases in the sample the agent would have cleared | Set separately from agreement, and set it low. Most teams land at zero for the sample | The disagreements are concentrated in exactly the cases that matter |
Unexplainable disagreement | Share of disagreements a reviewer still cannot account for after reading the agent's output | Set it as a ceiling, not a floor. A disagreement you can explain is usable information | You cannot supervise what you cannot reconstruct |
Population completeness | Share of the sampled queue the agent processed end to end without a human filling a gap | Close to the whole sample, or the exercise only tested the easy alerts | The result describes a subset you did not choose |
Evidence sufficiency | Share of outputs containing what a reviewer needs to act, including enough to support a filing decision | Judge against your own procedure, not against a general standard | Faster output that a reviewer still has to rebuild by hand |
Time per review | Median analyst minutes per alert, measured before and after on the same alert types | Measure your current number first. A reduction target without a baseline is not a target | Either no gain, or a gain that came from skipping steps |
Behavior on missing data | What the agent does when a required field is absent: abstain and escalate, or proceed on an assumption | Decide the required behavior in advance and test it deliberately | Silent assumptions in the cases you would most want flagged |
Auditability | Whether you can reconstruct, months later, the input, the reasoning and who decided | Reconstruct one case from the exercise as a test of the record, not of the agent | A decision you cannot defend after the fact |
Six of those eight rows are decision quality and the evidence behind it; time per review and auditability are operational fit. The commercial case has no row there at all, which is the most common of the four omissions and the one that stalls an exercise that worked.
The most concrete test in this whole set came from a buyer rather than a vendor. An engineering lead proposed running the same population through the incumbent system and the candidate and comparing them, with analysts watching how each case works through and evolves. It is measurable, needs no vendor cooperation to define, and generalizes to any agent. Insist on it.
Two things stay outside the scope of these thresholds. Who holds which review right once the agent is live is a separate design question, settled per queue against your own procedure rather than during the exercise.
The second is what you do with the disagreement data after deployment, once overrides accumulate and patterns appear in them. During the exercise the disagreements are a measurement. Afterwards they become a feedback loop, and that needs its own owner.
While you set the evidence threshold, check what an agent attaches to each recommendation. The useful question is whether the reasoning, the evidence and the typologies arrive in a form a reviewer can act on, rather than whether the recommendation itself is right.
Now the honest part. No published benchmark for agent acceptance thresholds exists, in supervisory guidance or anywhere else we have been able to find, this page included. We can tell you which dimensions to set thresholds on. We cannot tell you what number to pick, and anyone who hands you a number without having seen your queue is guessing.
So measure your own baseline first. Find out how often two of your analysts reach the same disposition on the same alert, because that figure is the ceiling on what agreement with an agent can mean. If your own analysts agree with each other less often than the target you are about to set for an agent, that target is incoherent rather than ambitious.
One deployment shows the shape a threshold takes when somebody has set one. A bank in deployment shows 92% agent agreement with its analyst team on alert disposition and 75% less time per review. Read that as an existence proof rather than an expectation: it is one institution, with its own queues, its own procedure and its own definition of a match.
The reason to look at it is the pairing. An agreement rate and a time figure together say something a time figure alone cannot.
A worked example: AML level-one alert triage
What follows is one team's criteria for the scope fixed earlier on this page, written out because a list of dimensions is easier to agree with than to use. Every figure in it is a target in this example, chosen before that team saw a single result. The 92% above and the 85% below are not the same kind of thing: one is a measurement taken from a deployment, the other a threshold somebody set in advance for their own queue. Read the columns, not the values.
Decision quality: what the agent has to get right. The benchmark is your analysts' historical dispositions rather than a vendor accuracy figure, which describes someone else's alert mix.
Criterion | Target in this example | How it is measured |
|---|---|---|
Agreement with analyst disposition | At least 85% across the 1,200-case set | Blind comparison, in which the agent does not see the recorded outcome |
Missed true positives | Zero alerts the analysts escalated to a SAR that the agent would have cleared | Manual review of every disagreement where the analyst escalated |
Behavior on ambiguity or missing data | Low-confidence cases escalate rather than guess, and no more than 15% of cases land in that bucket | Review of every low-confidence output |
Evidence sufficiency | An analyst can act on the stated reason without re-opening source systems, in 9 of 10 sampled cases | Blind spot-check of 50 recommendations by two analysts |
The second row has a hard zero in it, because on high-risk misses this team's tolerance genuinely was zero. Saying so at kickoff is what prevents an argument about averages at the review.
Operational fit: whether anyone can use it. This is where an exercise most often looks good and still fails to convert. An agent producing excellent recommendations somewhere analysts do not work produces nothing.
Criterion | Target in this example | How it is measured |
|---|---|---|
Time per review | From a 22-minute baseline to under 8 minutes | Timed sample of 40 alerts, before and after |
Where output lands | In the existing case queue. No new interface for analysts to adopt. | Observation during the exercise |
Workflow to be built around it | Configuration only, with no custom routing, approval or escalation code | Count of engineering tickets raised during setup |
Override path | An analyst can reject a recommendation, record why, and the rejection is retained | Tested explicitly in week one, not assumed |
Self-service change | A team member can add or amend a rule and test it in under five minutes, with no engineering ticket | One analyst attempts it unaided, observed |
Auditability | A colleague who was not involved can reconstruct why a disposition was reached, from the trail alone | Pull three closed cases, hand them to someone else, ask them to explain the decision |
That last test is deliberately awkward and it is the one worth keeping. The question is not whether a log exists. It is whether the log answers an examiner's question in eleven months, when the analyst who worked the case has moved teams.
The commercial case: whether it is worth buying. Skipping this is why technically successful exercises stall before purchase.
Criterion | Target in this example | How it is measured |
|---|---|---|
Analyst hours returned per month | At least 280 hours at current volume | Handling-time delta times monthly alert volume |
Cost per agent action against the human equivalent | Agent cost per alert below the loaded analyst cost per alert at 1,200 alerts a month | Per-action price times volume, against loaded hourly cost times handling time |
Use of returned capacity | Named in advance: clearing the level-two backlog, not headcount reduction. Agreed with the operations lead at kickoff. | Recorded in the criteria document before the start date |
Business case owner | Named, and presents at the review | n/a |
The second row deserves a blunt check before the exercise rather than after. At low case volumes with cheap per-case handling, an agent can cost more per action than the work it replaces, and if that arithmetic does not clear at your volume you will succeed technically and fail commercially, which wastes everyone's month.
The third row matters more than it looks. Efficiency is not an outcome anyone can point at in six months. A level-two backlog going from three weeks to five days is.
Come out of the first week holding a table like that, one row per queue. Which alert types allow a high-confidence recommendation to be cleared in one click, and which always route to manual review, is a question of risk appetite rather than of technology. Defining those thresholds is part of setting the exercise up, not a result of it, and a plan that leaves them to be decided later is missing a deliverable.
The phases, and the gates between them
Ask any vendor for the same four-part shape: data setup, a backtest against your historical alerts, a shadow run, and a readout. That is the structure of a serious pre-production evaluation, and it is worth asking for by name so that you can tell whether what you are being offered has all four parts.
The gates matter more than the sequence. Each transition needs an exit condition written in advance.
Data setup closes when the agreed sample is loaded and either the dispositions are present or the team has consciously accepted a blind exercise.
The backtest closes when the agreement and missed-true-positive numbers exist and have been read against the thresholds you set, not against how they feel.
The shadow run closes when the agent has seen production data without triggering an alert, a case or any downstream action, and its output has been compared against what your analysts actually did on the same items.
The readout closes the exercise, and only when a decision and an implementation timeline come out of it.
Ask about the states an agent passes through on its way to a live decision, too. The shape to require is graduated: tested in isolation, then run in shadow against production data without triggering alerts or cases, then a live experiment allocating a share of volume to the new version against a control, then live. Treat that as a question to put to any vendor rather than an assumption, and get the answer in writing, because shadow running and versioned promotion are the difference between a controlled rollout and a switch.
Two constraints on how you write this up. Do not accept a week-by-week calendar as the substance of the plan: the shape is the reusable part, and a calendar mostly reflects how a vendor staffs.
And treat integration dependencies as a scoping gate rather than a detail. If something you want proven depends on a live connection to a third-party tool that does not already exist, the exercise gets longer or the response gets simulated. Both are acceptable; discovering in week two which one you are getting is not. Ask which of your criteria require a third-party connection before the timeline is agreed.
What to leave out on purpose
Every additional thing you prove pulls in another team whose time you then have to justify.
On one deal the vendor argued to keep end-to-end case management out of the exercise, specifically because proving it would bring in a whole operations team that otherwise need not be involved. The engineering lead ruled it out in one line: it was not necessary, the core users were the compliance team, and proving what that team needed was enough.
Note the direction of that argument, because it is the useful part. The vendor proposed narrowing the scope. A vendor pushing to prove everything is optimizing for how impressive the exercise looks, not for how usable your decision is.
The rule to apply: include what the core user needs proven, and exclude what you already believe. If a demo has already convinced you that something works, do not spend the exercise re-proving it.
Some things are better sequenced than cut. On a second deal, evaluating one specific agent was deliberately deferred until the platform itself had cleared. That is the same boundary drawn earlier: which dimensions an agent gets judged on is one exercise, and the pass conditions on this page are another. Running them in the wrong order wastes both.
One thing stays in regardless of how far you narrow. Keep enough of an overview in the readout that everyone in the room understands what they are looking at. The attendees will not all share the same context, and a result nobody can follow is not a result.
Who signs off, and what happens next
Sign-off is a named list agreed at the start. Whoever happens to be in the room at the end is not a sign-off.
For who belongs on that list, borrow the regulator's own test. Supervisory guidance on model risk management was rewritten on 17 April 2026, when the Federal Reserve, FDIC and OCC jointly issued revised guidance as SR 26-2, OCC Bulletin 2026-13 and FDIC FIL-15-2026, superseding SR 11-7 from 2011 and the 2021 interagency statement on model risk management for BSA and AML systems. The revised guidance keeps the idea of effective challenge and says who can perform it: individuals with "the appropriate expertise to conduct a critical and objective challenge, sufficient independence to maintain objectivity, as well as the organizational standing and influence to effect any change."
The third condition is the one that gets skipped. A reviewer who can raise an objection and cannot change the outcome fails it, and that is the regulator's framing rather than an opinion.
Worth knowing while you build the list: the revised guidance places generative and agentic AI outside its own scope, stating that such models "are not within the scope of this guidance," with a request for information on banks' use of AI only planned. So there is no supervisory template for how much human involvement an agent decision needs. Your institution has to decide, write it down and defend it, which is precisely why the criteria are yours to set.
None of that lifts the obligations on the decision itself. What your position looks like once the agent is live, and the audit record a risk agent leaves behind, is the next question after this one.
End on a readout. Walk the agreed list, say what was met and what was not, and come out with a decision and an implementation timeline. That readout is what makes the up-front executive alignment bind, because it is the moment the commitment gets called in. Name the artifact that carries the result forward too: a written scope covering the use cases, the volumes, the workflow elements and the third-party connections involved.
Then write the failure branch, in the same document as the criteria and before the exercise starts. It is short. One name from the sign-off list holds the call, and every branch is decided in advance.
Branch | The rule, written before the start date |
|---|---|
Review | Fixed at kickoff. In this example, the end of week four, with the BSA officer deciding. |
Proceed if | All decision-quality criteria are met, and at least four of the six operational criteria |
Extend once if | Decision-quality criteria are met but integration problems blocked the operational measures |
Stop if | The zero-missed-escalations criterion fails, or the commercial case does not clear at current volume |
Either way | Every missed criterion is documented, with the gap quantified |
Writing the failure branch is what makes honest reporting possible. Without it, an exercise that underperforms tends to get quietly extended rather than concluded, and a clean stop is cheaper than a third extension. A proof of concept with no defined failure outcome is a proof of concept nobody can fail.
How long it should take, and why you will hear different numbers
You will be quoted different durations by different people, including inside the same vendor. Two are worth holding onto because they describe different exercises.
A proof of concept can run in about a week when it runs against historical data with the success criteria agreed in advance. A pre-production evaluation, with data setup, a backtest against historical alerts, a shadow run and a readout, is a four-week shape, which is the duration the worked example above assumes.
The point is not which figure is correct. It is that these are two different exercises, and a buyer quoted one week for what they pictured as the second will be disappointed by a perfectly successful engagement. So the criterion is simple to state and rarely written down: say which one you are buying, in writing, before the start date.
Read the one-week figure correctly while you are at it. A week is possible because the criteria were settled beforehand, which is this page's argument arriving from the other direction.
Why these fail
The two mechanisms at the top of this page were criteria written after the results and criteria that only measure speed. Five more failures are worth naming, and every one of them turned up in the deals behind this page.
No named owner. Observed live: generic requirements, and the person who would make them specific not yet hired.
A blind exercise judged as a calibrated one. The quietest failure on the list and the hardest to notice afterwards, because the output still looks impressive.
A pass nobody was obliged to act on. This is how an exercise passes and never gets operationalized: nobody agreed what a pass obliged anyone to do, so a pass and a fail lead to the same place, which is another meeting.
Scope inflation. A two-week test becomes a quarter of coordination and spends the goodwill you were going to need for implementation.
Nobody with standing on the sign-off list. The criteria are met, the reviewers agree, and nothing moves, because nobody in the room could move it.
Frequently asked questions
What is a proof of concept?
A proof of concept is a time-boxed exercise that tests whether a specific capability can do a specific job under your own conditions, judged against criteria agreed before the work starts. It answers a feasibility question, not a value question. The output is a decision about whether to proceed, not a working product.
What should the acceptance criteria for an AI risk agent POC be?
Four families, in writing, before the start date: decision quality, operational fit, the commercial case, and a decision rule naming who decides and what happens if the criteria are missed. Inside those, set a threshold on agreement with your analysts on disposition, missed true positives, disagreements a reviewer cannot explain, how much of the sampled queue was processed end to end, whether the output contains enough for a reviewer to act on, time per review, behavior when data is missing, and whether the record can be reconstructed later. Set the missed-true-positive threshold separately and set it low. Add the non-technical criteria too: named signatories and settled pricing.
What is the difference between a proof of concept and a proof of value?
A proof of concept asks whether it can work. A proof of value asks whether it is worth what it costs. They need different criteria, and most disappointment traces back to an exercise sold as the first and judged as the second.
How long should an AI risk agent POC run?
About a week if it runs against historical data with the success criteria agreed in advance. A full pre-production evaluation, covering data setup, a backtest against historical alerts, a shadow run and a readout, is a four-week shape. Those are two different exercises, so agree in writing which one you are buying before the start date.
What data does a vendor need for a POC, and what should you not hand over?
For a calibrated exercise: roughly 25 historical alerts to start, 60 to 90 days of transaction history for those entities, and basic entity and identity data. If your analysts' dispositions sit in documents rather than fields, a sample of 25 to 50 alerts with the matching rationale is enough. You do not need to hand over a live production connection, and synthetic data formatted in your own schema will prove that the workflow handles your data shapes, though not that the judgment matches your analysts.
How do you set a threshold before you see any results?
Measure your own baseline first. Find out your current median review time and, more importantly, how often two of your own analysts reach the same disposition on the same alert. Analyst-to-analyst agreement is the ceiling on what agreement with an agent can mean, so a target above it is incoherent rather than ambitious. No published benchmark for agent acceptance thresholds exists, which is why the baseline has to be yours.
What do you measure besides speed?
Agreement with your analysts, missed true positives, disagreements a reviewer cannot explain, completeness of the population reviewed, whether the output contains enough evidence to act on, behavior when a required field is missing, and whether the decision record can be reconstructed months later. Missed true positives deserve their own threshold, because an agreement rate can look excellent while the disagreements sit in the cases that matter most.
Who signs off that a POC passed?
A list of named people agreed at the start, on both sides. Apply the test supervisory guidance uses for effective challenge: appropriate expertise, sufficient independence to stay objective, and the organizational standing to actually effect a change. The third condition is the one most often missing and the one that makes a sign-off mean something.
What happens if the POC fails?
Whatever you agreed in advance: a second round with a changed scope, a different vendor, or no purchase. Write the branches into the criteria document as a short table (proceed if, extend once if, stop if) and name who decides. An exercise with no defined consequence for failing produces no consequence at all.
Can an AI risk agent POC run before we migrate platforms?
Yes. A calibrated exercise runs against exported historical alerts and their dispositions, so it needs no live connection to the system you are leaving. The caveat is the baseline: if your dispositions live in unstructured notes, pull a sample with the matching rationale or accept that you are running blind and adjust what you claim to have proven.
Write it down before week one
Everything on this page reduces to one habit. The thresholds, the data choice, the signatories and the failure branch all get decided before the exercise starts, or they get decided by whatever the results happen to look like. A proof of concept that picks its criteria afterwards cannot fail, and something that cannot fail cannot tell you anything.
None of this needs vendor cooperation to define. You can draft the threshold table, name the signatories and write the failure branch this week, then hand the list to whoever you are evaluating and see whether the exercise they propose can produce those answers. That question is often more informative than the exercise.
If you would rather see an agent work a real alert before you set a threshold against it, Oscilar's risk agents run a one-week proof of concept against your own historical data, with the criteria agreed first.

Oscilar Team
The Oscilar Team is comprised of experts from many domains of risk operations. These articles express viewpoints and knowledge from a variety of sources and contributors across the organization.
DISCLAIMER
The content on this website is provided for informational purposes only and does not constitute legal, tax, financial, investment, or other professional advice. Any views or opinions expressed by quoted individuals, contributors, or third parties are solely their own and do not necessarily reflect the views of our organization.
Nothing herein should be construed as an endorsement, recommendation, or approval of any particular strategy, product, service, or viewpoint. Readers should consult their own qualified advisors before making any financial or investment decisions.
Oscilar makes no representations or warranties as to the accuracy, completeness, or timeliness of the information provided and disclaims any liability for any loss or damage arising from reliance on this content. This website may contain links to third-party websites, which Oscilar does not control or endorse.


