Payments

Fraud Models Now Compete on Approval Yield

Blackrock Research
June 18, 2026
4 min read

Fraud Models Now Compete on Approval Yield

A fraud model can post an impressive detection rate and still destroy value if it rejects too many good customers. Visa’s June product announcement frames its Large Transaction Model around a joint objective: improve fraud detection while increasing authorization performance and reducing false declines. That is the right operating frame. Merchants and issuers should evaluate payments AI by approved legitimate value net of fraud, friction and review cost.

What the evidence shows

Legacy fraud programs often optimize within organizational silos. Risk teams focus on chargebacks and loss rates. Payments teams focus on authorization. Customer teams see complaints after a card is declined. Those measures describe the same decision from different angles, but they are rarely reconciled in one economic model.

Network-scale machine learning makes that separation less defensible. Visa says its new model is trained on billions of transactions. It is also enriching payment tokens with context about transaction type, channel and actor, then generating an assurance signal from provisioning and behavioral history. The promise is not merely a more accurate fraud score. It is more information at the moment an issuer decides whether a specific transaction is trustworthy.

That changes the competitive question. A model that catches one more unit of fraud but rejects several units of good demand may look strong on a risk dashboard and weak on a profit-and-loss statement. The useful frontier plots fraud avoided against legitimate value approved, including downstream service and loyalty effects.

The operating consequence

False declines are not evenly distributed. New customers, travelers, high-value orders and unusual but legitimate behavior can look risky precisely because history is thin. Merchants may lose the immediate sale and the future relationship. Issuers may push a customer toward another credential. Support contacts and manual reviews add cost even when the transaction is eventually recovered.

A richer token assurance signal may help, but it can also concentrate influence in network-level models that merchants cannot inspect. Operators need to know which decisions changed, whether the change improved outcomes for different cohorts and how quickly a bad pattern can be reversed. Aggregate approval improvement can hide localized harm.

The economics also depend on who controls thresholds. A model provider can supply a score, yet an issuer, acquirer or merchant often chooses the policy applied to it. Good model performance cannot rescue a blunt rule layered on top.

What operators should do now

Create a single approval-yield scorecard. At minimum, track approved legitimate value, confirmed fraud, false-decline rate, step-up rate, manual-review cost, customer contact and recovered transactions. Segment by new versus returning customer, channel, geography, credential type, basket value and merchant category.

Use champion-challenger testing with economic guardrails. A challenger should not graduate because its global accuracy rises by a fraction. Require evidence that net approved good spend improves without unacceptable loss or cohort degradation. Keep holdouts long enough to observe chargebacks and repeat behavior.

Treat explanations as operational tools, not cosmetic model outputs. Reviewers need reason codes that point to actionable evidence. Customer messages need a different translation: specific enough to resolve the purchase, but not detailed enough to teach adversaries how to evade controls.

Finally, define rollback ownership before launch. Network features and token signals can change quickly. Teams need versioned thresholds, monitoring for distribution shifts and a documented path to restore the prior policy when approval or fraud moves outside tolerance.

Set the economics before tuning thresholds. Put an explicit value on an approved order, a chargeback, a review, a step-up challenge and a customer contact. That makes tradeoffs visible when teams test a policy change. It also prevents a risk reduction that looks favorable in basis points from passing when it destroys more contribution margin than it saves.

Governance should follow the same joint objective. Growth and risk leaders should sign off on launch criteria together, and the post-launch review should show both loss and approval outcomes. If one side can declare victory using its own metric, the organization has not actually solved the optimization problem.

The decision

The next generation of fraud AI will be judged less by how much risk it identifies than by how efficiently it separates good demand from bad. Approval yield forces risk, payments and growth teams to look at the same economic outcome. That is a healthier competition than accuracy alone.