Flag Five Transactions a Day, Not Four Hundred.
← Chapter 191
Capstone 29 · Fraud Operations
Plain-language Brief

Flag Five Transactions a Day, Not Four Hundred.

The model works. The surprise is how few cases it should send you, and that working your full queue would cost thirteen times more than having no model at all.

To  Head of Fraud Operations
From  Analysis
Re  Four months of card transactions, 119,556 purchases, 318 confirmed frauds
Where this comes from
Chapter Chapter 191 · Imbalanced Classification: Credit Card Fraud
Part Part XXXI · Capstone Projects: Machine Learning
Dataset capstone-credit-card-fraud.xlsx
Notebook View the analysis

Recommendation

Bottom line

Hold a transaction when the model puts its chance of fraud above about 9 percent. That sends your team roughly five cases a day and cuts the monthly bill from fraud and reviews by 31 percent, from $14,476 to $9,959. It uses about one percent of your review capacity, and that is the right answer, not a sign the model is holding back.

Bar chart of total cost over 30 days. Flagging nothing costs $14,476, the cost-optimal threshold $9,959, and working the full review queue $190,667.
Figure 1. The three policies priced over a thirty-day window. Working every case the team could physically review costs thirteen times what doing nothing costs.

What the recommendation costs, and what the alternatives cost

We priced every option using the figures your team gave us: $16 for a transaction we hold that turns out to be legitimate, made up of analyst time and the annoyance of a customer declined at a checkout, and the transaction amount plus $25 for one we let through that turns out to be fraud. Over a 30-day window:

Table 1. Cost of each policy over a 30-day window, using the operations cost figures.
PolicyCases sent to reviewFrauds caughtCost over 30 days
Do nothing at all00 of 95$14,476
Hold above 9 percent145, about 5 a day37 of 95$9,959
Work the full 400-a-day queue12,00093 of 95$190,667
The line that will start an argument

The recommended policy misses 58 of the 95 frauds and is still the cheapest one available. Catching those 58 would mean holding thousands of legitimate payments at $16 each, which costs more than the fraud does. If your team is measured on the share of fraud it catches, this recommendation will look like a failure, and that is a conversation about the measure rather than about the model.

Why filling the queue is the worst option

Sending your analysts 12,000 cases does catch nearly every fraud there is. It also means about 11,900 of those 12,000 reviews are of customers who did nothing wrong. At $16 each, the reviews cost far more than the fraud they prevent. The capacity is real and it is not the constraint. The price of a false alarm is.

The practical consequence is that the spare capacity is genuinely spare. It is worth deciding deliberately what it should do instead, rather than letting a threshold be set low enough to fill it.

Paired bars comparing predicted against observed fraud rate for three approaches. Oversampling predicts 11.57 percent and class weights 11.53 against an observed 0.32; leaving the data alone predicts 0.23.
Figure 2. Why the choice of imbalance remedy mattered. Two of the three produce a fraud rate thirty-six times the truth, which cannot be multiplied by a transaction amount.

How confident we are

The cost curve is shallow around its minimum. Anything between roughly 4 and 15 percent lands within a few hundred dollars of the best answer, so you do not need to hit the number exactly and you can move it for operational reasons without much penalty.

Two cautions. First, these patterns were learned from the first three months and tested on the fourth; fraud is adversarial and people adapt, so this needs re-measuring on a schedule rather than when somebody wonders. Second, a hold is a real person declined at a checkout, and holds do not fall evenly. Customers who travel, shop late or have recently replaced a phone will see more of them, and none of those things is wrongdoing.

One thing we found on the way

The warehouse table you gave us contains a column recording whether a chargeback was later filed. A model using it looks almost perfect. It is also useless, because that column is filled in weeks afterwards, once an analyst has worked the case. At the moment you have to decide whether to hold a payment, it is empty. We left it in the shared dataset with a note, because it is the kind of column that quietly finds its way into the next model somebody builds.

From Statistics, Data Science and AI: A Visual Handbook by John Fisher. Every statistic, table, and figure in this report is reproduced by the companion notebook.