Flag Five Transactions a Day, Not Four Hundred.
The model works. The surprise is how few cases it should send you, and that working your full queue would cost thirteen times more than having no model at all.
Recommendation
Hold a transaction when the model puts its chance of fraud above about 9 percent. That sends your team roughly five cases a day and cuts the monthly bill from fraud and reviews by 31 percent, from $14,476 to $9,959. It uses about one percent of your review capacity, and that is the right answer, not a sign the model is holding back.

What the recommendation costs, and what the alternatives cost
We priced every option using the figures your team gave us: $16 for a transaction we hold that turns out to be legitimate, made up of analyst time and the annoyance of a customer declined at a checkout, and the transaction amount plus $25 for one we let through that turns out to be fraud. Over a 30-day window:
| Policy | Cases sent to review | Frauds caught | Cost over 30 days |
|---|---|---|---|
| Do nothing at all | 0 | 0 of 95 | $14,476 |
| Hold above 9 percent | 145, about 5 a day | 37 of 95 | $9,959 |
| Work the full 400-a-day queue | 12,000 | 93 of 95 | $190,667 |
The recommended policy misses 58 of the 95 frauds and is still the cheapest one available. Catching those 58 would mean holding thousands of legitimate payments at $16 each, which costs more than the fraud does. If your team is measured on the share of fraud it catches, this recommendation will look like a failure, and that is a conversation about the measure rather than about the model.
Why filling the queue is the worst option
Sending your analysts 12,000 cases does catch nearly every fraud there is. It also means about 11,900 of those 12,000 reviews are of customers who did nothing wrong. At $16 each, the reviews cost far more than the fraud they prevent. The capacity is real and it is not the constraint. The price of a false alarm is.
The practical consequence is that the spare capacity is genuinely spare. It is worth deciding deliberately what it should do instead, rather than letting a threshold be set low enough to fill it.

How confident we are
The cost curve is shallow around its minimum. Anything between roughly 4 and 15 percent lands within a few hundred dollars of the best answer, so you do not need to hit the number exactly and you can move it for operational reasons without much penalty.
Two cautions. First, these patterns were learned from the first three months and tested on the fourth; fraud is adversarial and people adapt, so this needs re-measuring on a schedule rather than when somebody wonders. Second, a hold is a real person declined at a checkout, and holds do not fall evenly. Customers who travel, shop late or have recently replaced a phone will see more of them, and none of those things is wrongdoing.
One thing we found on the way
The warehouse table you gave us contains a column recording whether a chargeback was later filed. A model using it looks almost perfect. It is also useless, because that column is filled in weeks afterwards, once an analyst has worked the case. At the moment you have to decide whether to hold a payment, it is empty. We left it in the shared dataset with a note, because it is the kind of column that quietly finds its way into the next model somebody builds.