PaySim / SCM 3.2.0 / 01 August 2026 / Test run
The catches.
The misses.
The cost.
6,362,620 synthetic transactions through the engine. The useful finding is not just how much fraud the screen catches. It is what it misses, and what reviewing its alerts would require.
01 / The hidden gap
The missed frauds are larger.
5.7×
Mean missed fraud versus mean caught fraud
7,504,264 versus 1,316,532 simulator units. Only 201 frauds were missed, but their value makes the gap consequential.
The drain-to-zero rule can miss capped transfers that leave a balance. That is a limitation of this screen on this synthetic source, not proof of causal insight.
02 / The operating burden
Recall is only half the decision.
1,520,581 alerts
At an assumed 30 seconds per alert, that is 12,672 review hours, or approximately 55 full-time reviewers over a 31-day window at 7.5 productive hours a day.
This is workload arithmetic, not observed staffing or a forecast. Actual handling time, productive hours and payment friction are not measured here.
03 / The comparison
Better than which alternative?
| Screen | Recall: count | Recall: value | Alert rate | Alerts / fraud |
|---|---|---|---|---|
| CADi scored detector | 97.6% | 87.5% | 23.9% | 190 |
| Random at the same alert rate (expected) | 23.9% | Not reported | 23.9% | Not reported |
| Flag every TRANSFER + CASH_OUT | 100.0% | 100.0% † | 43.5% | 337 |
| Dataset-native isFlaggedFraud † | 0.19% | Not reported | 0.00025% | 1 |
| July run, SCM 3.0.0 | 97.6% | Not measured | 23.9% | 190 |
† Historical CSV recomputations recorded in the study readout. These cells cannot be re-derived from the sealed aggregate and have not been recomputed for this report. All other current-run numerical results are calculated from the committed aggregate.
45.1% fewer alerts than the structural rule, at lower count and value recall. The July matrix is identical: this run improved instrumentation, not detector performance. Random lift is 4.08× by count.
04 / Where review accumulates
Some alerts cannot be fraud in this simulator.
332,507 alerts (21.9% of all alerts) occur in PAYMENT and DEBIT. The source generates fraud only in TRANSFER and CASH_OUT. No channel filter was fitted after seeing this result; doing so would tune to PaySim, not establish performance on real payments.
05 / Coverage and boundaries
All 27 detectors scored. One informative result.
Only BALANCE_EMPTYING_RISK produced a non-degenerate matrix. The other 26 need facts this single-event feed does not supply. Zero firings are a source limitation, not proof that those detectors work or fail on real transactions.
Inspect all 27 detector matrices
| Detector | TP | FP | FN | TN |
|---|---|---|---|---|
| PSD2_SCOPE_APPLIES | 0 | 0 | 8,213 | 6,354,407 |
| SCA_MISSING_RISK | 0 | 0 | 8,213 | 6,354,407 |
| EXEMPTION_INELIGIBLE | 0 | 0 | 8,213 | 6,354,407 |
| LOW_VALUE_EXEMPTION_MISUSE | 0 | 0 | 8,213 | 6,354,407 |
| MIT_CHAIN_EVIDENCE_MISSING | 0 | 0 | 8,213 | 6,354,407 |
| CIT_MIT_MISCLASSIFICATION | 0 | 0 | 8,213 | 6,354,407 |
| SOFT_DECLINE_AUTH_REQUIRED | 0 | 0 | 8,213 | 6,354,407 |
| AML_REVIEW_REQUIRED | 0 | 0 | 8,213 | 6,354,407 |
| STRUCTURING_PATTERN | 0 | 0 | 8,213 | 6,354,407 |
| HIGH_RISK_CORRIDOR | 0 | 0 | 8,213 | 6,354,407 |
| SANCTIONS_NAME_MATCH_REVIEW | 0 | 0 | 8,213 | 6,354,407 |
| PEP_MATCH_REVIEW | 0 | 0 | 8,213 | 6,354,407 |
| PAYOUT_HOLD_RECOMMENDED | 0 | 0 | 8,213 | 6,354,407 |
| EXCESSIVE_RETRY_RISK | 0 | 0 | 8,213 | 6,354,407 |
| REFUND_DELAY_COMPLIANCE_RISK | 0 | 0 | 8,213 | 6,354,407 |
| CHARGEBACK_RATIO_THRESHOLD_RISK | 0 | 0 | 8,213 | 6,354,407 |
| LATE_PRESENTMENT_RISK | 0 | 0 | 8,213 | 6,354,407 |
| PROCESSOR_DEGRADATION | 0 | 0 | 8,213 | 6,354,407 |
| AUTH_AGING_EXCEPTION | 0 | 0 | 8,213 | 6,354,407 |
| MANDATE_REFERENCE_MISSING | 0 | 0 | 8,213 | 6,354,407 |
| PCI_SCOPE_EXPANSION | 0 | 0 | 8,213 | 6,354,407 |
| SENSITIVE_FIELD_VIOLATION | 0 | 0 | 8,213 | 6,354,407 |
| TOKENIZATION_MISSING | 0 | 0 | 8,213 | 6,354,407 |
| CONFIRMATION_OF_PAYEE_MISMATCH | 0 | 0 | 8,213 | 6,354,407 |
| FIRST_TIME_PAYEE_RISK | 0 | 0 | 8,213 | 6,354,407 |
| BALANCE_EMPTYING_RISK | 8,012 | 1,512,569 | 201 | 4,841,838 |
| ROUND_NUMBER_TO_NEW_PAYEE | 0 | 0 | 8,213 | 6,354,407 |
6,362,620 rows read; 6,362,620 decided; 0 skipped. Independent counters agree. No second detector became non-degenerate. The structural baseline does not win on both recall and alert rate.
06 / What this supports
A reproducible test, not a deployment claim.
The run demonstrates processing and instrumentation on a synthetic benchmark. It does not establish causal effects, prevented losses, production calibration or a viable deployment operating point. Nobody outside CADi has checked this run.
Protocol caveat: prior exploratory output from a killed 420,000-row attempt was inspected. A later 20,000-row smoke test also printed a headline. Both are disclosed in the run record; the preregistration is not a pristine unseen-data exercise. The full run's outputs are sealed, and the registered negative condition remains triggered.
Run identity and calculation provenance
Run: paysim-full-6m-scm320 · SCM 3.2.0 · 01 August 2026. This report reuses completed outputs; no new engine run was performed.
Engine commit: 6c9a3a83ec3a724b01ecea169701ccf48a309ce2
Input SHA-256: 16910f90577b0d981bf8ff289714510bb89bc71bff7d3f220f024e287e4eea6b
Results SHA-256: a6b1be615e30e5b483daf42d372214a20cb71795312edbed2cbcf2f6e42d1647
Source: PaySim (Kaggle paysim1). Current-run counts, channel rates, means and value ratios are calculated from the hash-checked results.json aggregate. Historical CSV-only baselines are marked † above; the 72:1 value comparator has the same limitation. No customer or transaction records are embedded.