Self-Supervised Knowledge Representation for Rare Fraud and Operational Failure Detection in Multi-Channel Payment Systems
Abstract
In modern high-volume payment systems, detecting fraud is still essentially confined by abhorrent class imbalance, changing transaction patterns, and lack of dependably labelled fraud occurrences. The current research questions the issue of whether self-supervised learning (SSL) can add to the extraction of the knowledge related to fraud in comparison with the capability of strong supervised baselines in the multi-channel payment setting. Using a real world banking dataset of over 13.3 million transactions in the 2010-2019 period, we perform an extensive analysis, including supervised machine learning, anomaly-based SSL, and methods of integrating knowledge into machine learning strategies. Gradient-boosting models are able to build a strong base (F1 = 0.86, ROC-auc = 0.99) that suggests that the trained model has a near-saturation discriminative ability that is solely based on tabular transaction characteristics. We show that naive, generic, SSL-based anomaly detectors lead to reduced precision, and task-adapted representations of supervised models, stacked with task-adapted representations, can increase fraud recall by up to 4.9 with a small F1 increase ( +0.6). However, with strict operationally imposed conditions of accuracy ≥ 0.90, the added benefits of the use of SSL are not experienced, highlighting inherent thresholds of representation based improvement. Such results enhance a knowledge based perspective of when self-supervised representations add value to the decision making and when supervised models have already acquired adequate information about fraud meaning thereby guiding the design of financial fraud knowledge-management models in a robust way.