2018· International Journal of Commerce, Finance and Digital Economy· Vol 1, pp. 1-11· 0 citations
TL;DR
This paper investigates the adversarial robustness of various machine learning models applied to synthetic identity fraud detection in credit risk settings, and proposes and assess defense mechanisms, including adversarial training and robust feature engineering, to enhance model resilience.
Abstract
Synthetic identity fraud represents a growing threat in credit risk management, where attackers create fictitious identities by combining real and fabricated information to bypass traditional detection systems. Machine learning models have demonstrated effectiveness in detecting such fraudulent activities; however, they remain vulnerable to adversarial attacks that can manipulate input data to evade detection. This paper investigates the adversarial robustness of various machine learning models applied to synthetic identity fraud detection in credit risk settings. We simulate diverse adversarial attack strategies on benchmark datasets and evaluate the impact on model performance, highlighting critical vulnerabilities. Furthermore, we propose and assess defense mechanisms, including adversarial training and robust feature engineering, to enhance model resilience. Our results reveal significant trade-offs between accuracy and robustness, underscoring the need for balanced solutions tailored to financial fraud contexts. The study provides actionable insights for practitioners aiming to deploy more secure and trustworthy fraud detection systems, contributing to improved credit risk management in an increasingly adversarial environment.
Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-specific constraints, severe class imbalance, and asymmetric attacker capability. We argue that, in this setting, robustness is not only an attribute of the model, but also an attribute of the evaluation protocol. Different ways of enforcing constraints and capability can lead to substantially different robustness conclusions. This paper presents FraudBench, a protocol-sensitive benchmark for adversarial robustness evaluation in financial fraud and credit-risk detection. Rather than treating domain constraints as post-hoc validity checks, FraudBench evaluates the same dataset--model--attack--defence setting under three matched protocols: unconstrained attacks, post-hoc feasibility filtering, and deployment-aware constraint-integrated attacks. FraudBench covers four public financial datasets, and evaluates neural, tree-based, and ensemble models using three attack settings. Our results show that robustness conclusions are highly protocol-sensitive. On Lending Club Loan Data under the white-box setting, post-hoc filtering leaves only 3.7 feasible-flipped examples on average, whereas in-attack projection with attacker mutability masking produces 2,832.3 feasible-flipped examples under the same perturbation budget. The results on IEEE-CIS further show that feasibility and attacker capability are separate axes, while black-box evaluation shows that protocol choice can alter model-family rankings. These findings suggest that fraud robustness evaluation should report predictive degradation and attack feasibility jointly, and should incorporate domain constraints into attack generation rather than treating them as post-processing checks.
Xitong Zeng, Zhaoge Bi, Yi-Tian Yang et al.· 0 citations
This paper proposes a unified dual-defense framework that jointly integrates adversarial training with a denoising autoencoder (DAE)-based filtering mechanism, specifically designed for imbalanced tabular financial data under adversarial conditions, and explicitly targets adversarial robustness in financial fraud detection.
Mohammed Saad Javeed, Jannatul Maua, M. Mridha et al.· Computers, Materials & C...· 0 citations
A detailed overview of the security risks associated with adversarial attacks is offered, including evasion attacks carried out at inference time, data poisoning that corrupts the training process, backdoor insertion that hides dormant triggers inside a model, and model inversion that leaks private information back out of a trained system.
Harsh Verma· International Journal of Sci...· 0 citations
This study evaluated the adversarial robustness of machine learning-based fraud detection systems by comparing classifier vulnerability profiles and assessing adversarial training as a mitigation strategy. Using the IEEE-CIS Fraud Detection dataset, comprising 590,540 transactions with a fraud incidence of 3.5%, four classifiers—logistic regression, random forest, gradient boosting, and a feed-forward neural network—were trained under identical preprocessing and class-weighting conditions and then subjected to Fast Gradient Sign Method and Projected Gradient Descent attacks at a perturbation budget of 0.02. Adversarial examples were constructed directly using closed-form and backpropagated gradients for the differentiable classifiers and using a logistic regression surrogate for the non-differentiable ensembles, before adversarial training was applied as a post-attack mitigation stage. Logistic regression proved the most adversarially vulnerable architecture, sustaining a 31.12-percentage-point recall loss under Projected Gradient Descent, while adversarial training subsequently restored its recall from 0.39 to 0.999 at an accuracy cost of 0.10 percentage points. Random forest and gradient boosting were not degraded by the surrogate-based attack, indicating that comparative robustness claims for tree-based ensembles require attack methods suited to their non-differentiable structure rather than transfer-based evaluation alone. Within the scope of this single-dataset evaluation, the findings support the adoption of adversarial training for gradient-based fraud detection models and suggest that robustness claims should be accompanied by disclosure of the attack methodology used to establish them.
Ololade Zainab Adesokan, Abiola Omolola Bamsa, O. Obioha-Val et al.· Journal of Engineering Resea...· 0 citations
A novel framework for adversarial machine unlearning is introduced to enable privacy-preserving threat intelligence sharing and lays the foundation for secure and compliant knowledge transfer in federated security operations and collaborative defence ecosystems.
R. Polishetty· Journal of Intelligent Decis...· 0 citations
Machine learning-based network intrusion detection systems (ML-based NIDS) are vulnerable to adversarial evasion, where malicious samples are perturbed to evade detection and be misclassified as benign. Despite growing research on adversarial attacks and defenses for ML-based NIDS, comparative evaluations of multiple attack types, detection models, and defense strategies under a common setting remain limited. In this paper, we evaluate eight adversarial evasion attacks, fifteen detection models, and three representative defense strategies using the NF-UQ-NIDS dataset, which includes recent traditional and IoT network traffic with twenty distinct attack categories. The evaluation compares model performance on clean test data and on robustness evaluation sets that include adversarial samples, analyzes attack success consistency across models, and examines the effect of defense strategies on adversarial robustness. To support model comparison, we introduce the Robustness Index (RI), a compact comparative metric that rewards high balanced accuracy and macro-F1 score computed on the robustness evaluation set while penalizing high attack success rate (ASR). We further present AR-NIDS, a two-stage framework that uses an adversarially trained ensemble to distinguish normal, attack, and adversarial samples, followed by an adversarial attack classifier to identify the attack type. Under the evaluated transfer-based setting, the proposed adversarially trained ensemble achieves the strongest overall trade-off between classification performance on the robustness evaluation set and evasion resistance, reducing the average ASR from 0.41 to 0.03 and achieving an RI of 0.98.
Unknown authors· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.