Intelligent Real-Time Fraud Detection in Financial Institutions
Abstract
Financial institutions process high transaction volumes through digital payment channels, but conventional rule-based fraud detection systems often fail to adapt to changing fraud patterns and may produce excessive false alerts. This study developed and evaluated an intelligent real-time stacked ensemble prototype for credit card fraud detection in financial institutions. The system combined supervised fraud probability estimates from Random Forest and XGBoost with anomaly score evidence from Isolation Forest, and used Logistic Regression as a meta-classifier to generate the final transaction-level decision. The prototype was implemented in Python using Pandas, NumPy, scikit-learn, imbalanced-learn, XGBoost, Seaborn, Matplotlib, and Joblib. The Credit Card Fraud Detection dataset was used for experimentation, containing 284,807 transactions, 30 input features, and 492 fraudulent cases. The data were split using stratified 80:20 sampling; Time and Amount were standardised using training set statistics only, and SMOTE was applied only to the training partition to prevent test-set leakage. The held-out test set retained the original imbalanced distribution of 56,864 legitimate transactions and 98 fraudulent transactions. Random Forest achieved the highest precision and F1-score, with precision of 0.8454, recall of 0.8367, F1-score of 0.8410, false positive rate of 0.00026, and AUC-ROC of 0.9731. XGBoost achieved the highest base-model recall of 0.8673 and AUC-ROC of 0.9750, but produced 136 false positives. Isolation Forest achieved AUC-ROC of 0.8532 but failed to detect fraudulent transactions as a standalone binary classifier under the implemented configuration. The stacked ensemble achieved accuracy of 0.9993, precision of 0.7778, recall of 0.8571, F1-score of 0.8155, false positive rate of 0.00042, and AUC-ROC of 0.9758. The results show that the stacked ensemble produced a usable prototype-level fraud decision layer by combining supervised and anomaly based evidence, although threshold calibration, explainability, and validation on local institutional data remain necessary before operational deployment.