Interpretable Fuzzy Software Fault Prediction
Abstract
This repository contains the complete source code and statistical analysis results for the paper: "A Data Driven Interpretable Fuzzy System with Three-Dimensional Optimization for Software Fault Prediction" Authors: Behrooz Shahi, Hooman TahayoriAffiliation: Department of Computer Science and Engineering & IT, School of Electrical and Computer Engineering, Shiraz University, Shiraz, Iran ============================================================FILES============================================================ 1. final.py - Main implementation of the proposed five-phase architecture: (1) Automatic Fuzzification using Fuzzy C-Means (FCM) (2) Automatic Rule Extraction using Random Forest (3) Three-Dimensional Optimization (features, rules, rule length) (4) Sugeno (TSK) Fuzzy Inference Engine with normalization and isotonic calibration (5) Defuzzification and Output - Runs all 27 datasets (24 numerical + 3 Eclipse) with 5 runs each - Compares the proposed Fuzzy method against 7 baseline methods: Random Forest, SVM, Naive Bayes, Decision Tree, Logistic Regression, XGBoost, and Deep Learning (MLP) - Outputs: results_mean_std.csv, statistical_tests_wilcoxon.csv 2. final1.py - Testing and evaluation script - Runs the full experimental pipeline and generates statistical test results - Used to produce the raw Wilcoxon signed-rank test results 3. Holm-Bonferroni.py - Post-processing script for statistical analysis - Applies the Holm-Bonferroni correction for multiple comparisons - Removes 3 Eclipse datasets (ECLIPSE_JDT_CORE, ECLIPSE_PDF_UI, EQUINOX_FRAMEWORK) to align with the 24 numerical datasets reported in the paper - Outputs: wilcoxon_with_holm.csv, wilcoxon_summary.csv, table_X_final.csv ============================================================DATASETS============================================================ The 24 numerical datasets are from three public repositories: - NASA MDP (5 datasets): JM1, KC1, KC2, PC1, PC2 - PROMISE (9 datasets): Ant, Camel, JEdit, Log4j, Lucene, Poi, Velocity, Xalan, Xerces - TravisTorrent (10 datasets): Cloudify, Graylog2, Jackrabbit, JRuby, Metasploit, Open-Build, OpenProject, Rails, Ruby, Sonarqube The qualitative Fenton dataset (31 projects) is used for direct comparisonwith existing fuzzy methods. ============================================================SUMMARY OF RESULTS============================================================ Total pairwise tests (24 datasets x 7 baselines): 168Significant (raw, p < 0.05): 140 (83.3%)Significant (Holm-Bonferroni corrected): 112 (66.7%) The optimized fuzzy system uses only 8 features and 50 rules, making ithighly interpretable, reproducible, and scalable for industrial deployment. ============================================================REQUIREMENTS============================================================ Python 3.12numpy, pandas, scikit-learn, scikit-fuzzy, scipy, statsmodels,xgboost, tensorflow, matplotlib Install with: pip install -r requirements.txt ============================================================USAGE============================================================ 1. Place datasets in: ./datasets/nasa/, ./datasets/promise/, ./datasets/travis/, ./datasets/eclipse/ 2. Run the main script: python final.py 3. Run the testing script: python final1.py 4. Run the statistical analysis: python Holm-Bonferroni.py 5. Results are saved to: ./results/csv/ ============================================================LICENSE============================================================ Code: MIT LicenseData: Creative Commons Attribution 4.0 International (CC BY 4.0) ============================================================CONTACT============================================================ Corresponding author: Behrooz ShahiEmail: b.shahi@shirazu.ac.irGitHub: https://github.com/Behsh123/Interpretable-Fuzzy-Software-Fault-Prediction