Skip to content
#graph neural networks Dataset Open access

Dataset: Integrating Structural and Semantic Analysis for Code Smell Refactoring Prediction

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

The project bridges traditional static code analysis (Object-Oriented metrics) with advanced embedding representation learning to predict refactoring interventions on code-smelly methods. The predictive framework models a supervised binary classification task across 10 distinct target refactoring operations. Data Source The underlying dataset (the source for the data collection pipeline) is made of 79 open-source Java systems. Full list included in dataset.md. Dataset (pipeline output) & Feature Space Overview Object-Oriented Metrics & Code Smells: code smells mined via DesigniteJava alongside a comprehensive suite of 74 Object-Oriented metrics computed by DesigniteJava and the CK Analysis Tool. Control Flow Topologies (Embeddings Option A): Granular Control Flow Graphs (CFGs) parsed via Joern and translated into 64-dimensional dense vectors using the LINE graph embedding algorithm. Joint Syntactic-Semantic Bytecode (Embeddings Option B): Normalized token streams and Program Dependence Graphs (PDGs) extracted directly from the compiler intermediate representation layer (Jimple bytecode) via GraphCode2Vec. Replication Artifacts IncludedThe package is organized to ensure complete scientific reproducibility and the code zip file includes: src/core/: Core execution modules for chronological commit mining (RefactoringMiner alignment), smell resolution detection, and target label processing. src/runners/: Batch automation wrappers and full Linux/WSL execution pipelines (Joern parsing, edge list generation, and node embedding training). src/utils/: Feature engineering helpers, mean-pooling scripts for graph node aggregation, and multi-source CSV mergers. src/models_pipeline/: Machine learning workflows implementing dataset balancing/scaling, Bayesian hyperparameter tuning, model training (Logistic Regression, SVM, XGBoost, Deep Neural Networks), and Late Fusion stacking ensembles, along with ablation and explainability outputs (ROC, Calibration, UMAP, SHAP). For detailed environment setup, dependency configurations (such as Joern, GraphCode2Vec, RefactoringMiner, CK, and Designite), and step-by-step reproduction instructions, please refer to the included README.md file. Additionally, you can find the raw datasets corresponding to the different data collection pipeline configurations within the classification_report zip file, allowing you to work directly on model configuration and training. Alternatively, you can use the files in the refactoringminer zip file, which provide chronological commit sequences defining the analysis time windows, as pre-computed outputs to save you from running the tool yourself, enabling you to operate on the entire pipeline and perform the collection independently.

View source

Similar papers

#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoption of Agile methods in general, and Scrum in particular. Little, if anything, is empirically known about the application and adoption of Scrum in a multi-team and multi-project situation. The authors carried out an ethnographically informed longitudinal case study in industrial settings and closely followed how the Scrum method was adopted in a 20-person department, working in a simultaneous multi-project R&D environment. Altogether 10 challenges pertinent to the case of multi-team multi-project Scrum adoption were identified in the study. The authors contend that these results carry great relevance for other industrial teams. Future research avenues arising from the study are indicated.

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11
#computer vision Open access Sep 2012

Making the leap to a software platform strategy: Issues and challenges

A comprehensive taxonomy of the challenges faced when a medium-scale organization decided to adopt software platforms is provided, namely: business challenges, organizational challenges, technical challenges, and people challenges.

Yaser Ghanam, F. Maurer, P. Abrahamsson · 41 citations · ⚡3
#machine learning Open access Mar 2024

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction

MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently, offers an effective and efficient solution for PPI overall property predictions.

Yang Yue, Shu Li, Yihua Cheng et al. · 15 citations

PepPCBench is a Comprehensive Benchmarking Framework for Protein-Peptide Complex Structure Prediction

PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction, and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy.

Si-Long Zhai, Huifeng Zhao, Ji-Ke Wang et al. · 13 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

OmniMol is presented, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability, and achieves state-of-the-art performance in properties prediction.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 11 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.