Skip to content

Identifying key classes using biterm-sensitive PageRank with motif-based coreness

Aug 2026 · ACM Transactions on Software Engineering and Methodology · 0 citations · 70 references

Abstract

Key classes serve as excellent starting points for developers, particularly newcomers, to comprehend unfamiliar software systems. Although numerous unsupervised learning approaches for identifying key classes have been proposed in the literature—typically by representing software as class dependency networks (aka software networks) and leveraging various network metrics (e.g., \(h\) -index, \(a\) -index, coreness, and PageRank value) —these approaches fail to account for both the information available in diverse software artifacts, such as user manuals, and higher-order information captured by small network subgraphs (e.g., motifs) in software networks. In this paper, we introduce BiMo, a novel approach for identifying key classes in Java projects. First, we utilize a c lass-level a ggregated n etwork ( \({\rm CAN}\) ) to precisely represent classes and their inter-class couplings. Second, we extract biterms (i.e., co-occurred word pairs) from both user manuals and class source code, and further propose a b iterm-based s e mantic similarity metric, BE, to quantify the semantic similarity between user manuals and classes. Third, we detect all 3- and 4-node motifs within the \({\rm CAN}\) and further utilize them to construct a m otif-based c lass c oupling n etwork (MCCN); based on the MCCN, we propose a m otif-based c oreness metric, MC, to characterize the higher-order information of classes in the MCCN. Fourth, based on both the BE and MC metrics, we develop a novel PageRank variant, BiCoRank, to measure the PageRank value of classes in the \({\rm CAN}\) as the importance of classes. Finally, we sort classes in descending order according to their importance, and a cutoff is used to filter out non-key classes: the top-ranked classes are the recommended key classes. Empirical experiments performed on a set of 18 open-source Java projects show that i) BiMo significantly outperforms all baselines with significant differences; ii) BiMo is robust against different weighting mechanisms used to assign strengths for different relationship types in the \({\rm CAN}\) ; iii) both the MC and BE metrics significantly enhance the effectiveness of BiMo; iv) BiMo exhibits strong scalability and is well-suited for application in large-scale systems; and v) BiMo shows significant promise in real-world software comprehension tasks.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.