Project-Scale Statement-Level Fault Localization via Multi-view Semantic Learning and Pairwise Reranking
Abstract
Statement-level fault localization (FL) is critical for effective software debugging, as it enables developers to precisely identify faulty lines of code. While traditional spectrum-based, mutation-based, and deep learning-based FL techniques have achieved notable progress, they remain limited in modeling rich fault semantics. Recent advances in large language models (LLMs) offer new opportunities for FL due to their strong capacity for semantic understanding and reasoning over bugs. However, existing LLM-based FL approaches largely treat fault localization as isolated, code-centric prediction, limiting their ability to perform the holistic fault reasoning required for precise statement-level localization. In this paper, we propose FaultScape, a novel LLM-based framework for project-scale, statement-level FL that formulates FL as a multi-view semantic learning and reasoning problem. FaultScape addresses the limitations of existing approaches through three key components. First, we introduce a joint contrastive fine-tuning strategy that trains LLMs on large-scale bug-fix data to explicitly learn fault semantics from multiple complementary views. View-specific fault semantics, including code semantics, fault type, root cause, and repair intent, are learned via supervised binary classification, while cross-view semantic consistency is enforced through contrastive learning. These two objectives are jointly optimized within a unified training loss. The fine-tuned models extract multi-view fault likelihoods as semantic features for each suspicious statement. Second, we adopt a dynamic feature integration module that combines these semantic features with spectrum-based and mutation-based execution features, producing an initial suspiciousness statement ranking. Third, we design an LLM-based, test-guided pairwise re-ranking strategy that explicitly compares highly suspicious candidate statements using failing-test context. By estimating relative fault likelihoods through pairwise comparison rather than independent scoring, the model produces a refined statement-level ranking. We evaluate FaultScape on Defects4J v1.2.0, where it localizes 112/171/196 bugs at Top-1/3/5 out of 395, outperforming state-of-the-art DL-based and LLM-based baselines. On leakage-free benchmarks, FaultScape further localizes 11/18/20 bugs at Top-1/3/5 on ConDefects (31 bugs) and 12/21/22 bugs on GHRB (34 bugs), demonstrating strong generalization to unseen projects. These results show that combining multi-view fault semantics with contrastive, failure-guided reasoning substantially improves the effectiveness and robustness of statement-level fault localization.Statement-level fault localization (FL) is critical for effective software debugging, as it enables developers to precisely identify faulty lines of code. While traditional spectrum-based, mutation-based, and deep learning-based FL techniques have achieved notable progress, they remain limited in modeling rich fault semantics. Recent advances in large language models (LLMs) offer new opportunities for FL due to their strong capacity for semantic understanding and reasoning over bugs. However, existing LLM-based FL approaches largely treat fault localization as isolated, code-centric prediction, limiting their ability to perform the holistic fault reasoning required for precise statement-level localization. In this paper, we propose FaultScape, a novel LLM-based framework for project-scale, statement-level FL that formulates FL as a multi-view semantic learning and reasoning problem. FaultScape addresses the limitations of existing approaches through three key components. First, we introduce a joint contrastive fine-tuning strategy that trains LLMs on large-scale bug-fix data to explicitly learn fault semantics from multiple complementary views. View-specific fault semantics, including code semantics, fault type, root cause, and repair intent, are learned via supervised binary classification, while cross-view semantic consistency is enforced through contrastive learning. These two objectives are jointly optimized within a unified training loss. The fine-tuned models extract multi-view fault likelihoods as semantic features for each suspicious statement. Second, we adopt a dynamic feature integration module that combines these semantic features with spectrum-based and mutation-based execution features, producing an initial suspiciousness statement ranking. Third, we design an LLM-based, test-guided pairwise re-ranking strategy that explicitly compares highly suspicious candidate statements using failing-test context. By estimating relative fault likelihoods through pairwise comparison rather than independent scoring, the model produces a refined statement-level ranking. We evaluate FaultScape on Defects4J v1.2.0, where it localizes 112/171/196 bugs at Top-1/3/5 out of 395, outperforming state-of-the-art DL-based and LLM-based baselines. On leakage-free benchmarks, FaultScape further localizes 11/18/20 bugs at Top-1/3/5 on ConDefects (31 bugs) and 12/21/22 bugs on GHRB (34 bugs), demonstrating strong generalization to unseen projects. These results show that combining multi-view fault semantics with contrastive, failure-guided reasoning substantially improves the effectiveness and robustness of statement-level fault localization.