Sep 2026· International journal of software engineering and knowledge engineering· 0 citations
Software Testing and Debugging Techniques
TL;DR
An LLM-based, mutation testing-driven approach for test case generation by integrating the semantic understanding of large language models with the precise evaluation mechanism of mutation testing, paving a new path for intelligent test case enhancement.
Abstract
High-quality test cases are crucial for ensuring software reliability. However, automatically generating test cases that can precisely detect semantic defects remains challenging. Mutation testing, by evaluating a test suite's ability to detect artificially seeded faults (mutants), provides an effective benchmark for measuring and improving test quality. Nevertheless, leveraging the results of mutation testing, especially surviving mutants, to automatically guide test case generation and enhancement faces two major obstacles: 1) the semantic understanding gap-inferring high-level fault semantics from low-level syntactic variations of surviving mutants; and 2) the context integration challenge-generated test cases must align with the project's existing coding style, testing framework, and domain-specific knowledge. To address these challenges, this paper proposes an LLM-based, mutation testing-driven approach for test case generation. Our method first employs an Analysis Agent to perform semantic clustering and fault attribution on surviving mutants, identifying genuine fault patterns. Subsequently, a Generate Agent produces targeted test cases that are stylistically consistent and domain-relevant, leveraging the context of the original code and existing test suite. Finally, a Check Agent validates the effectiveness of the newly generated cases. We conducted experiments on a specialized code dataset from the coal industry. The results demonstrate that our approach achieves an accuracy of 84.41% in identifying the causes of surviving mutants, a passing rate of 81.71% in generated test cases, and eliminates 55.66% of the surviving mutants, significantly enhancing the fault detection capability of the test suite. This work paves a new path for intelligent test case enhancement by integrating the semantic understanding of large language models with the precise evaluation mechanism of mutation testing.
Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-bas...
Nils Kiele, Zainab Saad, Zi-Rui Wang et al.· 0 citations
Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubri...
Jiayuxuan Yang, Jie M. Zhang, Yiling Lou et al.· 0 citations
Reliable assessment of LLM-generated tests should treat executability as a gate and combine coverage with mutation testing and structural quality indicators, and in practice, model selection should precede prompt tuning.
Bilal Al-Ahmad, M. Harshvardhan, Khaled El-Fakih et al.· 0 citations
An empirical study involving 5 Large Language Models and 4 benchmarks evaluates the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing, finding that mutation testing only marginally outperforms traditional coverage criteria in both triggering and d...
Asma Hamidi, Michael Konstantinou, R. Degiovanni et al.· 0 citations
Large language models (LLMs) have demonstrated significant potential for airborne embedded code generation, yet existing evaluation methods based primarily on the Pass@k metric focus on functional correctness and fail to adequately assess robustness degradation under semantic perturbations or identify domain-specific c...
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.