Skip to content

LLM-Mestcase: A Test Case Modification Method Based on LLMs and Mutation Testing

Sep 2026 · International journal of software engineering and knowledge engineering · 0 citations
Software Testing and Debugging Techniques

TL;DR

An LLM-based, mutation testing-driven approach for test case generation by integrating the semantic understanding of large language models with the precise evaluation mechanism of mutation testing, paving a new path for intelligent test case enhancement.

Abstract

High-quality test cases are crucial for ensuring software reliability. However, automatically generating test cases that can precisely detect semantic defects remains challenging. Mutation testing, by evaluating a test suite's ability to detect artificially seeded faults (mutants), provides an effective benchmark for measuring and improving test quality. Nevertheless, leveraging the results of mutation testing, especially surviving mutants, to automatically guide test case generation and enhancement faces two major obstacles: 1) the semantic understanding gap-inferring high-level fault semantics from low-level syntactic variations of surviving mutants; and 2) the context integration challenge-generated test cases must align with the project's existing coding style, testing framework, and domain-specific knowledge. To address these challenges, this paper proposes an LLM-based, mutation testing-driven approach for test case generation. Our method first employs an Analysis Agent to perform semantic clustering and fault attribution on surviving mutants, identifying genuine fault patterns. Subsequently, a Generate Agent produces targeted test cases that are stylistically consistent and domain-relevant, leveraging the context of the original code and existing test suite. Finally, a Check Agent validates the effectiveness of the newly generated cases. We conducted experiments on a specialized code dataset from the coal industry. The results demonstrate that our approach achieves an accuracy of 84.41% in identifying the causes of surviving mutants, a passing rate of 81.71% in generated test cases, and eliminates 55.66% of the surviving mutants, significantly enhancing the fault detection capability of the test suite. This work paves a new path for intelligent test case enhancement by integrating the semantic understanding of large language models with the precise evaluation mechanism of mutation testing.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models

Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-bas...

Nils Kiele, Zainab Saad, Zi-Rui Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation

Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubri...

Jiayuxuan Yang, Jie M. Zhang, Yiling Lou et al. · 0 citations
Preprint Sep 2026

Evaluating the effectiveness of class-level LLM-generated test suites in Python

Reliable assessment of LLM-generated tests should treat executability as a gate and combine coverage with mutation testing and structural quality indicators, and in practice, model selection should precede prompt tuning.

Bilal Al-Ahmad, M. Harshvardhan, Khaled El-Fakih et al. · 0 citations
#software testing Preprint Sep 2026

How effective are traditional test criteria at detecting bugs in large language models generated code?

An empirical study involving 5 Large Language Models and 4 benchmarks evaluates the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing, finding that mutation testing only marginally outperforms traditional coverage criteria in both triggering and d...

Asma Hamidi, Michael Konstantinou, R. Degiovanni et al. · 0 citations
Conference Aug 2026

A Mutation-Driven Trustworthiness Evaluation Method for LLM-Based Airborne Code Generation

Large language models (LLMs) have demonstrated significant potential for airborne embedded code generation, yet existing evaluation methods based primarily on the Pass@k metric focus on functional correctness and fail to adequately assess robustness degradation under semantic perturbations or identify domain-specific c...

Tong Zhang, Xing-Zhi Wang, Kang-Fa Xu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.