Skip to content
Preprint

Evaluating the effectiveness of class-level LLM-generated test suites in Python

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

Reliable assessment of LLM-generated tests should treat executability as a gate and combine coverage with mutation testing and structural quality indicators, and in practice, model selection should precede prompt tuning.

Abstract

Context: Large language models (LLMs) can generate unit tests quickly, but high structural coverage does not establish that those tests execute reliably or detect faults. Existing evidence often treats coverage as the principal outcome and rarely compares prompt strategies and models through mutation testing at class level. Objective: This study examines how prompt strategy and model choice shape the executability, structural coverage, fault-detection effectiveness, and structural quality of LLM-generated Python test suites relative to human-written suites. Method: We evaluate multiple prompt strategies across a diverse set of current LLM configurations on the ClassEval benchmark. The evaluation combines execution outcomes, line and branch coverage, Cosmic Ray mutation scores, and structural quality indicators. Primary analyses treat successful execution as a prerequisite; paired comparisons use only classes shared by the relevant executable subsets. Results: Structural coverage is consistently near its ceiling and offers little discrimination among configurations. Executability varies substantially. The proposed prompt performs strongly for mutation score, but no prompt dominates across models. Model choice explains more variation than prompt choice, and their interaction shows that prompt effectiveness depends on the selected model. Human and LLM suites are evaluated on unequal executable subsets, so their relative mutation scores do not establish superiority. Conclusion: Reliable assessment of LLM-generated tests should treat executability as a gate and combine coverage with mutation testing and structural quality indicators. In practice, model selection should precede prompt tuning.

View source

Similar papers

#software testing Preprint Sep 2026

How effective are traditional test criteria at detecting bugs in large language models generated code?

An empirical study involving 5 Large Language Models and 4 benchmarks evaluates the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing, finding that mutation testing only marginally outperforms traditional coverage criteria in both triggering and d...

Asma Hamidi, Michael Konstantinou, R. Degiovanni et al. · 0 citations

Evaluation of large language models in api testing using open API documentation

Conventional automated REST API testing approaches often depend on rule-based logic, extensive configuration, or source-code access, which limits their adaptability in rapidly evolving development environments. Recent advances in Large Language Models (LLMs) offer new possibilities for automating API test generation th...

Phyu Thet Thet Kyaw · 0 citations
#artificial intelligence Preprint Sep 2026

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snipp...

Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband et al. · 0 citations
Open access Aug 2026

Optimizing Context and Cost in LLM‐Based Unit Test Generation: A Study on External Dependency Retrieval Strategies

A systematic empirical study of multiple strategies for context enrichment and optimization in LLM‐based unit test generation, conducted on seven diverse projects (three open‐source and four proprietary industrial systems), encompassing 261 distinct methods establish this optimized context strategy as a cost‐effective...

Javier Ferrer, Francisco Chicano · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models

Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-bas...

Nils Kiele, Zainab Saad, Zi-Rui Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.