Preprint
Sep 2026
Evaluating the effectiveness of class-level LLM-generated test suites in Python
Reliable assessment of LLM-generated tests should treat executability as a gate and combine coverage with mutation testing and structural quality indicators, and in practice, model selection should precede prompt tuning.
Bilal Al-Ahmad, M. Harshvardhan, Khaled El-Fakih et al.
· 0 citations