Jul 2026· Annual International Computer Software and Applications Conference· pp. 406-415· 0 citations· 28 references
Computer Science
Abstract
User stories serve as the core requirement carriers in agile development, and their quality directly determines the development team's understanding of requirements, as well as software delivery efficiency and quality. With the expansion of software specifications, user story quality evaluation is constrained by issues such as subjectivity and low efficiency in manual assessment, as well as lack of comprehensiveness and insufficient accuracy in the evaluation dimensions of automated methods, such as Natural Language Processing(NLP), Machine Learning(ML), and Large Language Models (LLM). This paper proposes an LLM-interpretable user story quality evaluation framework. Constructed by integrating the 3C principle, IN-VEST criteria, and IEEE 830 standards, the framework divides 40 core quality attributes into four dimensions: Structural Specification, Content Specification, Logical Soundness, and Actionability & Traceability (collectively SCLAT framework). To complement the framework, an LLM-interpretable application paradigm is proposed (known as K-CoT). This paradigm provides detailed, LLM-tailored designs for core attributes, introduces a negative example guidance mechanism, and a Chain-of-Thought (CoT) prompting strategy. Experimental validation on the NFDI4Cat dataset shows that the SCLAT framework enables LLMs to achieve an average F1-score of 89.9%, representing a significant improvement over traditional frameworks. The experimental results indicate that the interpretability design for LLM can comprehensively improve the dimensions, efficiency, and quality of user story defect detection.
This work presents the first cross-task empirical evaluation of LLMs spanning five RE-related activities, as well as replication materials supporting reproducibility, and a broader understanding of the capabilities, limitations, and practical readiness of current LLMs for RE.
Jacek Dabrowski, Manjeshwar Aniruddh Mallya, Alessio Ferrari et al.· 0 citations
Taxonomies provide a shared conceptual framework for organizing heterogeneous observations in software engineering (SE) research. Manually constructing such taxonomies is labor-intensive and requires annotators with expertise in the SE domain. While advances in Large Language Models (LLMs) have led to the emergence of...
Sota Nakashima, Yuta Ishimoto, Masanari Kondo et al.· 0 citations
Many requirements engineering (RE) tasks, such as requirements elicitation, documentation, and quality assessment, are inherently context-sensitive: what counts as a missing, wellwritten, or defective requirement varies by stakeholder-intent, domain, process, as well as countless other potential factors. Existing autom...
Max Unterbusch· IEEE International Requireme...· 0 citations
Interviews with sixteen early-adopter software professionals who integrated LLM-based tools into their day-to-day work in early to mid-2023 offer actionable implications for developers, organizations, educators, and tool designers seeking to integrate LLMs responsibly into professional software practice.
Benyamin T. Tabarsi, Heidi Reichert, Sam Gilson et al.· Empirical Software Engineeri...· 22 citations· ⚡1
Requirements engineering is a critical phase of software development that directly affects project scope, cost, and quality. In software development companies, a requirements list is typically prepared before creating a commercial proposal and signing a contract. EventStorming workshops are widely used for requirements...
Mantas Jurgelaitis, Antanas Ramanauskas, Gvidas Ambrozaitis et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.