This work proposes LLM-based methods for verifying semantically complex NL requirements on static GUI prototypes and introduces a multimodal LLM-based agent for verifying complex functional and non-functional requirements in dynamic GUI applications through automatically generated and evaluated interaction trajectories.
Abstract
Requirements elicitation is essential for developing interactive software systems, as it helps ensure that the resulting product meets stakeholder needs. Since elicitation typically relies on natural language (NL), misunderstandings can arise from its inherent ambiguity. Formal specifications can reduce ambiguity but require technical expertise. GUI prototyping therefore provides a valuable alternative by turning requirements into tangible visual artifacts that support communication, elicitation, and validation. However, creating high-fidelity prototypes remains time-consuming and costly. Similarly, requirements verification, which ensures that implementations conform to specified requirements, is still largely manual, while existing automated approaches are often limited to static, rule-based techniques. This work addresses two challenges: (C1) reducing the effort required to transform NL requirements into GUI prototypes, and (C2) reducing the effort required for requirements verification in GUI applications and prototypes. For C1, we introduce novel NL-based GUI retrieval and reranking methods, new benchmarks, and techniques for efficiently adapting LLMs to GUI generation, including proprietary GUI representations. Their effectiveness is demonstrated on a large benchmark with human annotations. For C2, we propose LLM-based methods for verifying semantically complex NL requirements on static GUI prototypes and introduce a multimodal LLM-based agent for verifying complex functional and non-functional requirements in dynamic GUI applications through automatically generated and evaluated interaction trajectories. Overall, the proposed methods substantially reduce manual effort in GUI prototyping and requirements verification.
Context and Motivation] Semi-formal syntax templates for natural language requirements positively impact various requirements metrics, such as singularity. Using templates such as MASTeR or EARS also improves understandability. Requirements that conform to templates are easier for humans to understand than unrestricted requirements. [Question/Problem] However, converting requirements into templates is time-consuming and requires substantial prior training and in-depth domain knowledge. Thus, most requirements are still written in unrestricted natural language. [Principal Ideas and Results] Our approach is to use large language models (LLMs) to convert unrestricted natural language requirements into templates. Our evaluation demonstrates the proficiency of LLMbased systems. LLM-converted requirements are rated similarly to human rephrasings, especially for shorter requirements. Thus, they can be used to accelerate the requirements rephrasing process. [Contribution] In this paper, we present an approach for automatically parsing free-text requirements into templates with LLMs. We also provide an overview of metrics for automatically validating such systems and use them to compare our rephrased requirements with a ground truth. In a practitioner survey comparing LLM- and human-converted requirements, we assess the validity of these metrics.
Julian Roßkothen, Dominik Fuchß, Florian Erdösi et al.· IEEE International Requireme...· 1 citation
A systematic mapping study of 74 peer-reviewed primary studies on AI-based automated requirements elicitation published between 2021 and 2025, identified from five databases following PRISMA 2020 and classified by AI technique, textual source, elicitation activity, and application domain gives researchers and practitioners guidance on which techniques the evidence supports for each elicitation task and textual source.
Safaa Eltahier, S. Al-Ghuribi, Mawal A. Mohammed et al.· Information· 0 citations
An automated framework is proposed based on Natural Language Processing (NLP) techniques to parse the software requirements syntactically using a set of heuristic rules that facilitate the extraction of actors, use cases, entities, relationships, and attributes from software requirements documents written in natural language.
Thamer A. Alrawashdeh, Adnan Hnaif, Mustafa Alrifaee et al.· Journal of Communications So...· 1 citation
Elicitation activities such as workshops generate rich qualitative data; however their reuse in Requirements Engineering (RE) remains limited due to the lack of explicit structure and traceable interpretation. While modeling approaches such as i* and the NFR Framework define modeling constructs, they do not make explicit how model elements are derived from elicitation data. In this paper, we present an industrial case based on workshop transcripts and conduct an exploratory exercise of how RE engineers derive goals, tasks, and quality concerns from the same textual input. From a limited but representative sample, we identify recurring patterns in their interpretive reasoning and derive an initial set of heuristics that make this reasoning explicit. We further illustrate how these heuristics can be operationalized within a Retrieval-Augmented Generation (RAG) pipeline, enabling semantic and semi-automated processing grounded in RE engineers' interpretations while preserving traceability. Our results suggest that, in this exploratory setting, making interpretative reasoning explicit supports the reuse of elicitation data through requirements models.
R. Portugal, J. Leite, L. Silva et al.· Anais do Workshop em Engenha...· 0 citations
This work presents the first cross-task empirical evaluation of LLMs spanning five RE-related activities, as well as replication materials supporting reproducibility, and a broader understanding of the capabilities, limitations, and practical readiness of current LLMs for RE.
Jacek Dabrowski, Manjeshwar Aniruddh Mallya, Alessio Ferrari et al.· 0 citations
The results indicate that current general-purpose LLMs can achieve practically significant performance on the unstructured NL-to-LTL task without task-specific fine-tuning, and suggest that modern LLMs are becoming viable front-end assistants for semi-automated formalization workflows.
Alexandra Newcomb, Omar Ochoa· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.