Skip to content
Preprint

Human-AI Collaboration in Requirements Engineering: Evidence of the Negative Effect of LLMs on Requirements Inspection

Aug 2026 · 0 citations · 61 references
Computer Science

TL;DR

The findings provide empirical evidence that LLM support does not necessarily improve performance and may, instead, hinder it for novice inspectors, and suggest that learning RI with LLM-support from the beginning may slow down the skill acquisition process.

Abstract

Background. Requirements inspection (RI) is a well-established practice for detecting potential defects in requirements artifacts early in the software lifecycle. Recent advances in large language models (LLMs) have stimulated interest in their potential to support requirements engineering (RE) tasks. However, empirical evidence on the effects of LLMs when used as collaborative assistants in human-performed RI remains scarce. Aims. We aim to investigate the impact of LLM support on human-performed RI, considering inspection effectiveness in terms of smell identification and severity classification (i.e., nocuous vs innocuous), as well as inspection duration. Method. We conducted a controlled crossover design experiment with 34 participants, who inspected textual specifications with and without LLM support, identifying and classifying requirements smells while recording inspection time. We analyzed the data using one Bayesian regression model per outcome variable, accounting for validity threats induced by the crossover design as well as covariates and mediators. Results. Results show that LLM support negatively affects smell detection accuracy but has no significant effect on smell classification or task duration. A learning effect is present across experimental periods, but reduced when RI is first performed with LLM support. Conclusions. Our findings provide empirical evidence that LLM support does not necessarily improve performance and may, instead, hinder it for novice inspectors. Moreover, the results suggest that learning RI with LLM-support from the beginning may slow down the skill acquisition process, implying threats for LLM-supported learning.

View source

Similar papers

Sep 2026

Automated LLM-based Classification of Software Requirements

The adoption of large language models (LLMs) in software engineering has enabled the potential to automate complex activities such as requirements analysis. This paper presents an empirical performance analysis of four modern LLMs: GPT-4o, Aya, Gemma and Phi-4 on the task of automated classification of atomic software requirements. The LLMs were invoked under two different scenarios to solve a multilabel classification of 296 requirements extracted from the PROMISE[Formula: see text] dataset. The baseline scenario relies solely on internal model knowledge, whereas the rubric-augmented scenario uses formal definitions derived from the SQuaRE product quality model. The results indicate that GPT-4o consistently attains the highest overall classification accuracy under both scenarios. Moreover, all LLMs exhibit strong and stable performance in identifying functional, performance efficiency and security requirements. Inter-rater agreement assessments using Cohen’s and Fleiss’ Kappa coefficients further demonstrate moderate to substantial agreement among the outputs of the evaluated LLMs.

Nourchène Elleuch Ben Ayed, Jaber Jemai, Keletso J. Letsholo et al. · 0 citations
Review Sep 2026

The Impact of GenAI on the Future of Requirements Engineering

Recent advances in artificial intelligence (AI), particularly large language models (LLMs), are transforming how we design and build systems by increasing access to domain knowledge and by providing automation support to software engineering (SE). As implementation becomes less expensive through generalist SE agents, engineering effort shifts away from writing correct code and toward expressing, curating, verifying, and evaluating requirements. In this paper, we survey the state of the art in AI for requirements engineering (RE) research leading up to the transformation, before reviewing advances in LLMs. We survey two subsequent research areas: prompt programming, which treats LLM instructions as a program in SE vernacular, and generalist SE agents, which combine multiple LLM advances to yield semi-autonomous processes that complete SE tasks. Finally, we explore the future of requirements engineering along two axes: matters changing how we interact with requirements through the SE process, and matters changing how requirements are experienced by software developers and stakeholders more broadly, including end-users. This article aims to inform how RE researchers can navigate this transformation in the selection of future research priorities.

Travis D. Breaux, Anmol Singhal · 0 citations
Preprint Aug 2026

Large Language Models for Requirements Engineering: A Cross-Task Empirical Evaluation

This work presents the first cross-task empirical evaluation of LLMs spanning five RE-related activities, as well as replication materials supporting reproducibility, and a broader understanding of the capabilities, limitations, and practical readiness of current LLMs for RE.

Jacek Dabrowski, Manjeshwar Aniruddh Mallya, Alessio Ferrari et al. · 0 citations
Open access 2026

A Comparative Study of LLMs and Human Judgment in UML Diagram Evaluation

This paper investigates the reliability of LLMs in evaluating UML diagrams generated through reverse engineering processes (source code) and asks: do LLM assessments align with those of human experts?

Olena Chebanyuk, Carles Sierra · 0 citations

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.

Miguel Zabaleta, Bai-Han Lin · 0 citations
Review Open access Aug 2026

Productivity vs. Compliance: The New Engineering Challenge of AI Coding Assistants in Regulated Codebases

A codepath-aware governance framework for AI-assisted engineering in regulated codebases, with emphasis on financial services, payments, healthcare, and other domains where software changes may affect legal, operational, privacy, and audit obligations is developed.

Ashutosh Pal · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.