Automated LLM-based Classification of Software Requirements
The adoption of large language models (LLMs) in software engineering has enabled the potential to automate complex activities such as requirements analysis. This paper presents an empirical performance analysis of four modern LLMs: GPT-4o, Aya, Gemma and Phi-4 on the task of automated classification of atomic software requirements. The LLMs were invoked under two different scenarios to solve a multilabel classification of 296 requirements extracted from the PROMISE[Formula: see text] dataset. The baseline scenario relies solely on internal model knowledge, whereas the rubric-augmented scenario uses formal definitions derived from the SQuaRE product quality model. The results indicate that GPT-4o consistently attains the highest overall classification accuracy under both scenarios. Moreover, all LLMs exhibit strong and stable performance in identifying functional, performance efficiency and security requirements. Inter-rater agreement assessments using Cohen’s and Fleiss’ Kappa coefficients further demonstrate moderate to substantial agreement among the outputs of the evaluated LLMs.