Skip to content

From Regulation to Requirements: An Automated Requirement Derivation and Explanation Pipeline

Jul 2026 · arXiv.org · Vol abs/2607.04448 · 0 citations · 34 references
Computer Science

TL;DR

An automated regulation-to-requirements pipeline that identifies requirement-bearing clauses in regulatory documents and derives system-agnostic software requirements, accompanied by plain-language explanations, traceable to their legal sources is presented.

Abstract

Ensuring software compliance with regulations such as the General Data Protection Regulation (GDPR) and the Artificial Intelligence Act (EU AI Act) poses a significant challenge, as requirements engineers must translate complex legal text into actionable software requirements - a process that remains largely manual and error-prone in practice. We present an automated regulation-to-requirements pipeline that identifies requirement-bearing clauses in regulatory documents and derives system-agnostic software requirements, accompanied by plain-language explanations, traceable to their legal sources. We evaluate the pipeline on the full clause sets of the GDPR (398 clauses) and the EU AI Act (574 clauses). For requirement-bearing clause identification, the approach achieves macro-averaged F1 scores of 0.82 and 0.78, respectively, outperforming a SetFit-based baseline. Human evaluation shows high completeness (4.60 and 4.45) and correctness (3.74 and 3.54) of derived requirements, while explanation clarity scores are near-ceiling (4.92 and 4.94) on a 1-5 scale. We implement the approach in Reg2Req, a publicly released tool that further supports requirement classification, use case seeding, cross-reference analysis, definition indexing, and a traceability matrix to operationalize regulatory compliance in practice. A user study with 25 practitioners shows that the plain-language explanations significantly improve comprehension of derived requirements and confidence in acting on them (p<0.001), and that all participants would use Reg2Req as a starting point for deriving software requirements from a regulation.

View source

Similar papers

Open access Jul 2026

RAG-Based AI Compliance Monitoring and Report Generation System

CompVault, an Enhanced Retrieval-Augmented Generation (ERAG)-based Artificial Intelligence Compliance Monitoring and Report Generation System for intelligent regulatory compliance assessment, and results indicate that the ERAG-based framework can be used as an efficient, scalable, and explainable solution for regulatory compliance monitoring and automated report generation.

S. N., Sathyapriya P., Vishnu Priya R. M. et al. · 0 citations
Review Open access Jul 2026

When may LLM outputs influence software requirements? A human-in-the-loop governance framework

Large language models are increasingly used to review, clarify, rewrite, and trace software requirements. These applications create a governance problem that output-quality assessment alone cannot resolve: a fluent proposal may rely on inadmissible evidence, alter stakeholder intent, introduce unsupported specificity, or imply an organizational commitment that the model has no authority to make. Existing work on retrieval-augmented generation, controlled natural language, formal verification, human oversight, and AI governance supplies relevant controls, but it does not specify the procedural status of an individual LLM proposal relative to a controlled requirements artifact. This article develops an artifact-centered, human-in-the-loop framework in which the permitted influence of a proposal is the primary object of governance. The framework combines five governance functions—governed evidence, bounded context construction, controlled LLM analysis, pre-commit verification, and accountable human approval—with four artifact-influence states: A0 advisory observation, A1 evidence-linked candidate, A2 verified recommendation, and A3 approved and committed change. Its central theoretical claim is that output quality, evidential legitimacy, verification status, and authority to commit a change are distinct properties and should not be collapsed into a single confidence judgment. Seven falsifiable hypotheses translate the model into measurable comparisons involving source admissibility, context leakage, unsupported specificity, semantic drift, reviewer agreement, unreviewed changes, governance cost, and organizational maturity. Human review is treated as both a necessary decision boundary and a potential source of automation bias, anchoring, and fatigue. The framework is conceptual rather than empirically validated and provides a basis for controlled experiments, field studies, and longitudinal evaluation.

Chuanjin Zhu · 0 citations
Conference Jul 2026

From Regulatory Text to Executable Constraints: Operationalising Compliance for Medicine Labelling

Regulatory requirements governing safety-critical health domains, such as medicine labelling, are predominantly expressed in legal prose with strong deontic modality (e.g., obligations and prohibitions). While suitable for human interpretation, such requirements are not directly machine-interpretable, limiting their use as deterministic, executable constraints in label design workflows. We introduce a machine-verifiable methodology that operationalises medicine labelling regulations into structured, executable compliance contracts. Legal provisions are systematically extracted, filtered, and translated into normative representations that specify content, structural, and format constraints. These representations are compiled into deterministic checks, enabling regulations to serve as executable constraints on design artefacts and supporting systematic analysis of compliance. We instantiate the methodology using U.S. Over-the-counter and Prescription labelling requirements (21 CFR § 201.66 and 21 CFR § 201.100). Under this process, we operationalise 16 from 81 design-related requirements in OTC regulations, and 7 from 12 in prescription regulations. We further introduce a tiered notion of evaluability over the operationalised subset, distinguishing requirements that are directly artefact-evaluable under a structured SVG representation from those that depend on unresolved product context or representation constraints. Our analysis reveals that only 7.4% of OTC requirements (6/81) and 16.7% of operationalised prescription requirements (2/12) are fully automatable, with the majority requiring additional product-context modelling or information outside of the regulation document. This exposes a fundamental limitation of artefact-level evaluation: many regulatory requirements are not intrinsically evaluable without resolving dependencies on product context, representation, and external regulatory sources. These findings highlight the need for joint modelling of artefact structure and product context, providing a foundation for scalable, auditable, and trustworthy validation of generative healthcare artefacts.

Frank Tran, Paul Benjamin Ramirez, Vinesh George et al. · 0 citations
Jul 2026

Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring

Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate structurally invalid requirements and logical inconsistencies, lacking formal correctness guarantees. Objectives: This study aims to eliminate logical inconsistencies and enforce structural conformance in LLM-generated requirements while quantifying the LLM's pre-validation decision uncertainty within a formal domain model. Methods: We present a neuro-symbolic multi-agent architecture that operationalizes the Object-Oriented Method for Requirements Authoring and Management (OOMRAM) lattice. The LLM acts as a non-deterministic heuristic for lattice traversal, while a deterministic symbolic validator enforces all structural constraints. We introduce a three-valued (T, I, F) -- Truth, Indeterminacy, Falsity -- framework to classify and score the LLM's requirement decisions before and after validation. Results: Evaluated across 37 natural-language project visions in eleven application families, the system completely eliminated structural inconsistencies in 35 out of 37 cases (94.6%), with the remaining two containing only 6 unresolved structural errors (0.39% of decisions) due to iteration limits. Three-valued analysis revealed that 24.7% of all decisions are indeterminate -- structurally valid but discretionary choices not explicitly mandated by the stakeholder. Conclusion: Offloading structural integrity to a deterministic symbolic layer successfully guarantees structural conformance, while the three-valued classification provides a formal way to measure neural uncertainty, facilitating safe LLM deployment in formal requirements engineering.

A. Ibrahim · 0 citations
Case report Open access Jul 2026

Controlled, Not Correct A Computer Software Assurance framework for domain-specific regulatory AI agents

Current discourse on artificial intelligence in regulatory affairs is organised around a single variable: model accuracy. Benchmarks assert accuracy thresholds, vendors claim accuracy figures, and adopters ask whether the model is correct. This paper argues that accuracy is the wrong variable, and that the regulated question — the one that has governed every other tool in a quality management system for fifty years — is whether the tool's output is subject to adequate control proportionate to its intended use and risk. The paper derives a three-layer human-in-the-loop control architecture from existing regulatory instruments (FDA Computer Software Assurance draft guidance 2022; ISO 13485:2016 Clause 4.1.6; MDCG 2019-11 rev.1; EU AI Act Article 10) rather than proposing a new framework. It then quantifies the architecture using a defect-escape model, and reports a result with direct consequences for how regulatory AI should be evaluated: A model at the 3σ baseline of human expert cognition (93.3% accuracy, 66,800 DPMO), placed under three mature control layers, yields a residual defect rate of 401 DPMO — an effective process yield of 99.960%. This is functionally equivalent to accuracy thresholds currently asserted as unreachable by AI systems. The threshold is reachable. It was never a property of the model. The same model under a single control layer yields 13,360 DPMO — worse than an unvalidated 99% model. The number of control layers, not the accuracy of the model, is the dominant term.

Rudolf Wagner · 0 citations
Review Open access Aug 2026

Integrating large language models and knowledge graphs for adaptive design review

This paper presents a framework integrating Knowledge Graphs and Large Language Models to support a more extensible design review environment, and demonstrates its ability to retrieve and execute existing rules from the KG, capture new requests during design, and maintain a verifiable, adaptive compliance checking system.

Maen Alnuzha, Tanya Bloch · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.