Aug 2026· ARES· pp. 401-419· 0 citations· 49 references
Computer Science
TL;DR
A lifecycle model for LLM systems is proposed that supports security analysis by structuring it around security-relevant boundaries rather than workflow optimisation, and is supported by a 12-stage LLMOps pillar and a 9-category governance pillar.
Abstract
Large language models are being integrated into critical infrastructure and enterprise workflows at unprecedented scale,yet the lifecycle frameworks governing their development and operations were designed for operational efficiency rather than security analysis. As a result, security-relevant activities such as data provenance verification, artifact signing, agentic permission control, and decommissioning are often left implicit or assumed to receive due care. Governance frameworks, in turn, organise requirements around risk levels or management processes without clearly linking them to the lifecycle stages where they apply. This paper addresses both deficiencies. We propose a lifecycle model for LLM systems that supports security analysis by structuring it around security-relevant boundaries rather than workflow optimisation. The model comprises 32 stages across four core pipeline layers (Data, Model, Distribution, Application), supported by a 12-stage LLMOps pillar and a 9-category governance pillar. Thirteen stages are introduced here as separate units because they expose distinct security concerns that existing frameworks do not clearly distinguish. A governance mapping synthesising the NIST AI RMF, the EU AI Act, and ISO/IEC 42001 reveals a structural property of the current regulatory landscape: governance evidence concentrates at deployment-facing stages, where systems are visible to regulators, while the most consequential decisions, data selection, alignment strategy, and capability boundaries, are made at development-facing stages, where regulatory visibility is lowest.
The rapid evolution of large language models (LLMs) has introduced new challenges in model development, deployment, monitoring, and governance. Traditional software-focused Continuous Integration (CI) pipelines are insufficient for managing the iterative and data-intensive lifecycle of LLMs, which require continuous data validation, model retraining, bias and safety auditing, reproducibility checks, and scalable deployment. This paper proposes a comprehensive CI pipeline architecture tailored to the unique requirements of LLM lifecycle management. The framework integrates automated data quality assessment, modular training workflows, version-controlled model artifacts, continuous evaluation against multi-dimensional metrics, and responsible AI checks including fairness, robustness, and alignment. We discuss implementation patterns using modern MLOps tooling, highlight operational challenges, and present best practices for ensuring reliability, traceability, and ethical compliance in LLM-centric systems. The proposed approach facilitates faster iteration cycles, safer model updates, and more efficient long-term governance of LLM deployments.
H. Mohamed· International Journal of Art...· 0 citations
The results show that lifecycle integrity can be evaluated as a structural property of the artifact graph, enabling early, machine-checkable detection of missing relations.
Padma Iyenghar, Christopher Zimmerman, C. Gregorio· International Conference on...· 0 citations
Abstract—Federal and regulated organizations continue to rely on document-centric Authorization to Operate (ATO) processes even as the NIST Risk Management Framework (RMF), continuous monitoring guidance, Zero Trust Architecture (ZTA), and continuous authorization initiatives require more continuous, evidence-driven risk management [1]-[3], [13], [15]. Manual System Security Plan (SSP) updates, spreadsheet-based Plan of Action and Milestones (POA&M) tracking, and disconnected assessment evidence create governance latency: the delay between operational security events and authorization-ready governance response. This paper presents OpenGRCRMF, a proposed open, vendor-neutral reference framework that models RMF lifecycle activities as workflow states, treats authorization artifacts as structured governance objects, and maps DevSecOps and Zero Trust telemetry into authorization-relevant evidence. Using Design Science Research, the study develops the OpenGRCRMF architecture, formalizes its data and risk model, and evaluates expected governance effects through a synthetic simulation of 1,500 findings across 180 assets and 320 controls [9]. OpenGRCRMF is evaluated as a reference framework rather than a production platform using a self-contained simulation specification and sensitivity analysis. In the simulation, the OpenGRCRMF-enabled workflow reduced modeled governance processing time by 36.8 percent, increased modeled evidence completeness by 41.5 percent, and increased modeled control-to-evidence traceability by 52.7 percent compared with a document-centric baseline. These results are modeled outcomes under stated assumptions, not production deployment proof. The paper contributes a governance object model, governance latency construct, reproducibility-oriented simulation design, threat-to-validity analysis, and education-oriented framework for teaching how operational telemetry becomes authorization evidence.
Anand Janjal· Journal of Cybersecurity Edu...· 0 citations
Cyber ranges are complex environments comprising many interacting components and stakeholders with different security concerns. The Service-Oriented Cyber Range (SOR) is no exception, particularly when it comes to training scenarios targeting critical infrastructure. Security concerns are translated into security requirements, the elicitation of which is usually difficult and time-consuming. This work examines how large language models can assist in eliciting security requirements for a service-oriented range and help produce a useful baseline for designers and developers. The approach follows a SEBoK-guided process in which security mission objectives and stakeholder needs were first identified and then provided as a prompt context along with architectural guidelines to five LLMs: GPT-5.2, Gemini 3.1 Pro, Grok 4.1, Sonar, and Kimi K2.5. The models generated 84 security requirements in total, which were consolidated into a comprehensive set of 27 requirements and then mapped to the architectural layers of the service-oriented range. The final set was evaluated by five cybersecurity experts against the criteria of necessity, clarity, completeness, feasibility, and testability, with an additional rejection option. The results showed a high acceptance rate, specifically for necessity with 98.5%, clarity with 87.4%, completeness with 85.2%, feasibility with 78.5%, and rejection with 0.7%. Testability was lower at 44.4%, indicating a slight lack of information on how these requirements could be tested. These findings show that LLMs can support early stages of the elicitation of security requirements, although human review is still needed, especially to improve or adjust certain aspects of the requirements.
Michail Takaronis, Athanasia Kollarou, G. Kavallieratos et al.· 0 citations
A hybrid framework that takes a BPMN process model and a security requirements document as input and automatically generates security annotations adhering to the SecBPMN2 specification is presented, providing a scalable foundation for security-by-design BPM.
Enterprise use of large language models creates architectural challenges involving control, traceability, state management, output conformance, and operational governance. This design-science paper introduces the Deterministic– Probabilistic Responsibility Allocation Framework, which distinguishes a model-first architecture from a hybrid code-first architecture by specifying where validation, routing, state management, contextual reasoning, output enforcement, recovery, and telemetry should reside. The framework contributes an explicit responsibility-assignment method, a boundary contract for each model invocation, and lifecycle control points that connect probabilistic generation to deterministic validation and escalation. It is examined through an application programming interface sandbox-generation illustration and a modelagnostic reference procedure covering prompt structure, OpenAPI parsing, schema validation, retry and fallback behavior, and output acceptance criteria. The analysis shows how the framework can structure fault isolation, observable execution paths, and enforceable output contracts without claiming measured improvements in reliability, scalability, latency, or cost. Controlled multi-model benchmarking remains necessary to quantify these properties.
Unknown authors· AI and Machine Learning Adva...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.