Jul 2026· International Scientific Journal of Engineering and Management· Vol 05, pp. 1-11· 0 citations
TL;DR
The Prompt-Level AI Risk Governance Framework (PLAIRGF), a four-phase architectural model covering Prevention, Detection, Response, and Continuous Improvement, is introduced, with a structured alignment between PLAIRGF controls and international compliance standards, offering organizations a practical, audit-ready foundation for deploying LLM-based systems securely.
Abstract
Generative Artificial Intelligence systems—particularly those built on Large Language Models (LLMs)—have become central to modern enterprise computing, yet they carry with them a class of vulnerabilities that traditional cybersecurity models were never designed to address. Decoder-only transformer architectures process system instructions and untrusted user inputs as a single undifferentiated sequence of tokens, which makes them susceptible to direct and indirect prompt injections, jailbreaking attacks, and retrieval poisoning. Existing governance frameworks—including the NIST AI Risk Management Framework, the EU AI Act, and ISO/IEC 42001—offer valuable compliance guidance at the organizational level, but they stop short of providing the kind of execution-level security blueprints that engineering teams actually need.
This paper introduces the Prompt-Level AI Risk Governance Framework (PLAIRGF), a four-phase architectural model covering Prevention, Detection, Response, and Continuous Improvement. Rather than depending on simulated metrics to evaluate the framework, we validate PLAIRGF through mathematical formalization of key risk indicators—including Attack Success Rate (ASR), Risk Severity Score (R), the Governance Readiness Index (GRI), and a newly proposed Human-in-the-Loop Alignment Coefficient (η)—combined with rigorous defensive capability mapping and step-by-step operational trace walkthroughs of representative adversarial attack scenarios. The paper concludes with a structured alignment between PLAIRGF controls and international compliance standards, offering organizations a practical, audit-ready foundation for deploying LLM-based systems securely.
Keywords: Generative AI Security, Large Language Models, Prompt Injection, AI Governance, Retrieval-Augmented Generation, Human-in-the-Loop, PLAIRGF
Artificial intelligence has introduced new capabilities for enterprise quality engineering, yet the absence of structured governance frameworks limits its practical adoption in regulated software environments. This paper presents the Probabilistic Augmentation and Governance Model (PAGM), a three-tier framework that formally allocates responsibility between AI agents and human testers across autonomous execution, confidence-gated escalation, and human-led verification. PAGM draws on a structured review of literature spanning AI quality assurance, autonomous system testing, UI-resilient automation, explainable AI, and defect prediction to identify governance gaps in existing approaches. The model addresses eight documented AI testing challenges, including interpretability, absence of formal specifications, oracle determination, and dynamic operational environments. Unlike prior approaches that treat AI test execution as a performance optimization problem, PAGM positions governance, traceability, and bounded autonomy as first-class design requirements. Its application is demonstrated through a regulated property insurance release pipeline in which PAGM enables a full quarterly regression cycle within a five-day service level agreement while preserving mandatory human accountability for premium calculation and claims adjudication workflows. Quantitative evidence from the literature supports the model's constituent mechanisms, including a 95% UI change handling rate for contextual recognition methods compared to 40 to 80 percent achieved by conventional tools and AI-based classifiers outperforming classical approaches by 25 to 50 percent across software quality metrics. PAGM provides a governance-ready foundation for organizations seeking to deploy AI-assisted testing at enterprise scale within regulated or auditable environments.
Rejenish Kiran· International journal of com...· 0 citations
This paper proposes a structured protocol to automate AI risk mitigation through a taxonomy-driven analysis of open-source LLM evaluation and security tools, and presents a taxonomy-driven framework applicable to open-source and proprietary solutions.
Afreen Alam, Evgenija Popchanovska, Ana Gjorgjevikj et al.· 0 citations
AITD-MAP is introduced, an integrated framework that connects the AITD taxonomy, quality and risk impacts, and mitigation strategies into a unified structure for risk-aware AI engineering, and aims to assist AI software engineers in making AI safety and security technical debts visible, understanding their root causes, and mitigating their presence.
Muhammad Tukur, H. Adeyemo, Tao-An Chen et al.· arXiv.org· 1 citation
— While Generative AI (GenAI) is set to make a mark in financial services, the use of this technology in banking risk and compliance throws up the critical questions of interpretability, auditability and trustworthiness in the highly regulated sector. Large language models (LLMs) such as GPT-4 and Gemini Pro are not specialized for banking and do not comply with regulatory rules and standards. It proposes a novel GenAI based banking risk and compliance framework, namely a purpose-built GRACE (Generative Risk and Compliance Evaluation Framework), which integrates Explainable AI (XAI), a cryptographically secured Immutable Audit Trail, a Human-in-the-Loop (HITL) oversight layer and dedicated compliance alignment layers for Basel III, IFRS 9, AML/CFT and GDPR. Beyond the architecture, we suggest a method for evaluating GRACE and representative comparator systems (GPT-4, Gemini Pro, and BloombergGPT) by six criteria: interpretability, compliance readiness, trustworthiness, regulatory auditability, bias and fairness, and domain specificity. A proposed evaluation protocol is presented to assess the adaptability of the systems through an illustrative architectural capability assessment. The present assessment is theoretical rather than empirical and is intended to demonstrate the potential value and discriminative capability of the proposed methodology. Because no expert panel evaluation or empirical dataset has yet been established, the illustrative values should not be interpreted as measured performance scores. Whereas purpose-built systems are more architecturally flexible, domain tuned GenAI systems are less flexible. We propose a prototype of how this work could be implemented in the real world. This encompasses a Banking Compliance Evaluation Suite, scoring protocol devised by a panel of experts, and a statistical
Anamika Singh· Iconic research and engineer...· 0 citations
It is concluded that validation and governance of grounded and agentic AI must be treated as a first-class enterprise reliability engineering discipline — auditable, thresholddriven, and embedded across the inference lifecycle — rather than as an extension of conventional model evaluation.
Suresh Babu Narra· International Journal of Int...· 0 citations
The rapid development of Large Language Models (LLMs) and Artificial Intelligent (AI) powered autonomous agents has fundamentally changed the existing forms of software governance. In spite of the rigorous standards of transparency and account ability required according to the international frameworks such as the European Union's AI Act, there is a considerable gap between theory and reality. The present study discusses the inherent drawbacks of currently utilized platforms for LLM evaluation, machine learning workflow, and application performance monitoring in general. It has been shown that current disjointed solutions fail to protect unbound state space agentic architecture from serious threats such as alignment drift, SaaS security concerns, and unauthorized deployment of shadow AI systems. Moreover, a solution is proposed for overcoming the discussed challenges in form of a coherent multi-level AI governance stack Traccia built on the top of OpenTelemetry infrastructure platform. Traccia resolves the last mile for AI Alignment by adding the telemetry data, passive semantic guardrail assessment, and execution lineage into a hashed trace ledger. Traccia automatically creates compliance evidence packages by appending tamper-resistant fingerprints and SHA-256 content hash, that map to regulatory requirements (Articles 12, 14, 19, 26(6), and 50 of the EU AI Act) without invading any data privacy. By performing this evaluation in a methodical manner, a solid machine-readable base has been created for enterprise-wide management of autonomous AI systems.