Skip to content
Review

Engineering Trustworthy Agentic AI for Critical Systems

Jul 2026 · arXiv.org · Vol abs/2607.18548 · 1 citation · 131 references
Computer Science Engineering

TL;DR

Synthesizing across four constraint-bound engineering domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.

Abstract

Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences. This survey addresses a gap in current literature by treating trustworthiness, whether agentic behavior can be verified, audited, and trusted under the constraints that engineering practice actually requires, as a first-class engineering property, rather than evaluating agentic AI by task capability alone. The study adopts a trustworthiness model organized around five cross-cutting dimensions: safety and constraint satisfaction; robustness and reliability; transparency and interpretability; accountability and auditability; and privacy and security. This is mapped onto an agentic assurance workflow spanning perception through audit. Building on this foundation, agentic systems architectures, threats, concrete trust mechanisms, and quantitative metrics are surveyed for direct application in agentic systems development and evaluation. These principles are then examined across four constraint-bound engineering domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks, identifying recurring design patterns, shared failure modes, and domain-specific gaps. Synthesizing across those domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.

View source

Similar papers

Preprint Aug 2026

Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities

It is argued that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms.

Timothy Merritt, Alejandro Jarabo-Peñas, Juan Bravo-Arrabal et al. · 0 citations
Review

Trustworthy AI for Safety-Critical Perception and Decision Systems

It is asserted that trustworthiness is a systems property, not a single algorithmic feature, and that realising it requires coordinated advances in explainability, uncertainty quantification, human–AI interaction design, failure detection, and domain-specific data governance.

Shruti Kshirsagar · 0 citations
Conference Jul 2026

A Survey on Autonomous Compliance Enforcement using Agentic AI

Infrastructure compliance enforcement is increasingly considered for agentic AI systems that can plan, act and self-correct over many steps. A structured examination of the literature reveals several significant gaps. Multi-agent compliance pipelines have been architecturally proposed in several works but none report a functioning prototype with measurable compliance outcomes. Theoretical discussions extensively cover the safety hazards arising from the granting of autonomous control to an LLM agent over remediation of infrastructure. However, there are no documented real-world cases of an LLM agent outputting an operationally dangerous output in a compliance setting. The concept of utilising cross-run memory for compliance agents has been recognised but no lightweight implementation has yet been demonstrated to change the behaviour of agents. This paper surveys the field along seven dimensions: agentic architectures, compliance automation, LLM output safety, anomaly detection, multi-agent coordination, statefulness, and cloud-native deployment drawing on 46 representative works. Six specific gaps are identified through structured analysis. A hybrid architecture is then proposed that integrates the Isolation Forest anomaly detection with a four-agent LLM pipeline consisting of an Analyser that interprets system state and prior run history, a Planner that generates remediation strategies, a Verifier that applies LLM safety constraints, and an Explainer that produces human-readable audit reports. The architecture further incorporates deterministic value-level validation and persistent SQLite-based cross-run memory. In controlled experiments, the system improved compliance scores from 60% to 100%. Notably, a concrete instance of LLM overreach was observed during testing: the Planner agent generated a remediation plan that would have locked out SSH access by closing all network ports, a failure mode not previously reported in empirical literature. The paper concludes with a feature-by-feature comparison across twelve prominent works, a discussion of open challenges, and a proposed hybrid cloud extension.

Saleha Soudagar, V. Rajpurohit, Arati Shahapurkar et al. · 0 citations
Preprint Aug 2026

Securing Agentic AI: From Per-Action Checks to Trajectory Assurance

Charting these challenges provides a roadmap toward trustworthy autonomous agent deployment: security must become a verifiable property of the architectures, protocols, and runtimes that govern agent behavior, rather than an optional layer of guidance.

Alireza Lotfi, Subangkar Karmaker Shanto, Imtiaz Karim et al. · 1 citation
Review Jul 2026

Beyond Component Testing: Validating Agentic AI Systems

This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems, and concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures.

Fabio Orazio Mirto, L. D’Agati, Giuseppe Tricomi et al. · 0 citations
Open access 2026

Agentic AI: Architectures, Types, Capabilities, Mathematical Equations and Governance in the Era of Autonomous Intelligence

A comprehensive framework for the design, evaluation, and responsible deployment of Agentic AI is proposed, emphasizing safety, explainability, human-in-the-loop supervision, and ethical compliance and aims to maximize the benefits of Agentic AI while minimizing potential risks.

Nitin S. Shrirao, Dnyaneshwar S. Jadhav, Sarita B. Patil · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.