It is argued that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms.
Abstract
Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real-world adoption remains constrained not only by autonomy performance, but by the difficulty of integrating agentic behavior into professional work: operators must understand, trust, and govern automation under uncertainty, time pressure, and accountability. This position paper synthesizes the ambitions and lessons from two ongoing efforts: NAMUR, which explores LLM-supported robot control in SAR and firefighting contexts, and PERSIST, which explores persistent drone operations for monitoring and security at critical infrastructure sites. We argue that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms. We outline a human-centered, participatory, and iterative research approach aimed at uncovering stakeholder needs, shaping agent capabilities through successive prototypes, and producing transferable proof-of-concept systems and evaluation strategies for other safety-critical contexts.
Synthesizing across four constraint-bound engineering domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.
Omar Al-Refai, Ibrahim Shahbaz, A. Husseinat et al.· arXiv.org· 1 citation
It is asserted that trustworthiness is a systems property, not a single algorithmic feature, and that realising it requires coordinated advances in explainability, uncertainty quantification, human–AI interaction design, failure detection, and domain-specific data governance.
A mixed-criticality architectural framework that applies SAE ARP4754B methods to swarm reconfiguration, enabling intelligent swarm behaviors without compromising flight-critical isolation is proposed.
Luiz Giacomossi, Zafer Yigit, Marwan Shakarna et al.· 0 citations
Charting these challenges provides a roadmap toward trustworthy autonomous agent deployment: security must become a verifiable property of the architectures, protocols, and runtimes that govern agent behavior, rather than an optional layer of guidance.
MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints and identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability.
B. Alsinglawi, Wei-Zheng Wang, Jun-Yi Wu et al.· arXiv.org· 0 citations
Modern power grids require safer and more reliable field operations, yet conventional robots often face limitations in unstructured environments because of rigid pre-programming and weak perception–action coupling. This review examines Embodied Intelligence (EI) as an emerging direction for enhancing power-system field operations. We first evaluate the environmental adaptability of morphological carriers, including quadrupeds, humanoids, and unmanned aerial vehicles, and then define the perception–cognition–execution closed-loop architecture used in this review. Three application domains are then examined. Intelligent inspection focuses on active perception and potential open-vocabulary object detection. Live-line maintenance emphasizes Sim-to-Real methods and shared autonomy, while disaster-response applications involve heterogeneous air–ground robotic coordination. The review also discusses the potential for EI to reduce human exposure to hazardous tasks and influence labor structures, while a regional text-based proxy illustrates differences in policy attention to digital infrastructure. Finally, we analyze major constraints, including hardware endurance under extreme climates, edge-computing latency, foundation-model uncertainty and hallucination, cybersecurity, and safety certification. Overall, EI should not be interpreted as a mature replacement for current utility practice; it is a developing technological direction whose safe deployment will require field validation, standardized evaluation, cybersecurity assurance, and continued human supervisory authority.
Yuxin Wen, Peixiao Fan, Zhiyu Mao et al.· Applied Informatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.