Skip to content
Preprint

Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

It is argued that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms.

Abstract

Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real-world adoption remains constrained not only by autonomy performance, but by the difficulty of integrating agentic behavior into professional work: operators must understand, trust, and govern automation under uncertainty, time pressure, and accountability. This position paper synthesizes the ambitions and lessons from two ongoing efforts: NAMUR, which explores LLM-supported robot control in SAR and firefighting contexts, and PERSIST, which explores persistent drone operations for monitoring and security at critical infrastructure sites. We argue that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms. We outline a human-centered, participatory, and iterative research approach aimed at uncovering stakeholder needs, shaping agent capabilities through successive prototypes, and producing transferable proof-of-concept systems and evaluation strategies for other safety-critical contexts.

View source

Similar papers

Review Jul 2026

Engineering Trustworthy Agentic AI for Critical Systems

Synthesizing across four constraint-bound engineering domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.

Omar Al-Refai, Ibrahim Shahbaz, A. Husseinat et al. · 1 citation
Review

Trustworthy AI for Safety-Critical Perception and Decision Systems

It is asserted that trustworthiness is a systems property, not a single algorithmic feature, and that realising it requires coordinated advances in explainability, uncertainty quantification, human–AI interaction design, failure detection, and domain-specific data governance.

Shruti Kshirsagar · 0 citations
Preprint Aug 2026

Securing Agentic AI: From Per-Action Checks to Trajectory Assurance

Charting these challenges provides a roadmap toward trustworthy autonomous agent deployment: security must become a verifiable property of the architectures, protocols, and runtimes that govern agent behavior, rather than an optional layer of guidance.

Alireza Lotfi, Subangkar Karmaker Shanto, Imtiaz Karim et al. · 1 citation
Jul 2026

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints and identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability.

B. Alsinglawi, Wei-Zheng Wang, Jun-Yi Wu et al. · 0 citations
Review Open access Aug 2026

Embodied Intelligence for Safer Power-System Field Operations: A Critical Review of Technologies, Applications, and Challenges

Modern power grids require safer and more reliable field operations, yet conventional robots often face limitations in unstructured environments because of rigid pre-programming and weak perception–action coupling. This review examines Embodied Intelligence (EI) as an emerging direction for enhancing power-system field operations. We first evaluate the environmental adaptability of morphological carriers, including quadrupeds, humanoids, and unmanned aerial vehicles, and then define the perception–cognition–execution closed-loop architecture used in this review. Three application domains are then examined. Intelligent inspection focuses on active perception and potential open-vocabulary object detection. Live-line maintenance emphasizes Sim-to-Real methods and shared autonomy, while disaster-response applications involve heterogeneous air–ground robotic coordination. The review also discusses the potential for EI to reduce human exposure to hazardous tasks and influence labor structures, while a regional text-based proxy illustrates differences in policy attention to digital infrastructure. Finally, we analyze major constraints, including hardware endurance under extreme climates, edge-computing latency, foundation-model uncertainty and hallucination, cybersecurity, and safety certification. Overall, EI should not be interpreted as a mature replacement for current utility practice; it is a developing technological direction whose safe deployment will require field validation, standardized evaluation, cybersecurity assurance, and continued human supervisory authority.

Yuxin Wen, Peixiao Fan, Zhiyu Mao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.