Skip to content
Review Open access

Agentic AI Safety: A Structured Review of Open Problems and Their Regulatory Anchoring

Aug 2026 · Applied Informatics · Vol 7, pp. 298 · 0 citations · 33 references

TL;DR

A structured review and taxonomy of open scientific problems in agentic AI safety, mapped explicitly onto the EU AI Act and the NIST AI Risk Management Framework is presented, arguing that progress on inner alignment, interpretability for deceptive-alignment detection, and multi-agent safety would most directly reduce compliance uncertainty.

Abstract

The shift from passive predictive models to autonomous agents capable of tool use and multi-step planning moves the AI safety landscape from prediction error to control failure: small misjudgements become irreversible actions, and risks compound across long horizons and populations of interacting systems. We present a structured review and taxonomy of open scientific problems in agentic AI safety, mapped explicitly onto the EU AI Act and the NIST AI Risk Management Framework. The corpus follows a PRISMA-ScR scoping review, assembled through anchor-based citation chaining and curated reading lists across arXiv, the major machine-learning conferences, and selected security and fairness venues, with a primary March 2026 search cut-off (extended to May 2026 during revision for a small number of high-relevance governance and agentic-safety sources), explicit eligibility criteria, and an analytical distinction between open scientific problems and deployment risks. The taxonomy identifies eight problem families spanning reinforcement-learning policies and language-model planners: goal specification, inner alignment, safe learning and robustness, scalable oversight, interpretability, tool-use security, multi-agent safety, and evaluation and assurance. Mapping these onto the two frameworks shows close alignment for some families and notable absences for others, with multi-agent safety surfacing as a regulatory gap. We add a per-family research roadmap with concrete milestones and a practitioner-facing deployment-posture triage, arguing that progress on inner alignment, interpretability for deceptive-alignment detection, and multi-agent safety would most directly reduce compliance uncertainty.

Read PDF

Similar papers

Preprint Aug 2026

Multi-Agent AI Safety as an Institutional Design Problem

This is the first paper from POLIS, an ongoing research programme studying algorithmic institutions for multi-agent systems, and asks which parts of an AI institution produce safety and how they do it.

X. Abdullah · 1 citation
#reinforcement learning Review Open access Sep 2026

A survey of world models for physical AI with uncertainty representation and control

Physical AI systems must reason about real-world dynamics in order to perceive, predict, and act safely under partial observability and uncertainty. World models–learned predictive representations of environment dynamics and action consequences–have emerged as a unifying framework for integrating perception, prediction, planning, and control in embodied agents. This survey provides a comprehensive and technically grounded review of learning-based world models for Physical AI, with particular emphasis on closed-loop decision-making. We organize existing approaches along six compositional design dimensions: state abstraction, temporal dynamics, uncertainty source and treatment, structural prior, observation modality, and decision coupling. Beyond this design-oriented taxonomy, we analyze how world models interact with optimization–highlighting compounding error, planner exploitation, rollout horizon management, and uncertainty calibration as central design tensions. We further examine evaluation methodologies, benchmark ecosystems, and sim-to-real transfer challenges, and synthesize open problems in long-horizon consistency, physical constraint enforcement, data efficiency, and safety. By clarifying recurring trade-offs across robotics and model-based reinforcement learning, this survey outlines principled directions for building reliable and scalable Physical AI systems.

Sven Kirchner, Nils Purschke, Alois Knoll · 0 citations
Review Jul 2026

Beyond Component Testing: Validating Agentic AI Systems

This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems, and concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures.

Fabio Orazio Mirto, L. D’Agati, Giuseppe Tricomi et al. · 0 citations
Review Open access Aug 2026

From Language Models to Agentic AI: A Survey of Autonomous, Action-Enabled, and Collaborative LLM Agents

A unified, taxonomy-driven, and deployment-oriented survey of agentic AI systems, synthesizing recent advances through a modular reference architecture and a four-dimensional taxonomy that characterizes agents along the axes of autonomy, tool use, collaboration, and safety–governance is presented.

Sparsh Bajoria, Shreyanshu Ranjan, Adhitya M et al. · 0 citations
Preprint Jul 2026

Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents

A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al.,"Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive discrete tool universes (Majumdar,"Sparse Agentic Control", 2026), and regulatory-control decompositions of multi-agent systems into auditable feedback loops (Nogueira and Skogestad, 2026). We do not claim to introduce control theory to LLM agents -- that ship has sailed. Our narrower claim is about what the controlled variable is. Prior work controls tool selection, inter-agent message routing, or the agent's raw action stream. We instead treat context assembly itself -- which prompt template, which few-shot demonstrations, how much retrieved context, how many planning/verification passes -- as the controlled variable, learned online by a contextual bandit or REINFORCE policy sitting outside a frozen model. This paper develops the formal decomposition (inner frozen policy $\pi_\theta$, outer context policy $\pi_\phi$), gives a stability argument for the online controller in the sense used by Zhang et al. (2026) (non-decreasing expected reward under bounded policy change), and reports an uncertainty-calibration analysis of the controller's own confidence against realized task outcomes. The applied counterpart to this paper instantiates the same controller across three domains and two model providers and releases the dataset, trajectory logs, and a deployment recipe; here we focus on the formal framing and the stability/uncertainty evidence a control-theoretic claim requires.

Debjyoti Paul · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.