Aug 2026· Applied Informatics· Vol 7, pp. 298· 0 citations· 33 references
TL;DR
A structured review and taxonomy of open scientific problems in agentic AI safety, mapped explicitly onto the EU AI Act and the NIST AI Risk Management Framework is presented, arguing that progress on inner alignment, interpretability for deceptive-alignment detection, and multi-agent safety would most directly reduce compliance uncertainty.
Abstract
The shift from passive predictive models to autonomous agents capable of tool use and multi-step planning moves the AI safety landscape from prediction error to control failure: small misjudgements become irreversible actions, and risks compound across long horizons and populations of interacting systems. We present a structured review and taxonomy of open scientific problems in agentic AI safety, mapped explicitly onto the EU AI Act and the NIST AI Risk Management Framework. The corpus follows a PRISMA-ScR scoping review, assembled through anchor-based citation chaining and curated reading lists across arXiv, the major machine-learning conferences, and selected security and fairness venues, with a primary March 2026 search cut-off (extended to May 2026 during revision for a small number of high-relevance governance and agentic-safety sources), explicit eligibility criteria, and an analytical distinction between open scientific problems and deployment risks. The taxonomy identifies eight problem families spanning reinforcement-learning policies and language-model planners: goal specification, inner alignment, safe learning and robustness, scalable oversight, interpretability, tool-use security, multi-agent safety, and evaluation and assurance. Mapping these onto the two frameworks shows close alignment for some families and notable absences for others, with multi-agent safety surfacing as a regulatory gap. We add a per-family research roadmap with concrete milestones and a practitioner-facing deployment-posture triage, arguing that progress on inner alignment, interpretability for deceptive-alignment detection, and multi-agent safety would most directly reduce compliance uncertainty.
This is the first paper from POLIS, an ongoing research programme studying algorithmic institutions for multi-agent systems, and asks which parts of an AI institution produce safety and how they do it.
Physical AI systems must reason about real-world dynamics in order to perceive, predict, and act safely under partial observability and uncertainty. World models–learned predictive representations of environment dynamics and action consequences–have emerged as a unifying framework for integrating perception, prediction, planning, and control in embodied agents. This survey provides a comprehensive and technically grounded review of learning-based world models for Physical AI, with particular emphasis on closed-loop decision-making. We organize existing approaches along six compositional design dimensions: state abstraction, temporal dynamics, uncertainty source and treatment, structural prior, observation modality, and decision coupling. Beyond this design-oriented taxonomy, we analyze how world models interact with optimization–highlighting compounding error, planner exploitation, rollout horizon management, and uncertainty calibration as central design tensions. We further examine evaluation methodologies, benchmark ecosystems, and sim-to-real transfer challenges, and synthesize open problems in long-horizon consistency, physical constraint enforcement, data efficiency, and safety. By clarifying recurring trade-offs across robotics and model-based reinforcement learning, this survey outlines principled directions for building reliable and scalable Physical AI systems.
Sven Kirchner, Nils Purschke, Alois Knoll· Discover Artificial Intellig...· 0 citations
This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems, and concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures.
Fabio Orazio Mirto, L. D’Agati, Giuseppe Tricomi et al.· arXiv.org· 0 citations
A unified, taxonomy-driven, and deployment-oriented survey of agentic AI systems, synthesizing recent advances through a modular reference architecture and a four-dimensional taxonomy that characterizes agents along the axes of autonomy, tool use, collaboration, and safety–governance is presented.
Sparsh Bajoria, Shreyanshu Ranjan, Adhitya M et al.· Cognitive Computation· 0 citations
A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al.,"Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive discrete tool universes (Majumdar,"Sparse Agentic Control", 2026), and regulatory-control decompositions of multi-agent systems into auditable feedback loops (Nogueira and Skogestad, 2026). We do not claim to introduce control theory to LLM agents -- that ship has sailed. Our narrower claim is about what the controlled variable is. Prior work controls tool selection, inter-agent message routing, or the agent's raw action stream. We instead treat context assembly itself -- which prompt template, which few-shot demonstrations, how much retrieved context, how many planning/verification passes -- as the controlled variable, learned online by a contextual bandit or REINFORCE policy sitting outside a frozen model. This paper develops the formal decomposition (inner frozen policy $\pi_\theta$, outer context policy $\pi_\phi$), gives a stability argument for the online controller in the sense used by Zhang et al. (2026) (non-decreasing expected reward under bounded policy change), and reports an uncertainty-calibration analysis of the controller's own confidence against realized task outcomes. The applied counterpart to this paper instantiates the same controller across three domains and two model providers and releases the dataset, trajectory logs, and a deployment recipe; here we focus on the formal framing and the stability/uncertainty evidence a control-theoretic claim requires.
Debjyoti Paul· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.