This position paper outlines a common abstraction layer that can substantially lower the effort for designing agents with constrained creativity, and demonstrates early promise of this paradigm for two SysOps use cases: Root Cause Analysis and Network Configuration Generation.
This work formalizes minimum-sufficient execution and the Agent Cognitive Redundancy Ratio (ACRR), and proposes E3 (Estimate, Execute, Expand): the agent estimates an initial operating point, executes a minimum viable path, and expands scope only when verification fails.
This is the first paper from POLIS, an ongoing research programme studying algorithmic institutions for multi-agent systems, and asks which parts of an AI institution produce safety and how they do it.
A prototype of a Plan Mode for spreadsheet programming is built and evaluated against a non-planning baseline and it is found that using Plan Mode led to a reduction in refinement and a better perception of the tool across dimensions of creativity support and human-machine collaboration.
Aayush Kumar, Avik Dutta, Sumit Gulwani et al.· 0 citations
This paper synthesizes 27 benchmark, taxonomy, and audit papers (2023-2026), spanning 19 distinct benchmarks, into a cross-cutting taxonomy of agent limitations, the first synthesis that integrates evidence across tool use, planning, long-horizon reasoning, multi-agent coordination, safety, and measurement validity into a single, unified taxonomy of LLM agent limitations.
Wael S. Albayaydh, Rui Zhao, Ivan Flechais· 1 citation
This work presents the first systematic evaluation framework for agentic abstention, and identifies failure modes such as post-hoc abstention, in which agents execute irreversible actions before recognizing abstention triggers.
Xun Liu, Y. Zhang, Vira Kasprova et al.· arXiv.org· 2 citations
A unified, taxonomy-driven, and deployment-oriented survey of agentic AI systems, synthesizing recent advances through a modular reference architecture and a four-dimensional taxonomy that characterizes agents along the axes of autonomy, tool use, collaboration, and safety–governance is presented.
Sparsh Bajoria, Shreyanshu Ranjan, Adhitya M et al.· Cognitive Computation· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.