Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have also made them harder to defend against external adversaries. Defenses that restrict this arc...
Mayank Rathee, Alexander Stepanov, Shalin Madabhavi et al.· 0 citations
Software development follows an implementation-verification loop in which developers or agents iteratively revise an implementation until an evaluator, such as a test suite, accepts it. The evaluator checks the implementation against a set of requirements under a model of the deployment environment. Yet even a formal p...
Alexander Krentsel, Shubham Agarwal, M. Cemri et al.· 0 citations
It is argued that an agent system should maintain an explicit representation of how it fails, induced from its own behavior and reusable wherever failure feedback is needed, and AdaMAST builds this representation by converting a target system's traces into a compact, evidence-grounded failure taxonomy.
M. Cemri, Andrei Cojocaru, Melissa Z. Pan et al.· 0 citations
A three-level taxonomy inspired by autonomous driving that distinguishes degrees of autonomy along a roadmap from today’s AI-assisted development workflows to fully autonomous software development in which AI systems autonomously identify demands and design, implement, verify, and maintain software without human oversi...
Hao Wang, Ruijie Meng, Zhe Ye et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.