Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to go...
Hong-Yi Du, Tian-Yi Zhang, Wei-Jia Zhang et al.· 0 citations
AI systems can strengthen democracy by supporting deliberation at scale by addressing cognitive, social, platform-design, and market-driven frictions, while preserving human agency. Unlike proposals such as liquid democracy that restructure representation through vote delegation, in this position paper, we argue that A...
José Ramón Enríquez, Jia-Xin Pei, A. Pentland· 0 citations
A user-centric framework for systematically auditing system prompts in AI systems, AISPA is introduced, a user-centric framework for systematically auditing system prompts in AI systems that examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users.
Xiangning Lin, Shenzhe Zhu, Shu Yang et al.· arXiv.org· 0 citations
ChainSWE is introduced, the first benchmark for evaluating agents on sequential, dependent bug fixes within a shared codebase, and reveals a consistent performance drop by up to 70% as the chain length increases.
Qirui Jin, Lingching Tung, Kenan Li et al.· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.