Recursive self-improvement (RSI) relies on evaluation feedback to assess progress and guide further research, yet repeatedly running complex benchmarks is costly and slows iteration. Human experts reduce this cost by selecting benchmark subsets or designing compact suites. We ask whether AI agents can automate this des...
Yao Zhang, Tian-Yi Xu, Yu-Jie Zhao et al.· 0 citations
AsynCodeBench is introduced, a dependency-centric benchmark for asynchronous multi-agent software engineering that represents each task with an explicit dependency graph and executable Dependency Checkers and proposes two complementary measures: Asynchronous Dependency Pass Rate (ADPR), which measures how many cross-ag...
Kai-Tuo Zhang, Zhen Xiong, Zhi-Meng Jiang et al.· 0 citations
Extreme weather and volatile wholesale electricity markets expose residential consumers to catastrophic financial risks, yet demand response at the distribution level remains an underutilized tool for grid flexibility and energy affordability. While a demand-response program can shield consumers by issuing financial cr...
Jose E. Aguilar Escamilla, Lingdong Zhou, Xiangqi Zhu et al.· arXiv.org· 1 citation
This paper investigates LVLMs’ behavior in VAD from a visual-textual co-occurrence perspective, and proposes VAD-DPO, a direct preference optimization method supervised with counter-example pairs that enhances both anomaly detection and reasoning performance, particularly in scene-dependent scenarios.
Menghao Zhang, Huazheng Wang, Pengfei Ren et al.· Neural Information Processin...· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.