Assessing cybersecurity vulnerability awareness in coding agents requires evaluations that reveal capability gaps and remain informative as models evolve. Static benchmarks offer fixed coverage and difficulty, while scarce vulnerable repositories and costly expert authoring limit their renewal at scale. We introduce Se...
Xiao-Nan Luo, Yue Huang, Ke-Han Guo et al.· 0 citations
How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.
Yue Huang, Zhangchen Xu, Yu-Chen Ma et al.· 1 citation
The Graph Theory Agent (GTA), which pairs a preference-trained representation selector with plan-and-decompose scaffolding around a frozen executor LLM, is proposed, which lifts Phi-4 from 53.5% to 69.1% on the benchmark's easy split and from 33.0% to 41.5% on its hard split.
Zi-Xiang Xu, Yan-Bo Wang, Chenxi Wang et al.· 2 citations· ⚡1
Experiments on RWKU and MUSE across different LLM architectures show that ATU achieves a better balance between target forgetting and retained utility, making unlearning more robust under tool-augmented agent deployment.
Baicheng Chen, Zhe-Yuan Liu, Jingyu Zhang et al.· 0 citations
This paper examines how skills may help address bottlenecks of current agents and how they may expand agent capabilities through reusable domain procedures loaded at inference time and outlines open questions in skill construction, composition, evaluation, portability, governance, and security.
Hanwen Xing, Haomin Zhuang, Xuandong Zhao et al.· Proceedings of the 32nd ACM...· 7 citations· ⚡1
SocialMaze is introduced, a benchmark that organizes six tasks across social deduction games, daily-life interactions, and digital community platforms along three descriptive design axes: deep reasoning, dynamic interaction, and information uncertainty.
Zi-Xiang Xu, Yan-Bo Wang, Yue Huang et al.· 1 citation
DOG-DPO is proposed, a training-free data selection framework that treats preference pairs as structured geometric signals and recovers most of the safety gains of full-data training while remaining entirely teacher-free, training-free, and substantially faster than representative selection baselines.
Yi Nian, Tiankai Yang, Yudi Zhang et al.· arXiv.org· 1 citation
As Large Language Model (LLM) agents have demonstrated broad competence, but they still struggle in specialized, real-world workflows. Existing approaches such as RAG, fine-tuning and tool integration improve knowledge access, model adaptation, and external functionality, yet they do not fully address a central gap: th...
Hanwen Xing, Haomin Zhuang, Xuandong Zhao et al.· Proceedings of the 32nd ACM...· 6 citations· ⚡1
A knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced, which reduces the confound between lying and not knowing and enables more rigorous auditing and...
Zhe-Yuan Liu, Wei-Liang Zhao, Xiangchi Yuan et al.· 1 citation
KITE (Knowledge-boundary Instruction Tuning via Exploration), a two-stage framework that combines failure-guided data generation with boundary-aware uncertainty curation, is proposed, showing that KITE yields more stable improvement than strong synthetic-data baselines.
MemoHarness is introduced, an adaptive harness optimization framework that learns from its own executions and improves over the fixed harnesses it is compared against and shows selective transfer to unseen suites and base models.
Yue Huang, Wenjie Wang, Han Bao et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.