Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users'expressed views. In reality, sycophancy rarely happens in a single exchange; it m...
Sidharth Pulipaka, R. Binkytė, Ivaxi Sheth et al.· 0 citations
The framework and EvalAwareBench provide the tools to measure, attribute, and mitigate evaluation awareness, building the foundation for future solutions to ground evaluation awareness in social psychology.
Changling Li, Terry Jingchen Zhang, Jie M. Zhang et al.· arXiv.org· 1 citation
Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent about verbalizing evaluation awareness while avoiding to supe...
Usman Anwar, Sahar Abdelnabi, David Krueger· 0 citations
This work introduces SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment and shows that qualitatively different safety behaviors emer...
Saswat Das, Parvati Viswanathan, Daniel Donnelly et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.