Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs. We formalize this phenomenon as \textit{false memory}, wh...
Quan M. Tran, Zhuo-Jing Huang, Zhe-Min Fang et al.· 0 citations
It is shown that preference alignment preserves the human response distribution only under a restrictive condition, and no consistent evidence that real human preferences satisfy it, and human-likeness is established as an explicit dimension of alignment rather than something assumed to follow from preference alignment...
Su-Qin Yuan, Runqi Lin, Mu-Yang Li et al.· 0 citations
This work proposes Targeted Anti-obfuscation with Mechanistic Enforcement (TAME), which uses Sparse Autoencoders (SAEs) to combine behavioral feedback with targeted suppression of template-associated activations during RL, providing a path from behavioral monitoring to representation-level oversight for more auditable...
Xu-Tao Mao, Jianing Zhu, Jin-Man Zhao et al.· 0 citations
This work establishes a new paradigm for generated image detection by recasting the detection task as a problem of machine unlearning, and introduces two detection methods: data-free detection, which prunes model parameters to induce unlearning without data access, and data-driven detection, which optimizes LVMs to unl...
Jun Nie, Yonggang Zhang, Tongliang Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.