TULIP: Targeted LLM Unlearning at Layers Identified Per-Input
A hijacking experiment is designed that grafts hidden states of the target model into an oracle trained only on the retain set, which serves as a plug-and-play component that further improves existing methods.
Yejin Kim, William F. Shen, Seokwon Jung et al.
· 0 citations