Preprint
Aug 2026
ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
It is argued that effective unlearning must operate at the level of concepts, ensuring complete removal of unsafe applications while maintaining their correct and useful usage, thereby achieving conceptually meaningful and complete unlearning.
Sahil Kale, Ian G. Harris
· 0 citations