One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models
This work presents the first systematic, bidirectional study of cross-modal unlearning transfer across three VLM architectures: LLaVA-1.5 (MLP projection), InstructBLIP (Q-Former), and IDEFICS (gated cross-attention), finding that unlearning transfers across modalities, but the transfer is asymmetric and incomplete.