This work proposes ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state and introduces three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery.
Abstract
Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where stronger forgetting of target knowledge can degrade model utility, and unlearned knowledge may reappear under post-unlearning fine-tuning or prompt attacks. We propose ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state. ARIA uses sparse autoencoder (SAE) latents to train a lightweight linear detector, then applies an interpretable intervention on triggered states with negligible test-time overhead. Empirical evaluations on TOFU, R-TOFU, and WMDP show that ARIA improves the forget-retain trade-off over weight-based baselines across both a thinking model (DeepSeek-R1-Distilled-Qwen-1.5B) and an instruction model (Gemma-3-1B-it), e.g., reducing WMDP-cyber forget-set accuracy significantly while keeping MMLU within 1% of the pre-unlearning model. We further introduce three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery, and find that ARIA remains robust under all three, with forgetting changing by less than 1% under attack. A feature-level case study leveraging the interpretability of ARIA suggests that some retain degradation may reflect response styles underlying the unlearning data rather than leakage of the targeted knowledge itself, highlighting a potential source of bias in unlearning task construction.
This work proposes Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer-level unlearning framework that selects transformer layers using a forget-to-retain significance score and provides an empirical path toward quantization-resilient unlearning.
Ravi Ranjan, O. Kotevska, Agoritsa Polyzou· 1 citation
This work introduces SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features and shows theoretically that contrastive training promotes target-selective features and that the selection score controls expected background knowledge perturbation.
Approximate machine unlearning seeks to remove the influence of a forget set from a trained model without full retraining. Existing gradient-based methods require data-dependent hyperparameter search, struggle when forget and retain knowledge are entangled, and offer little insight into where unlearning actually happen...
Hao-Ran Tang, Andrew Tan, Rajiv Khanna· 0 citations
Whether unlearning runs exhibit exploitable structure in weight space is investigated, and it is observed that models from different runs still lie in a shared evaluation-performance basin, suggesting that stronger models may be recovered through an unlearning-tailored soup strategy, reducing the need for repeated tuni...
Pu-Ning Yang, Qi-Zhou Wang, Jun-Chi Yu et al.· 0 citations
Machine unlearning (MU) aims to remove the influence of selected data from trained models, offering an efficient alternative to full retraining. With the rise of increasingly stringent privacy regulations, including the right to be forgotten, machine learning models must incorporate mechanisms that ensure compliance wh...
This work introduces ERASE, Erasure via Reconstructive Adversarial Signal Editing, a framework for on-the-go forgetting that suppresses the observable influence of private data without modifying model weights, establishing a scalable, regulation-aligned pathway for continual, privacy-conscious learning.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.