Skip to content

Test-Time Unlearning via Sparse Autoencoder

Sep 2026 · 0 citations · 50 references
Computer Science

TL;DR

This work proposes ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state and introduces three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery.

Abstract

Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where stronger forgetting of target knowledge can degrade model utility, and unlearned knowledge may reappear under post-unlearning fine-tuning or prompt attacks. We propose ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state. ARIA uses sparse autoencoder (SAE) latents to train a lightweight linear detector, then applies an interpretable intervention on triggered states with negligible test-time overhead. Empirical evaluations on TOFU, R-TOFU, and WMDP show that ARIA improves the forget-retain trade-off over weight-based baselines across both a thinking model (DeepSeek-R1-Distilled-Qwen-1.5B) and an instruction model (Gemma-3-1B-it), e.g., reducing WMDP-cyber forget-set accuracy significantly while keeping MMLU within 1% of the pre-unlearning model. We further introduce three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery, and find that ARIA remains robust under all three, with forgetting changing by less than 1% under attack. A feature-level case study leveraging the interpretability of ARIA suggests that some retain degradation may reflect response styles underlying the unlearning data rather than leakage of the targeted knowledge itself, highlighting a potential source of bias in unlearning task construction.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

This work proposes Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer-level unlearning framework that selects transformer layers using a forget-to-retain significance score and provides an empirical path toward quantization-resilient unlearning.

Ravi Ranjan, O. Kotevska, Agoritsa Polyzou · 1 citation
#artificial intelligence Preprint Sep 2026

Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders

This work introduces SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features and shows theoretically that contrastive training promotes target-selective features and that the selection score controls expected background knowledge perturbation.

Itai Zehavi, Fanny Jourdan, Ulrich Aivodji · 0 citations
#artificial intelligence Preprint Sep 2026

Unmerge: Efficient Machine Unlearning via Task Arithmetic

Approximate machine unlearning seeks to remove the influence of a forget set from a trained model without full retraining. Existing gradient-based methods require data-dependent hyperparameter search, struggle when forget and retain knowledge are entangled, and offer little insight into where unlearning actually happen...

Hao-Ran Tang, Andrew Tan, Rajiv Khanna · 0 citations
#machine learning Preprint Sep 2026

UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?

Whether unlearning runs exhibit exploitable structure in weight space is investigated, and it is observed that models from different runs still lie in a shared evaluation-performance basin, suggesting that stronger models may be recovered through an unlearning-tailored soup strategy, reducing the need for repeated tuni...

Pu-Ning Yang, Qi-Zhou Wang, Jun-Chi Yu et al. · 0 citations
2026

DeepU: Deeper Granular Within-Layer Machine Unlearning

Machine unlearning (MU) aims to remove the influence of selected data from trained models, offering an efficient alternative to full retraining. With the rise of increasingly stringent privacy regulations, including the right to be forgotten, machine learning models must incorporate mechanisms that ensure compliance wh...

Anudeep Vurity, Zhi-Sheng Yan, Massimiliano Albanese · 0 citations

On-the-go Forgetting without Explicit Unlearning via ERASE

This work introduces ERASE, Erasure via Reconstructive Adversarial Signal Editing, a framework for on-the-go forgetting that suppresses the observable influence of private data without modifying model weights, establishing a scalable, regulation-aligned pathway for continual, privacy-conscious learning.

Kushal Chakrabarti, Mayank Baranwal · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.