#artificial intelligence
May 2026
Measuring the Depth of LLM Unlearning via Activation Patching
The Unlearning Depth Score (UDS), a metric that quantifies the mechanistic depth of unlearning via activation patching, is introduced, confirming the causal approach as the most reliable for unlearning evaluation.
Jaeung Lee, Dohyun Kim, Jaemin Jo
· arXiv.org · 1 citation