Decoy Direction Optimization (DDO), a fast, post-hoc weight-editing defense that requires no base-model finetuning, is introduced, based on a simple mechanistic insight: ablation attacks rely on contrastive estimators to find the refusal direction.
Aashiq Muhamed, Mona T. Diab, Virginia Smith· 2 citations· ⚡1
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introd...
This work proposes Dynamic SAE Steering for Preference Alignment (DSPA), an inference-time method that makes sparse autoencoder (SAE) steering prompt-conditional, and audits the SAE features DSPA modifies, finding that preference directions are dominated by discourse and stylistic signals.
J. Wedgwood, Aashiq Muhamed, Mona T. Diab et al.· arXiv.org· 1 citation
This work proposes the Pairwise Dictionary Mean Correlation Coefficient (PW-MCC) as an assignment-based metric to quantify consistency and demonstrates that high levels are achievable with appropriate architectural choices for TopK SAEs on LLM activations with appropriate architectural choices.
Xiangchen Song, Aashiq Muhamed, Yujian Zheng et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.