Decoy Direction Optimization (DDO), a fast, post-hoc weight-editing defense that requires no base-model finetuning, is introduced, based on a simple mechanistic insight: ablation attacks rely on contrastive estimators to find the refusal direction.
Aashiq Muhamed, Mona T. Diab, Virginia Smith· 2 citations· ⚡1
This work introduces RAPTOR - a Role-Aware Private Training framework, which alternates shared and expert optimization and targets each failure directly, using expert-specific clipping and noise together with a public expected-owner denominator and a count-independent update schedule that avoids conditioning on private...
Duc Dm, Khai Le-Duc, D. Nguyen et al.· 0 citations
This work proposes Dynamic SAE Steering for Preference Alignment (DSPA), an inference-time method that makes sparse autoencoder (SAE) steering prompt-conditional, and audits the SAE features DSPA modifies, finding that preference directions are dominated by discourse and stylistic signals.
J. Wedgwood, Aashiq Muhamed, Mona T. Diab et al.· arXiv.org· 1 citation
This work proposes the Pairwise Dictionary Mean Correlation Coefficient (PW-MCC) as an assignment-based metric to quantify consistency and demonstrates that high levels are achievable with appropriate architectural choices for TopK SAEs on LLM activations with appropriate architectural choices.
Xiangchen Song, Aashiq Muhamed, Yujian Zheng et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.