Skip to content

Author

Bernard Ghanem

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection

Rank-One Safety Injection (ROSI), a white-box method that amplifies a model's safety alignment by permanently steering its activations toward the refusal-mediating subspace, is proposed, suggesting that targeted, interpretable weight steering is a cheap and potent mechanism to improve LLM safety, complementing more resource-intensive fine-tuning paradigms.

H. Shairah, Hasan Abed Al Kader Hammoud, G. Turkiyyah et al. · 7 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.