Skip to content

Author

Sampad Mohanty

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal is encoded by transplanting weights from aligned m...

Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.