When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
To mitigate the refusal-cue shortcut, sparse complementary masking is adapted as a lightweight post-hoc intervention that identifies and suppresses a small set of shortcut-associated attention heads and MLP neurons without retraining and achieves an approximately 79% relative reduction in response-initial detection failures induced by refusal cues, while preserving standard detection performance.