When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
It is found that defenses rarely improve downstream capability, but instead vary in how they trade safety gains against usability and efficiency as well as performance impact, over-refusal on benign inputs, and inference cost.