Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge
LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models'hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-...