Skip to content

Author

Kyungsik Yang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models'hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-...

Wonjun Lee, Kyungsik Yang, Gaeun Ji et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.