Explaining Jailbreaks: Structured and Interpretable Safety Assessment for Large Language Models
Sunghee Dong, Sungwon Yi, Kangmin Bae et al.
· 0 citations
1 paper indexed here
Fetches their full publication history.
Not the right person? Other researchers publish under this name.