Preprint
Aug 2026
Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning
The central finding is a zero-refusal phenomenon: modern LLMs decline to discuss restricted books in only 0.07% of cases, effectively invalidating the premise of jailbreaking research for this content class.
Xuchen Yu, Emily Knox, Haohan Wang
· 0 citations