No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios
This work presents the most comprehensive evaluation of LLM safety capabilities to date, systematically testing models across datasets that are organized into four distinct categories, and uncovers critical blind spots.
Afshin Orojlooyjadid, Hitesh Laxmichand Patel
· 0 citations