SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models
System-level results provide system-level evidence that deployment-time safety depends not only on the capability of individual guardrails, but also on how their decisions are organized and coordinated.
Xing-Ru Zhou, Luis Sentis, Aarti Choudhary
· 0 citations