Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails
Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch...