Guard Models Are Overconfident Where Base Models Are Uncertain
It is found that although several guard models are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections, highlighting a mismatch between guard confidence and base model unc...