Skip to content

Author

Jong-Soo Hong

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Guard Models Are Overconfident Where Base Models Are Uncertain

It is found that although several guard models are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections, highlighting a mismatch between guard confidence and base model unc...

Jong-Soo Hong, M. Jung, Minwoo Kim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.