Skip to content

Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

Safety-Flag is introduced, which places seven widely used safety benchmarks (BeaverTails, XSTest, Ethics, WildGuard, Aegis, ToxiChat, and ToxiGen) into a single balanced flag / do-not-flag protocol.

Abstract

Large language models are increasingly used for content moderation, but most evaluations still report aggregate accuracy on individual benchmarks. We introduce Safety-Flag, which places seven widely used safety benchmarks (BeaverTails, XSTest, Ethics, WildGuard, Aegis, ToxiChat, and ToxiGen) into a single balanced flag / do-not-flag protocol. We release item-level decisions and confidence scores for six general-purpose LLMs and four dedicated guards, together with three reference models, evaluated on the same items. Safety-Flag measures three dimensions of moderator reliability: error direction, probability calibration, and confidence-based error ranking for human review. They often disagree. Aggregate accuracy does not reveal error direction: one model flags $85\%$ of benign content, whereas another misses $54\%$ of harmful content. All six general-purpose models are overconfident; fitting one temperature per model reduces calibration error by $2.8$--$6.0\times$ without changing predicted labels or confidence ordering. Confidence-based abstention lowers selective risk for every model, although the gains depend on how well confidence ranks errors. Dedicated guards produce fewer false alarms and are better calibrated, but several have higher miss rates outside their documented coverage. We release the benchmark, fixed item lists, evaluation code, per-item model outputs, and leaderboard at: https://github.com/yibo-hu-lab/safety-flag-benchmark.

View source

Similar papers

Preprint Aug 2026

Item Response Theory for AI Safety

Overall, it is shown IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which frontier labs and evaluators adopt.

J. Rivera, Neil Shah, D. Africa et al. · 1 citation
Preprint Sep 2026

Safety Monitors Mostly Catch What the Model Already Refuses

Safety monitors are evaluated by recall on harmful prompts, regardless of whether the target model would answer them. Yet a monitor matters most on the prompts the model does answer. We measure recall on exactly those prompts, defined by sampling the target model and judging its responses. Across four text guards, two...

Sripad Karne · 0 citations
Open access Sep 2026

Development of a Framework for Evaluating Large Language Model Safety and Reliability: a Proof-of-Concept Evaluation

Large language models (LLMs) are entering clinical decision support faster than methodology can characterise their safety. Aggregate accuracy treats all errors as interchangeable and cannot support safe deployment under Software as a Medical Device (SaMD) and EU AI Act frameworks. To develop and demonstrate a framework...

Fang-Yan Liu, Zhi Liu, Xiaolu Fei et al. · 0 citations
#machine learning Preprint Sep 2026

SafeLLM4SE: Statistical Evaluation and Reporting for LLM-based Software Engineering Systems

Large language models (LLMs) are increasingly used for software engineering tasks, yet their stochastic behavior challenges the validity, reproducibility, and comparability of their evaluations. Conventional practices such as reporting a single output, an average score, best-of-N, or pass@k performance can obscure vari...

Francisco Ortin · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.