Large language model (LLM) safety judges from different model families can disagree by an order of magnitude on the same responses, with reported attack success rate (ASR) ranging from 0.04 to 0.42 on a single benign-only fine-tuned model. This is not ordinary classifier noise: the gap is driven by judges whose unsafe recall is too low to detect harmful outputs. Existing pipelines select judges by aggregate metrics such as balanced accuracy, which can look acceptable while false negatives silently suppress ASR. We propose a confidence-interval (CI) filtered two-stage judge-screening framework. Stage 1 retains only judges whose Clopper–Pearson 95% lower bound on unsafe recall exceeds a predefined floor; Stage 2 selects the highest balanced-accuracy judge among the survivors. When no judge has sufficient calibration evidence, the framework emits an explicit low-confidence or abort verdict, and downstream ASR can be accompanied by a Rogan–Gladen first-order prevalence correction. We evaluate the framework on seven labeled prompt sources and nine candidate judges spanning five base model families in two judge categories. In small-calibration regimes, it reduces the rate of selecting judges that fail the unsafe-recall screen by an order of magnitude. Across five English sources it produces a pass and reject map that is stable under the unweighted bound, with one of 45 judge-source cells reversing under design weighting; on a Chinese multilingual jailbreak source, every English-trained judge fails Stage 1 and the framework abstains. A controlled supervised fine-tuning (SFT) case study with varying benign-to-safety data ratios shows that on the same benign-only (1:0) responses the reported ASR spans a factor of 12.5 across judges, and that the lowest-recall judge underestimates the human-labeled ASR by a factor of 6.0; CI-filtered screening prevents the corresponding ASR underestimation.
Ji-Xiang Yang, Jun-Fei Yi, Jin-Han Li et al.· IEEE Access· 0 citations
Large-language-model (LLM) agents perform well in embodied benchmarks but are costly, stochastic, and difficult to audit. We propose an LLM-free zero-shot decision agent for ALFWorld that combines structured commonsense priors with adaptive calibration. The agent uses object-location priors, task templates, synonym mappings, and a hierarchical state controller over admissible commands. Because every decision is traceable to explicit knowledge entries, success and failure signals update only the responsible entries rather than all parameters. On 134 ALFWorld valid_unseen tasks, static priors obtain 67.2% success; symmetric calibration over P(obj,loc), M(target,entity), and S(word) improves this to 73.9%, with one-round convergence and CPU-only execution. We also observe that 34/134 tasks contain description-environment inconsistencies; on the consistent subset, our system reaches 93.0%. The results show that interpretable structured priors can be a practical alternative for well-specified embodied decision making.
Shengjie Ma, Jin-Han Li, Guo-An Zhang et al.· 2026 3rd World Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.