It is argued that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.
Abstract
Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment. We audit these assumptions for one of the most widely-used LLMs, comparing two modalities, ChatGPT's chat UI and OpenAI's API, with and without web search enabled. We use a stratified total sample of 401 prompts from two popular benchmarks, BBQ and SafetyBench, collecting 4,812 total responses across three repeated runs per prompt. Beyond standard performance measures, we evaluate model output dimensions including response consistency, response text similarity, citation grounding, and abstention behavior. For instance, chat UI responses were less accurate than API responses on both benchmarks with search disabled. Enabling web search reduced accuracy by up to 8 percentage points, and even reversed the direction of modality performance trends for one benchmark. Repeated runs of the same prompt produced inconsistent responses in up to 21\% of prompts. The two modalities also grounded answers in different citations, and abstention behavior was also inconsistent across both modalities. These results illustrate that, even within a model family, reporting only simple accuracy metrics can obscure important forms of model behavioral variation relevant to AI safety assessments. We argue that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmarks. Second, psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, and roughly ten adaptively chosen items suffice for several individual benchmarks, cutting evaluation cost by 97-99%. Third, IRT supports audits of individual models, showing that it can be used to detect naive sandbagging and changes of model behind APIs. Overall, we show IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which we recommend frontier labs and evaluators adopt.
J. Rivera, Neil Shah, D. Africa et al.· 1 citation
Power analysis is critical for assuring rigor and validity of quantitative research yet remains underutilized due to technical challenges associated with specialized software. At the same time, large language models (LLMs) are being rapidly integrated into research practice, raising interest in their potential to assist statistical and research design tasks. However, despite their widespread adoption, the reliability of LLMs in supporting statistically rigorous procedures has not been systematically evaluated, posing risks for unexamined or overly optimistic use. To address this gap, we evaluated four widely used LLMs—ChatGPT (GPT-3.5, GPT-4, GPT-4o) and Llama 3.2—across two experiments. Experiment 1 examined whether LLMs could calculate required sample sizes for common statistical tests (two-sample t-test, one-way ANOVA, and χ² goodness-of-fit test) under different prompting strategies, including direct calculation versus R/Python code generation. Experiment 2 assessed models’ ability to identify missing input parameters necessary for power analysis, which is a task that requires methodological understanding. Results revealed that GPT-4 and GPT-4o performed well when generating R code for sample size estimation, but struggled with direct numerical calculation. Furthermore, while LLMs were able to detect missing information, their reliability varied by statistical context. Findings suggest that while LLMs may offer support in structuring and initiating power analysis, they cannot substitute for expert judgment. Overall, the study underscores the importance of critically evaluating LLM performance in statistically demanding tasks. Responsible integration of LLM requires critical oversight, cross-verification, and methodological evaluation.
Hajung Kim, Jia Qi, Zhe Feng et al.· Journal of Behavioral Data S...· 0 citations
A fundamental disconnect is suggested between a model's capacity for factual accuracy and its ability to maintain social fairness, highlighting the need for multi-dimensional evaluation frameworks for small-scale systems.
M.J.F. Valdez, Arghir-Nicolae Moldovan· International Conference on...· 0 citations
This work presents the most comprehensive evaluation of LLM safety capabilities to date, systematically testing models across datasets that are organized into four distinct categories, and uncovers critical blind spots.
A large-scale assessment of the effectiveness and robustness of these automated pipelines is conducted by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which reveals a capability-safety confound that mixes model capability with apparent safety.
A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.
V. Rodionov, Shamil Assylbekov· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.