Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real questionnaires, however, ask each respondent a sequence of related questions. A simulated respondent should show coherent preferen...
This paper introduces RSI-router, a routing framework that constructs subtask-level model assignments and model-specific skills through recursive self-improvement over accumulated experience and establishes a stronger performance--cost Pareto frontier than 9 routing methods.
Hao Li, Hang-Fan Zhang, Zhi-Yao Cui et al.· 0 citations
Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature as endogenous self-reflection mechanisms mature.
ControlScope compares continuing generated code, editing the next tool call's data arguments, and replacing the unfinished workflow from the same public execution state to evaluate one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld.
Jing-Jie Ning, Xue-Qi Li, Yi-Bo Kong et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting, is introduced and ViDrA-adapted, an automated scorer that predicts human creativity ratings of agent drawings, is presented.
Shane K. A. Dalumura Hettige, J. Oppenlaender· 0 citations
This work presents SpecRead, a benchmark that isolates specification comprehension from generation ability, and is automatically scorable by deterministic checks, with gray-zone cases counted wrong under the conservative main scoring.
Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is evaluated only indirectly through downstream prediction. We introduce READ-Bench, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines...
Gerardo Pastrana, Hao-Jun Li, Dhruv Mehta et al.· 0 citations
Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understan...
Noah L. Schroeder, Yessy Eka Ambarwati, Yu-Ji Zhang et al.· 0 citations
Large Language Models (LLMs) perform well on medical examinations and question-answering benchmarks, but remain unreliable on medical calculation tasks that require exact numerical outputs. These calculations support high-stakes decisions such as medication dosing, organ-function assessment, and prognostic scoring, for...
The results show that PQC migration is not simply an algorithm-replacement exercise; it is an organisational transformation process requiring governance, cryptographic visibility, vendor coordination, phased implementation, and continuous monitoring.
Babatunde Oladoja, S. Tanev· Quantum Information Technolo...· 0 citations
Amyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disorder characterised by upper and lower motor neuron loss. Diagnosis is frequently delayed because of phenotypic heterogeneity, overlap with ALS mimics, and the absence of a single definitive biomarker. Artificial intelligence (AI) offers compleme...
Dhinesh Selvaraju, Kishore Durairaj, Krishna Ravi· The Egyptian Journal of Neur...· 0 citations