Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in which language must discriminate between several objects. We build SLIM, a pushing benchmark with several small objects and paired visual and language goals on identica...
Florian Strohm, Patrick Wagner, Jannik Schwab et al.· 0 citations
When two AI agents disagree, who persuades whom? As multi-agent systems increasingly combine language models of different families and sizes, the answer can determine which judgments survive interaction. Measuring persuasion as the probabilistic shift in an agent's decision after a single exchange with a dissenting pee...
Frida Nøhr Laustsen, Marie Haahr Petersen, Victoria Popa et al.· 0 citations
Large language model (LLM)-based self-evolving search is a promising approach to scientific discovery. However, high-fidelity evaluation of every candidate is prohibitively expensive in some domains. Self-evolving systems in such settings therefore rely on low-cost but imperfect proxy rewards, which may assign high sco...
Safety monitors help safeguard language-model agents interacting with external tools and environments, but conservative monitoring can generate many false alarms, consuming extensive review resources and weakening trust in alerts. Because false and genuine alarms often remain interleaved in native monitor scores, obtai...
Xi-Chen Yan, Chong-Yang Gao, Ke-Zhen Chen et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Suppressing a small set of routed experts can weaken the safety behavior of a sparse Mixture-of-Experts (MoE) language model without retraining. Which experts to suppress is therefore a security question, and the usual answer is activation frequency, but frequency measures use, not influence. We test an alternative: ro...
Md Nurul Absar Siddiky, Liu-Wan Zhu, Yi Dong· 0 citations
Language model agents are deployed with a harness, the software around the model that manages its context, tools, and feedback. When such an agent is distilled into a smaller one, the harness stays in place, so the student mainly needs the teacher-specific abilities that the harness cannot provide, such as acting corre...
Moonseok Choi, Taehong Moon, G. Nam et al.· 0 citations
Vision-language models (VLMs) are powerful listwise rerankers for multimodal retrieval, but high inference costs restrict them to evaluating small local candidate views. Existing multi-call strategies rely on fixed schedules, wasting expensive VLM calls on uninformative candidate pairs and easy queries. To address this...
Wen-Teng Chen, Jiachen Zhu, Shan Rong et al.· 0 citations
HIOF: Quantum Biodynamics (HIOF-QB) A Seven-Paper Series on When a Quantum Event Becomes Biological Function This is a short, non-technical companion to the 7-paper QB series (QB 0 to QB VI). It is written for readers with a background in quantum biology, biophysics, chemical physics, or systems biology. It is not writ...
Wai-Hung (Pan) Tam· Zenodo (CERN European Organi...· 0 citations
Small code language models are now easy to run on a developer's own laptop, and one thing people ask them to do is a quick security pass over code before it ships. I wanted to know how far that trust holds for one narrow but dangerous flaw family, cryptographic API misuse. I built CryptoBench, a set of thirty-six vulne...
Sunny Chokshi· Zenodo (CERN European Organi...· 0 citations
HIOF: Quantum Biodynamics (HIOF-QB) A Seven-Paper Series on When a Quantum Event Becomes Biological Function This is a short, non-technical companion to the 7-paper QB series (QB 0 to QB VI). It is written for readers with a background in quantum biology, biophysics, chemical physics, or systems biology. It is not writ...
Wai-Hung (Pan) Tam· Zenodo (CERN European Organi...· 0 citations
Prompt Smells and Optimization Dataset & Local AI Agent is an open-access research dataset and autonomous evaluation toolkit for studying Prompt Smells (anti-patterns) in Large Language Models (LLMs) and Prompt Engineering education. 📊 Dataset Overview The dataset contains 435 curated English examples structured speci...
Md.Ferdaus Alam· Zenodo (CERN European Organi...· 0 citations
Imagine removing everything from the Universe: the stars, the galaxies, matter, radiation, atoms and particles. Then remove space itself. Remove distance, directions, and the familiar distinction between “here” and “there.” Finally, remove even the clock on the wall of the Universe. What remains? This work begins there...
Marco Cacci· Zenodo (CERN European Organi...· 0 citations