Abstract Evaluation of reasoning language models gained importance after it was observed that they can combine their existing capabilities into novel traces of intermediate steps before task completion and that the traces can sometimes help them to generalize better than past models. We show that better performance res...
The biomedical literature contains a vast collection of omics studies, yet most published data remain functionally inaccessible for computational reuse. When raw data are deposited in public repositories, essential information for reproducing reported results is dispersed across main text, supplementary files, and code...
Alexandre Hutton, Jesse G. Meyer· PLoS Computational Biology· 0 citations
Data, R analysis code, and derived results for the psychometric evaluation of the German-language nine-item Transhumanism Scale (THS-9). The dataset comprises two independent samples (N = 200 and N = 1,051). Analyses include ordinal exploratory factor analysis, cross-validated ordinal confirmatory factor analysis, grad...
Harald Walach, Rainer J. Klement, David Martin et al.· Zenodo (CERN European Organi...· 0 citations
With the advent of large language models (LLMs), numerous software service providers are developing LLMs tailored for code generation, such as CodeLlama. However, these models can be exploited by malicious developers to generate malicious code, posing severe threats to the software ecosystem. To address this issue, we...
Kaiwen Ning, Jiachi Chen, Qingyuan Zhong et al.· ACM Transactions on Software...· 1 citation
While recent Transformer-based diffusion models have significantly advanced text-to-image (T2I) synthesis, they inherently lack explicit spatial inductive bias. Consequently, generating complex scenes with multiple objects and fine-grained attributes often leads to severe "attribute leakage" and "spatial misalignment....
Large language models (LLMs) are being explored for medical risk assessment, yet performance on standardized benchmarks does not necessarily establish clinical competency. The present study evaluated whether LLM-generated Alcohol Use Disorder (AUD) risk judgments aligned with epidemiological evidence and remained consi...
Charles M. Bolton, Kiwon Song, Katelyn A. Wasson et al.· 0 citations
Reliable spacecraft detection remains difficult due to scarce labels in extreme visual conditions, while most existing annotation-free pipelines produce indiscriminate pseudo-labeling. We propose an annotation-free detection framework that combines Vision Language Model (VLM)-based Grounded SAM 2 pseudo-labeling with a...
Semantic Verbal fluency is a core component of neuropsychological assessment,valued for its clinical sensitivity and ease of administration. Yet, its full potential is rarelyfully exploited: performance is typically reduced to a total word count, overlooking the richdynamics of semantic foraging, such as clustering (ex...
Lucie Vigreux, Victor Altmayer, Mathias Benedek et al.· 0 citations
Disaggregated large language model serving introduces a recovery problem that spans request execution, remote key-value (KV) cache reuse, and service lifecycle management. A successful response does not by itself establish that the remote cache path has recovered. This report describes KVQuake, an external fault-inject...
Yunfei Sun· Zenodo (CERN European Organi...· 0 citations
Agreement Between Large Language Models and Humans in Research Proposal Review — Data and Code This repository contains the data and code required to reproduce the analyses, statistical results, and figures presented in the associated manuscript. Files are organized by function and described below. All research proposa...
Christopher A. Gorski· Zenodo (CERN European Organi...· 0 citations