BixBench3, a benchmark that measures the capacity of AI agents to process raw biological data through to scientific results, reveals that LLMs vary substantially in their ability to execute multiple sequential analysis steps coherently, manage large quantities of raw data, and work across scientific domains.
Abstract
Artificial intelligence (AI) promises to accelerate biological research by automating computational analyses. Yet the ability of AI agents to carry out computational biology at the scale of complete research studies has not been systematically evaluated. Here we introduce BixBench3, a benchmark that measures the capacity of AI agents to process raw biological data through to scientific results. We designed BixBench3 tasks to mirror the delegation of work from a scientist to an agent: the scientist chooses the research question and high-level methods, then delegates implementation of all analyses to the agent. In each task, an agent receives a research objective, methodological guidance, and raw data derived from a published scientific study, and must execute a sequence of analyses to achieve the research objective. The data artifacts resulting from these analyses - such as peak call matrices or differential expression tables - are programmatically graded against the corresponding artifacts generated and reported in the original study. Across 20 BixBench3 tasks encompassing the generation of 138 unique artifacts, we find that 13 frontier models achieve scores ranging from 0.00 for Gemini 3.1 Flash Lite to 0.48 for GPT 5.6 Sol. Agents perform worse on tasks with larger raw datasets (0.36 on tasks with<100 GB versus 0.10 on tasks with>100 GB) and on analyses requiring more sequential steps (0.36 at 1-2 steps vs 0.24 at 3+). On average, agents use 6.8 hours, 102 million tokens, and $43 to complete each task, with the longest attempts consuming 24 hours, 1.07 billion tokens, and $525. Notably, the highest-scoring agents used fewer tokens and were cheaper than less performant options. These results reveal that LLMs vary substantially in their ability to (1) execute multiple sequential analysis steps coherently, (2) manage large quantities of raw data, and (3) work across scientific domains.
This article examines the emerging paradigm of agentic AI for scientific discovery, traces the conceptual shift from tools to agents, lays out a six-stage workflow spanning literature synthesis to manuscript generation, and reviews practical systems in chemistry, equation discovery, materials science, and general machine learning research.
Alexander Taktakidze· Longevity Horizon· 0 citations
Overall, it is found that the use of coding agents in scientific computing holds great promise for accelerating scientific research and increasing the reliability of critical systems, but that outstanding concerns remain.
Jeremiah H. Li, Alex Rubinsteyn, Sergey Feldman et al.· bioRxiv· 0 citations
Agentic artificial intelligence (AI) systems that are capable of planning and executing multi-step analytical tasks are increasingly available to environmental health researchers, but their reliability in real-world practice has not been fully explored. This paper describes a human-in-the-loop agentic framework for environmental health research, involving the review, verification, and correction of AI-generated data analysis code and results at each step - a process that mirrors the mentorship structure of traditional research teams. This approach offers a path toward rigorous, reproducible use of agentic AI in routine data-rich environmental health research. We illustrate this framework through a case study analyzing nitrogenous organic contaminants at U.S. Superfund sites, a chemical family linked to the emerging tire-derived contaminants 6PPD and 6PPD-quinone. Using an agentic large language model to generate R code for data filtering, spatial mapping, and cluster analysis, we document instances where initial agentic AI outputs benefited from a human-in-the-loop process to produce more rigorous and reproducible results. We conclude that effective use of agentic AI requires both domain expertise to frame questions and evaluate outputs, and coding literacy to guide the AI's approach, while outlining future opportunities for agentic AI workflow to advance the environmental health sciences.
Scientific progress depends not only on finding solutions, but on learning the rules that explain why they work and using that understanding to design better experiments. We introduce science sandboxes, a framework for studying this capability in AI agents through repeated cycles of experimentation, feedback, and hypothesis revision. Science sandboxes invite an agent to query the natural world in different ways, ranging from"wet"physical experiments, to"damp"predictive models trained on empirical data, to"dry"invented rules. By establishing a common experimental loop and a protocol for evaluating agents within it, science sandboxes allow assessment of both quantitative performance on specific metrics and qualitative scientific reasoning, across a spectrum of empirical verifiability. Here, we instantiate this framework in two biological settings, models of regulatory genomics and protein fitness prediction, and examine the capabilities of frontier agents. Across these settings, we could see when agents successfully optimized a quantitative metric without understanding the rules underlying the system. In particular, their scientific reasoning deteriorated when they encountered systems whose rules fell outside familiar biological priors. By highlighting such failure modes, science sandboxes make the frontier of scientific capability measurable and provide a controlled setting in which to study and ultimately expand it.
Arya S. Rao, Rodrigo I. Castro, Sager J. Gosai et al.· 0 citations
BioDataLab evaluates the capability of autonomous agents to transform raw, heterogeneous biological resources into structured, analysis-ready databases, and underscores that while LLMs are proficient in downstream reasoning, autonomous upstream curation remains a formidable frontier.
Jiaxian Yan, Xi Fang, Jintao Zhu et al.· Proceedings of the 32nd ACM...· 0 citations
This survey bridges the existing gap by presenting a comprehensive blueprint for scientific agents' design and introduces a unified taxonomy based on capability envelope and capability maturity, characterizing both the scope of scientific workflow coverage and the reliability of agent behavior under realistic research conditions.
Xinming Wang, Jian Xu, Sheng Lian et al.· IEEE Transactions on Pattern...· 9 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.