Skip to content
Book Open access

Benchmarking LLM Agents on Real-World Biological Database Curation for Data-Driven Scientific Discovery

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 10138-10149 · 0 citations · 12 references

Abstract

High-quality biological databases are the bedrock of data-driven scientific discovery. However, the construction of these resources remains a labor-intensive bottleneck, particularly for emerging research frontiers where structured data is non-existent. While LLM-based agents have catalyzed progress in downstream scientific modeling, their potential to automate the critical upstream challenge of database curation remains largely untapped. To bridge this gap, we introduce BioDataLab, a rigorous benchmark comprising 100 tasks meticulously derived from 57 high-impact database publications. BioDataLab evaluates the capability of autonomous agents to transform raw, heterogeneous biological resources into structured, analysis-ready databases. Unlike static evaluations, BioDataLab provides a fully interactive environment encompassing data retrieval, extraction, annotation, and integration, featuring process-oriented curation targets and contamination-control checks. We benchmark 11 state-of-the-art LLMs (including Gemini-3.0, GPT-5.2, and Claude-4.5) under different agent frameworks, revealing a substantial capability gap: the top-performing model achieves only a 40% success rate. Further error analysis identifies significant bottlenecks in multi-step tool orchestration and adherence to complex biological data formats. These findings underscore that while LLMs are proficient in downstream reasoning, autonomous upstream curation remains a formidable frontier. All data and codes are available at GitHub.

Read PDF

Similar papers

Review Open access Aug 2026

Detailed curation of biological samples and experimental designs for genomics using LLM-supported agentic workflows

An automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource, with performance near that of human curators, at approximately 1/20th the cost and at least 100 times the speed.

P. Pavlidis, B. O. Mancarci, A. Mãximo et al. · 0 citations
Jul 2026

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

SDABench is introduced, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics).

Chuhan Shi, Xiaoquan Ren, Sicheng Song et al. · 1 citation
Jul 2026

SciDataSailor: Deep Scientific Data Exploring

This work presents SciDataSailor, a framework for synthesizing tool-interactive trajectories by balancing broad exploration with targeted exploitation and presents SciDataSailor, a framework for synthesizing tool-interactive trajectories as Monte Carlo Tree Search (MCTS) with four task-specific mechanisms.

J. Rao, Yicheng Qiu, Chi Zhang et al. · 1 citation
Preprint Aug 2026

ChemReporter: A Framework for Curating and Exporting Large-Scale Chemical Datasets for MLIP Training

Training set quality and diversity are key determinants of the reliability of machine learning interatomic potentials (MLIPs), yet using massive datasets in full is often impractical and redundant, making intelligent data selection essential. A major bottleneck, however, is the lack of infrastructure for uniformly accessing, curating, and subsampling heterogeneous large-scale chemical datasets, which differ widely in structure, metadata, and file format. We address this gap with ChemReporter, a modular, method-agnostic framework that converts arbitrary molecular and materials datasets into a unified, queryable representation and exports the results directly into MLIP-ready training data. ChemReporter operates in three decoupled stages: processing, which parses raw datasets into a partitioned Apache Parquet repository enriched with structural, physical, and chemical metadata; querying, which filters and samples this repository via a CLI or Python API using arbitrary selection criteria, from simple physical constraints to custom, user-defined strategies; and exporting, which streams the selected subset into an HDF5 file ready for direct use in modern MLIP training frameworks. Throughout this process, every exported data point remains traceable to its original source entry, and dataset exports can be reliably reproduced given the same configuration and query database version. Because data is stored in a queryable, disk-backed format, ChemReporter can process datasets far larger than available memory, allowing it to scale to billion-structure datasets on standard compute infrastructure. ChemReporter is available on GitHub and PyPI under the Apache License 2.0.

M. Bluntzer, Jules Tilly, Christoph Brunken · 0 citations
Review Jul 2026

Evaluating Agentic Bioinformatics through Function, Evidence, and Validation

It is argued that agentic bioinformatics should be assessed through workflow correctness rather than final-answer correctness alone, and the Function--Evidence--Validation (FEV) framework is introduced, which separates demonstrated workflow operations, traceable support for actions and claims, and use-case-specific validation.

Phuc Pham, Truong-Son Hy · 0 citations
Book Open access Aug 2026

Automating End-to-End Hybrid Query Processing: Benchmark, Solution, and Insights

Hybrid queries—natural language questions over structured data that require both database capabilities and LLM reasoning—have recently emerged as a prominent research topic. However, existing solutions remain overly dependent on manual workflows, and current benchmarks are limited in scale and diversity. To bridge this gap, we present (1) HyQBench \xspace, a large-scale benchmark with 60\sim 90× more queries than prior work, built on 3× more databases; (2) AutoHyQ \xspace, an automated pipeline that can execute existing methods without manual intervention; (3) multi-dimensional, fine-grained evaluation metrics for comprehensive assessment. Through extensive experiments across multiple hybrid query approaches on diverse LLM backbones, we reveal their strengths and limitations, and identify research opportunities for advancing this emerging field. Our code and data are available at https://github.com/XMUDM/HyQBench.

Bo Li, Chenzhan Wang, Long-Kang Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.