An automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource, with performance near that of human curators, at approximately 1/20th the cost and at least 100 times the speed.
Abstract
We describe an automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource. Gemma is a hand-curated database of reprocessed transcriptomic studies, currently covering over 23,000 human, mouse and rat data sets largely drawn from the Gene Expression Omnibus (GEO). We developed a pipeline that uses both traditional (mechanical) and large-language models to produce detailed ontology-anchored, sample- and experiment-level annotations in accordance with our established curation guidelines. In this report, we describe benchmarking the pipeline and investigations aimed at evaluating readiness of the v1.1 Gemma curation agent for production use. Overall, performance is near that of human curators, at approximately 1/20th the cost and at least 100 times the speed. We also present preliminary exploration of triage methods for identifying agent curations that are more likely to contain errors, and thus can be forwarded for human review. We discuss the potential place of such curation approaches in bioinformatics ecosystems. Besides the software, our deliverables include the benchmark set of 500 studies and an evaluation framework that can be used to further develop the pipeline or compare to other approaches.
GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.
Xiaodan Zhang, S. Paithankar, Jing Pu et al.· bioRxiv· 0 citations
High-quality biological databases are the bedrock of data-driven scientific discovery. However, the construction of these resources remains a labor-intensive bottleneck, particularly for emerging research frontiers where structured data is non-existent. While LLM-based agents have catalyzed progress in downstream scientific modeling, their potential to automate the critical upstream challenge of database curation remains largely untapped. To bridge this gap, we introduce BioDataLab, a rigorous benchmark comprising 100 tasks meticulously derived from 57 high-impact database publications. BioDataLab evaluates the capability of autonomous agents to transform raw, heterogeneous biological resources into structured, analysis-ready databases. Unlike static evaluations, BioDataLab provides a fully interactive environment encompassing data retrieval, extraction, annotation, and integration, featuring process-oriented curation targets and contamination-control checks. We benchmark 11 state-of-the-art LLMs (including Gemini-3.0, GPT-5.2, and Claude-4.5) under different agent frameworks, revealing a substantial capability gap: the top-performing model achieves only a 40% success rate. Further error analysis identifies significant bottlenecks in multi-step tool orchestration and adherence to complex biological data formats. These findings underscore that while LLMs are proficient in downstream reasoning, autonomous upstream curation remains a formidable frontier. All data and codes are available at GitHub.
Jiaxian Yan, Xi Fang, Jintao Zhu et al.· Proceedings of the 32nd ACM...· 0 citations
An automated pipeline is developed that uses large language models to find single-gene perturbation experiments in NCBI-GEO and reconstruct their case-control sample groupings, along with the perturbed gene, perturbation type and cell line as structured, ontology-normalised fields.
GenesetGPT is proposed, an efficient, LLM-based framework that emphasizes both curated biological context and iterative prompt construction, thus enabling realistic summarization of heterogeneous gene sets at scale.
Jack R. Leary, Samantha Pattey, Rhonda L. Bacher· bioRxiv· 0 citations
LLMBDC (Large Language Model for Biological Domains Oriented Clustering of Gene Ontology) provides a scalable, reproducible, and interpretable route to context-aware, system-level interpretation of GO enrichment results while preserving biological specificity.
The rapid advancement of high-throughput technologies has led to an explosion of biological data and a subsequent surge in bioinformatics analysis tools, thereby creating an urgent demand for automated bioinformatics workflows. Recently Large Language Models (LLMs) and LLM-based agents show great potential in this area. However, existing benchmarks primarily focus on static question-answering (QA) tasks, failing to capture the knowledge-action gap between understanding tool usage and executing complex bioinformatics workflows. Furthermore, current evaluation paradigms often prioritize algorithmic success rates, while neglecting the biological validity. Moreover, the construction of execution benchmarks is challenging due to complex environmental dependencies and the high cost of manual annotation, leading to poor scalability. In this study, we propose BioFlowBench, a comprehensive benchmark designed to shift from static knowledge assessment to dynamic execution evaluation in bioinformatics tool utilization. First, we construct a multi-layered dataset consisting of 5,071 test samples, including Syntax Understanding, Contextual Application and Real-world Execution. Second, we introduce BioGen, an agent-based pipeline designed for the automated generation of executable benchmarks. By creating compact, low-overhead synthetic data, BioGen facilitates low-cost and large-scale testing. Third, we propose a multi-dimensional evaluation framework comprising static knowledge, structural integrity, functional validity, and efficiency metrics. Our experiments reveal that: (1) A significant gap exists between static QA and dynamic execution tasks, with top LLMs perform well on static QA but falter in real-world execution scenario; (2) specialized agents outperform general models in real-world execution through environmental interaction and iterative refinement; and (3) domain knowledge remains the primary bottleneck, often leading to executable but biologically inaccurate outputs. The code is available at: https://github.com/YufeiHouAnne/BioFlowBench and the dataset can be accessed at: https://www.scidb.cn/detail?dataSetId=aee284681d674f53bfc6dae44635e773.
Yufei Hou, Jiajia Wang, Ke Xiang et al.· Proceedings of the 32nd ACM...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.