Skip to content
Open access

Automated generation of a gene perturbation transcriptomic atlas using large language models

Aug 2026 · bioRxiv · 0 citations · 35 references
Biology

TL;DR

An automated pipeline is developed that uses large language models to find single-gene perturbation experiments in NCBI-GEO and reconstruct their case-control sample groupings, along with the perturbed gene, perturbation type and cell line as structured, ontology-normalised fields.

Abstract

Public transcriptomic repositories contain thousands of gene perturbation experiments, a valuable resource for understanding gene function, but perturbation metadata are not structured, which blocks systematic reuse. Existing perturbation atlases depend on expert manual curation, so they are costly to maintain and infrequently updated, while automated grouping approaches neither identify which samples form the perturbation arm nor recover the perturbed gene. Here we develop an automated pipeline that uses large language models to find single-gene perturbation experiments in NCBI-GEO and reconstruct their case-control sample groupings, along with the perturbed gene, perturbation type and cell line as structured, ontology-normalised fields. We manually curated 3,300 GEO experiments with sample-level case-control assignments and release these as an open benchmark (2,400 training, 600 validation, 300 temporally held-out test). Reasoning models and task-specific finetuning substantially improved identification of valid perturbation groups, with the best model reaching precision 0.925 and recall 0.836 on the test set. Applied at scale, the pipeline generated an atlas of 6,802 gene perturbation expression signatures from 4,453 GEO experiments, covering 2,907 uniquely perturbed genes. An R package, perturbMatch, supports exploration of the atlas and querying of user-supplied expression signatures against it using similarity scoring, so users can identify experiments that recapitulate a transcriptional state of interest.

Read PDF

Similar papers

Open access Aug 2026

Automating scientific annotations for open transcriptomic profiles via multi-stage agents

GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.

Xiaodan Zhang, S. Paithankar, Jing Pu et al. · 0 citations
Open access Aug 2026

PxFquery: A Bioinformatics Tool for Large Language Model-Assisted Functional Analysis of Large-Scale Perturbation Signatures

Background/Objectives: Genetic and chemical perturbation experiments provide a systematic approach to investigate cellular responses. Large-scale resources such as Connectivity Map (CMap) contain extensive transcriptomic perturbation signatures that support functional interpretation and perturbation retrieval. However, existing access to these resources mainly relies on structured inputs, making it challenging to connect natural-language perturbation questions with experimental evidence. Methods: To address this challenge, we developed PxFquery, an evidence-grounded tool that enables natural-language exploration of large-scale perturbation resources. PxFquery uses an LLM-assisted workflow to interpret biological questions and organize responses based on retrieved perturbation evidence. It converts over 500,000 CMap/LINCS perturbation signatures into a compact functional response space and supports bidirectional perturbation-function queries. Results: Despite the sparse perturbation coverage of CMap (5.9%), PxFquery integrated related perturbation evidence and improved access to perturbation resources. Across evaluated genetic and chemical perturbation-to-function and function-to-perturbation queries, PxFquery outputs showed closer agreement with experimental reference rankings than direct and PubMed-augmented LLM approaches (paired Wilcoxon tests; p < 0.05). Functional response representation reduced storage requirements to 0.068–0.24% of the original resources and enabled lightweight deployment through a Python package (v0.5.29), website, and AI workflow interfaces. Conclusions: PxFquery provides a natural-language interface for exploring large-scale perturbation resources while maintaining connections to experimental evidence. It lowers the barrier to accessing these resources and enables their integration into diverse AI workflows.

Jun Cao, Xiaoyue Wang · 0 citations
Review Open access Aug 2026

Detailed curation of biological samples and experimental designs for genomics using LLM-supported agentic workflows

An automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource, with performance near that of human curators, at approximately 1/20th the cost and at least 100 times the speed.

P. Pavlidis, B. O. Mancarci, A. Mãximo et al. · 0 citations
Open access Jul 2026

LLM-powered Functional Gene Set Summarization with genesetGPT

GenesetGPT is proposed, an efficient, LLM-based framework that emphasizes both curated biological context and iterative prompt construction, thus enabling realistic summarization of heterogeneous gene sets at scale.

Jack R. Leary, Samantha Pattey, Rhonda L. Bacher · 0 citations
Open access Jul 2026

MKMC enables reference-free transcriptomic analysis using k-mer representations

MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.

L. Mboning, Maciej Dlugosz, Marek Kokot et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.