Skip to content
Open access

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes

Aug 2026 · BMC Genomics · 0 citations

TL;DR

ScGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation, and supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Abstract

Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Read PDF

Similar papers

Open access Jul 2026

LLM-powered Functional Gene Set Summarization with genesetGPT

GenesetGPT is proposed, an efficient, LLM-based framework that emphasizes both curated biological context and iterative prompt construction, thus enabling realistic summarization of heterogeneous gene sets at scale.

Jack R. Leary, Samantha Pattey, Rhonda L. Bacher · 0 citations
Open access Jul 2026

scGraphVerse: a modular workflow for single-cell gene network inference

Abstract Motivation Inferring gene networks from single-cell RNA sequencing data is challenging due to high sparsity, dimensionality, and technical noise. Current pipelines lack the multi-dataset integration and comprehensive post-processing analysis. Results scGraphVerse is an R package that integrates multiple algorithms (GENIE3, GRNBoost2, ZILGM, PCzinb, and JRF) with extensive evaluation and visualization tools. Its modular workflow supports early, late, and joint integration strategies for multi-dataset analysis, providing standardized input/output interfaces and biological interpretation tools, including community detection, pathway enrichment, and literature mining. Benchmarking on simulated data showed model-based methods (PCzinb and ZILGM) perform well with limited sample sizes, while JRF performs best as the network size and dataset numbers increase. A PBMC case study demonstrates JRF’s ability to identify literature-supported regulatory communities across donors. Availability and implementation The package is available in Bioconductor 3.22 at https://bioconductor.org/packages/release/bioc/html/scGraphVerse.html. Code and examples: https://github.com/ngsFC/scGV_analysis.

Francesco Cecere, D. De Canditiis, Annamaria Carissimo et al. · 0 citations
Preprint Aug 2026

Uncovering Cellular Resolution in scRNAseq via Unbiased Cell and Gene Network Analysis

The results indicate that GmGM provides a unified, reproducible framework for joint cell clustering and gene-network inference, capable of revealing cellular structure beyond that captured by conventional pipelines.

O. Lanzetta, L. Cutillo, Bailey Andrew et al. · 0 citations
Open access Jul 2026

MKMC enables reference-free transcriptomic analysis using k-mer representations

MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.

L. Mboning, Maciej Dlugosz, Marek Kokot et al. · 0 citations
Open access Aug 2026

pysigscore: gene signatures scoring across bulk and single-cell transcriptomics

Summary High-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures. Availability and Implementation Source code is available at https://github.com/bioinformatics-hub/pysigscore. Contact: tommaso.giacomello@phd.unibocconi.it, francesca.buffa@unibocconi.it Supplementary information Supplementary data are available at Bioinformatics online.

Tommaso Giacomello, S. Mazzara, Gennaro Abbruzzese et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.