Skip to content

A statistical framework for disease classification with scRNA-Seq Data

Aug 2026 · bioRxiv · 0 citations · 25 references
Biology

TL;DR

This work introduces a two-stage statistical framework for interpretable patient-level disease classification from single-cell data, and recovered biologically coherent, cell-type specific gene signatures consistent with known disease mechanisms, demonstrating improved interpretability without sacrificing predictive accuracy.

View source

Similar papers

Aug 2026

A topology-based framework for robust cancer-associated gene signature identification from scRNA-seq data.

Cancer transcriptomics faces a fundamental challenge: conventional gene selection methods capture statistical variance but fail to decode the intrinsic geometric architecture of high-dimensional single-cell RNA sequencing data, leaving critical cancer-specific signals obscured by noise and biological heterogeneity. We present a topology-guided framework that harnesses persistent homology to extract structurally invariant, cancer-associated gene signatures from scRNA-seq data - moving beyond gene-level statistics toward shape-aware biological discovery. The framework integrates highly variable feature selection and dimensionality reduction with Vietoris-Rips filtration-based gene correlation topology, followed by a dual-stage stability-driven classification strategy that identifies samples exhibiting reproducible cancer-specific topological patterns. Topologically significant genes are rigorously validated through differential expression analysis, ROC/AUC evaluation, KEGG pathway enrichment, protein-protein interaction network analysis, and literature evidence. Against conventional HVF+PCA-based selection, the TDA framework delivers markedly superior discriminative power, substantially higher literature-supported biological relevance, and dramatically more focused cancer-specific pathway enrichment - while converging to compact, functionally coherent gene sets that conventional approaches cannot achieve. In breast cancer, the framework reveals a dominant mitotic regulatory module centered on cell cycle dysregulation, while colorectal cancer is characterized by extracellular matrix remodeling and tumor microenvironment-driven mechanisms - demonstrating cancer-type-specific biological fidelity. Critically, the framework identifies computationally prioritized novel candidate biomarkers absent from standard pathway databases yet exhibiting topological and statistical significance. This work establishes persistent homology as a transformative paradigm for transcriptomic biomarker discovery, offering a principled, structure-aware foundation for precision oncology and next-generation cancer diagnostics.

Sudarshana Gogoi, S. Bandyopadhyay, S. Bera et al. · 0 citations
Jul 2026

CDState Resolves Malignant Cell Heterogeneity from Bulk Tumor RNA-Sequencing Data.

Intratumor transcriptional heterogeneity (ITTH), defined as the coexistence of diverse cell states within a single tumor, complicates cancer treatment and contributes to variable therapeutic responses. Although single-cell RNA sequencing (scRNA-seq) can resolve this complexity, its cost and technical demands limit large-scale use. Bulk RNA sequencing (bulk RNA-seq) provides a scalable alternative but requires computational methods to deconvolve bulk transcriptomes into distinct cell states. Existing supervised approaches rely on accurate reference data, which are lacking for many cancer types, while unsupervised methods are not tailored to capture heterogeneity within the malignant compartment. To address these limitations, we developed CDState, an unsupervised deconvolution method based on nonnegative matrix factorization with a sum-to-one constraint and a cosine-similarity-based optimization, which infers malignant cell states using bulk RNA-seq data. CDState demonstrated robustness using pseudobulk scRNA-seq datasets from five cancer types, outperforming existing unsupervised methods in estimating both state-specific gene expression and cell proportions. Applied to 33 cancer types from The Cancer Genome Atlas, CDState revealed recurrent gene programs, including epithelial-mesenchymal transition, MYC targets, and oxidative phosphorylation, as major contributors to malignant ITTH. The malignant state proportions were linked to clinical features, including patient survival and therapeutic response. Finally, mutations and copy number alterations in genes such as TP53, KRAS, PIK3CA, SOX2, and SATB1 were identified as potential genetic drivers of malignant cell ITTH across cancer types. This study demonstrates the utility of CDState for characterization of malignant cell states from bulk RNA-seq data, establishing a framework for investigating malignant cell ITTH in large-scale cancer atlases.

Agnieszka Kraft, J. Yates, Florian Barkmann et al. · 0 citations
Preprint Aug 2026

Uncovering Cellular Resolution in scRNAseq via Unbiased Cell and Gene Network Analysis

The results indicate that GmGM provides a unified, reproducible framework for joint cell clustering and gene-network inference, capable of revealing cellular structure beyond that captured by conventional pipelines.

O. Lanzetta, L. Cutillo, Bailey Andrew et al. · 0 citations
Open access Jul 2026

Deconvolution-derived cell-type expression targets for personal genome sequence-to-expression prediction

Sequence-to-function models learn regulatory features from genomic sequence, but they remain limited in their ability to predict gene-expression differences among individuals. Cell-type-specific regulatory effects may be obscured in bulk RNA sequencing, whereas paired genotype and single-cell expression cohorts remain small. We evaluated whether deconvolution of bulk RNA-seq could provide scalable cell-type-specific targets for personal-genome expression prediction. GTEx v8 bulk RNA-seq from six tissues was deconvolved with BayesPrism using single-nucleus reference profiles, producing targets across 83 tissue–cell-type contexts. Deconvolved expression agreed with matched pseudobulked GTEx single-nucleus RNA-seq, with median donor-level Pearson correlations across genes ranging from 0.53 to 0.73 by tissue. We compared genotype-feature models, regressors trained on frozen Enformer representations, and fine-tuned Enformer and Borzoi models. Across random and nonlinear-enriched gene sets, sequence-derived approaches generally outperformed genotype-feature baselines, while frozen Enformer features were competitive with end-to-end fine-tuning. For the random gene set, Fisher-averaged Pearson correlations were 0.122–0.142 for sequence-derived approaches and 0.081–0.086 for genotype-feature baselines in a coverage-aware sensitivity analysis. Model performance was positively associated with deconvolution–pseudobulk agreement for sequence-derived models (r = 0.35–0.43 across tissue–cell-type contexts), suggesting that target reliability may constrain downstream prediction. Context-specific Enformer fine-tuning did not materially out-perform a shared, combined-context strategy. These results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints. Frozen pretrained representations provide a computationally efficient and competitive baseline for personal sequence-to-expression modeling.

Stephen Sim, Li Shen · 0 citations
Open access Aug 2026

pysigscore: gene signatures scoring across bulk and single-cell transcriptomics

Summary High-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures. Availability and Implementation Source code is available at https://github.com/bioinformatics-hub/pysigscore. Contact: tommaso.giacomello@phd.unibocconi.it, francesca.buffa@unibocconi.it Supplementary information Supplementary data are available at Bioinformatics online.

Tommaso Giacomello, S. Mazzara, Gennaro Abbruzzese et al. · 0 citations
Open access Aug 2026

Refining scRNA-Seq Clusters: The Power of Feature Selection

Feature selection is critical for resolving cell-type heterogeneity in single-cell RNA sequencing (scRNA-seq). DUBStepR (Determining the Underlying Basis using Stepwise Regression) is a widely used gene selection method for scRNA-seq designed to identify feature genes that maximize cell-type separation. DUBStepR has been reported to perform effectively in this domain; however, its reliance on linear Pearson correlation and rigid thresholding limits its effectiveness on complex, high-dimensional datasets. Three enhancements are presented in this study: RFCell-DUBStepR, which uses random forests to capture expression-level importance; Copula-DUBStepR, which models non-linear correlations via Gaussian Copulas; and Zqt-DUBStepR, which utilizes quantile-based selection for improved gene retention. Using both simulated and real-world datasets (scRNA-seq), these modifications are shown to resolve the biases of the original algorithm. The modified methods consistently select a more representative gene set and yield higher clustering accuracy across varying levels of biological complexity. These findings establish the modified DUBStepR frameworks as more reliable tools for high-fidelity subpopulation identification in downstream single-cell analysis.

Ching-Hsuan Chen, Chen-An Tsai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.