Skip to content
Open access

pysigscore: gene signatures scoring across bulk and single-cell transcriptomics

Aug 2026 · bioRxiv · 0 citations · 16 references
Biology

Abstract

Summary High-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures. Availability and Implementation Source code is available at https://github.com/bioinformatics-hub/pysigscore. Contact: tommaso.giacomello@phd.unibocconi.it, francesca.buffa@unibocconi.it Supplementary information Supplementary data are available at Bioinformatics online.

Read PDF

Similar papers

Open access Jul 2026

MKMC enables reference-free transcriptomic analysis using k-mer representations

MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.

L. Mboning, Maciej Dlugosz, Marek Kokot et al. · 0 citations
Open access Jul 2026

scGraphVerse: a modular workflow for single-cell gene network inference

Abstract Motivation Inferring gene networks from single-cell RNA sequencing data is challenging due to high sparsity, dimensionality, and technical noise. Current pipelines lack the multi-dataset integration and comprehensive post-processing analysis. Results scGraphVerse is an R package that integrates multiple algorithms (GENIE3, GRNBoost2, ZILGM, PCzinb, and JRF) with extensive evaluation and visualization tools. Its modular workflow supports early, late, and joint integration strategies for multi-dataset analysis, providing standardized input/output interfaces and biological interpretation tools, including community detection, pathway enrichment, and literature mining. Benchmarking on simulated data showed model-based methods (PCzinb and ZILGM) perform well with limited sample sizes, while JRF performs best as the network size and dataset numbers increase. A PBMC case study demonstrates JRF’s ability to identify literature-supported regulatory communities across donors. Availability and implementation The package is available in Bioconductor 3.22 at https://bioconductor.org/packages/release/bioc/html/scGraphVerse.html. Code and examples: https://github.com/ngsFC/scGV_analysis.

Francesco Cecere, D. De Canditiis, Annamaria Carissimo et al. · 0 citations
Open access Aug 2026

CoTRA: an integrated R/Shiny framework for transparent bulk and single-cell RNA-seq analysis

Bulk and single-cell RNA sequencing (scRNA-seq) have become essential for investigating disease mechanisms and identifying diagnostic biomarkers. However, the growing volume of transcriptomic data remains difficult to reuse efficiently for many researchers. Downstream analysis often requires multiple statistical, visualization, and reporting tools, creating fragmented workflows that reduce transparency and reproducibility, particularly when analyzing scRNA-seq data. To address this gap, we developed CoTRA (Comprehensive Toolbox for RNA-seq Analysis), an open-source R/Shiny package for bulk and scRNA analysis. CoTRA integrates established methods into modular workflows, exposes parameters, and offers alternatives at selected stages. It supports bulk RNA-seq quality assessment, differential expression, annotation, enrichment, and reporting, as well as scRNA quality control, dimensionality reduction, clustering, marker identification, cell-type annotation, differential abundance, trajectory inference, pathway activity, and cell-cell communication. CoTRA runs on workstations or HPC environments without mandatory external data submission and was tested on Linux, Windows, and macOS. Compared with 14 other platforms for bulk RNA-seq/scRNA-seq, CoTRA supported 46 of 49 predefined functionality criteria. Tool validation using published rd10 retinal bulk RNA-seq identified 1,947 shared differentially expressed genes with concordant direction and strong log2 fold-change agreement. A retinal scRNA-seq case study demonstrated appropriate clustering, cell-type resolved analysis, and pathway activity scoring. CoTRA provides a graphical environment for bulk and single-cell RNA-seq analysis while retaining parameter transparency, methodological flexibility, and reproducible outputs. Strong concordance with the published bulk RNA-seq analysis supports the workflow consistency, while the single-cell case study demonstrates its applicability to advanced scRNA-seq analysis. The source code is freely available at https://github.com/UmairSeemab/CoTRA.

Umair Seemab, Katri Vainionpaa, Ziaurrehman Tanoli et al. · 0 citations
Open access Aug 2026

HopRatio: Profiling single-sample transcriptomic dysregulation using stable gene-ordering relationships

Summary Single-sample transcriptomic analysis can provide gene-level views of how individual tumors deviate from matched reference states, complementing cohort-level differential expression. Here, we present HopRatio, a rank-based framework that quantifies, for each gene in each sample, the fraction of stable reference gene-ordering relationships inverted relative to a context-matched reference cohort. By relying on within-sample ranks rather than cross-sample expression magnitudes, HopRatio enables sample-specific scoring without cross-sample normalization for score calculation. Across 16 cancer types from The Cancer Genome Atlas with tissue-matched Genotype-Tissue Expression references, HopRatio generated individualized dysregulation profiles that supported tumor-normal discrimination using single genes and compact multi-gene panels. Recurrent high-performing features defined a 246-gene set enriched for developmental, membrane-associated, and ion-transport programs and associated with poor survival across cancers. Benchmarking in the Sequencing Quality Control dataset supported its robustness relative to commonly used differential expression methods, highlighting a scalable route for interpretable individualized transcriptomic analysis.

Yue Zhao, Bo Gao, Rui-Zhong Chen · 0 citations
Jul 2026

Modular RNA-seq Analytics for Exploratory Biomarker Discovery using Public Data.

Publicly available RNA sequencing (RNA-seq) data provide a cost-effective springboard for biomarker discovery. However, heterogeneity across studies often complicates analysis. This study presents a modular analytics pipeline that combines publicly available datasets with established open-source tools; standardizes quality control, differential expression analysis, and pathway analysis; and leverages competitive machine learning to unify disparate RNA-seq datasets for robust biomarker identification. The workflow is demonstrated across three disease contexts, ranging from small pilot datasets to larger, integrated analyses: (1) identifying differentially expressed gene signatures associated with COVID-19 severity, (2) combining differential expression and machine learning techniques to analyze multi-cohort sepsis datasets, resulting in concise biomarker panels, and (3) utilizing both bulk and single-cell data to examine tissue and cell-type specificity of N-acyl-phosphatidylethanolamine phospholipase D in atherosclerosis. These applications demonstrate how an adaptable, modular pipeline using open-source tools can repurpose public data to reduce noise, generate new hypotheses, and reveal meaningful biological insights, thereby establishing a foundation for future research and underscoring the importance of public data in exploratory biomarker discovery.

Cheryl L. Sesler, Lukasz S. Wylezinski, Guzel I. Shaginurova et al. · 0 citations
Aug 2026

A statistical framework for disease classification with scRNA-Seq Data

This work introduces a two-stage statistical framework for interpretable patient-level disease classification from single-cell data, and recovered biologically coherent, cell-type specific gene signatures consistent with known disease mechanisms, demonstrating improved interpretability without sacrificing predictive accuracy.

Zhi-Wei Xiao, William Torous, Jeffrey B. Cheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.