Summary High-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures. Availability and Implementation Source code is available at https://github.com/bioinformatics-hub/pysigscore. Contact: tommaso.giacomello@phd.unibocconi.it, francesca.buffa@unibocconi.it Supplementary information Supplementary data are available at Bioinformatics online.
MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.
L. Mboning, Maciej Dlugosz, Marek Kokot et al.· bioRxiv· 0 citations
Abstract Motivation Inferring gene networks from single-cell RNA sequencing data is challenging due to high sparsity, dimensionality, and technical noise. Current pipelines lack the multi-dataset integration and comprehensive post-processing analysis. Results scGraphVerse is an R package that integrates multiple algorithms (GENIE3, GRNBoost2, ZILGM, PCzinb, and JRF) with extensive evaluation and visualization tools. Its modular workflow supports early, late, and joint integration strategies for multi-dataset analysis, providing standardized input/output interfaces and biological interpretation tools, including community detection, pathway enrichment, and literature mining. Benchmarking on simulated data showed model-based methods (PCzinb and ZILGM) perform well with limited sample sizes, while JRF performs best as the network size and dataset numbers increase. A PBMC case study demonstrates JRF’s ability to identify literature-supported regulatory communities across donors. Availability and implementation The package is available in Bioconductor 3.22 at https://bioconductor.org/packages/release/bioc/html/scGraphVerse.html. Code and examples: https://github.com/ngsFC/scGV_analysis.
Francesco Cecere, D. De Canditiis, Annamaria Carissimo et al.· Bioinformatics Advances· 0 citations
Bulk and single-cell RNA sequencing (scRNA-seq) have become essential for investigating disease mechanisms and identifying diagnostic biomarkers. However, the growing volume of transcriptomic data remains difficult to reuse efficiently for many researchers. Downstream analysis often requires multiple statistical, visualization, and reporting tools, creating fragmented workflows that reduce transparency and reproducibility, particularly when analyzing scRNA-seq data. To address this gap, we developed CoTRA (Comprehensive Toolbox for RNA-seq Analysis), an open-source R/Shiny package for bulk and scRNA analysis. CoTRA integrates established methods into modular workflows, exposes parameters, and offers alternatives at selected stages. It supports bulk RNA-seq quality assessment, differential expression, annotation, enrichment, and reporting, as well as scRNA quality control, dimensionality reduction, clustering, marker identification, cell-type annotation, differential abundance, trajectory inference, pathway activity, and cell-cell communication. CoTRA runs on workstations or HPC environments without mandatory external data submission and was tested on Linux, Windows, and macOS. Compared with 14 other platforms for bulk RNA-seq/scRNA-seq, CoTRA supported 46 of 49 predefined functionality criteria. Tool validation using published rd10 retinal bulk RNA-seq identified 1,947 shared differentially expressed genes with concordant direction and strong log2 fold-change agreement. A retinal scRNA-seq case study demonstrated appropriate clustering, cell-type resolved analysis, and pathway activity scoring. CoTRA provides a graphical environment for bulk and single-cell RNA-seq analysis while retaining parameter transparency, methodological flexibility, and reproducible outputs. Strong concordance with the published bulk RNA-seq analysis supports the workflow consistency, while the single-cell case study demonstrates its applicability to advanced scRNA-seq analysis. The source code is freely available at https://github.com/UmairSeemab/CoTRA.
Summary Single-sample transcriptomic analysis can provide gene-level views of how individual tumors deviate from matched reference states, complementing cohort-level differential expression. Here, we present HopRatio, a rank-based framework that quantifies, for each gene in each sample, the fraction of stable reference gene-ordering relationships inverted relative to a context-matched reference cohort. By relying on within-sample ranks rather than cross-sample expression magnitudes, HopRatio enables sample-specific scoring without cross-sample normalization for score calculation. Across 16 cancer types from The Cancer Genome Atlas with tissue-matched Genotype-Tissue Expression references, HopRatio generated individualized dysregulation profiles that supported tumor-normal discrimination using single genes and compact multi-gene panels. Recurrent high-performing features defined a 246-gene set enriched for developmental, membrane-associated, and ion-transport programs and associated with poor survival across cancers. Benchmarking in the Sequencing Quality Control dataset supported its robustness relative to commonly used differential expression methods, highlighting a scalable route for interpretable individualized transcriptomic analysis.
Yue Zhao, Bo Gao, Rui-Zhong Chen· iScience· 0 citations
Publicly available RNA sequencing (RNA-seq) data provide a cost-effective springboard for biomarker discovery. However, heterogeneity across studies often complicates analysis. This study presents a modular analytics pipeline that combines publicly available datasets with established open-source tools; standardizes quality control, differential expression analysis, and pathway analysis; and leverages competitive machine learning to unify disparate RNA-seq datasets for robust biomarker identification. The workflow is demonstrated across three disease contexts, ranging from small pilot datasets to larger, integrated analyses: (1) identifying differentially expressed gene signatures associated with COVID-19 severity, (2) combining differential expression and machine learning techniques to analyze multi-cohort sepsis datasets, resulting in concise biomarker panels, and (3) utilizing both bulk and single-cell data to examine tissue and cell-type specificity of N-acyl-phosphatidylethanolamine phospholipase D in atherosclerosis. These applications demonstrate how an adaptable, modular pipeline using open-source tools can repurpose public data to reduce noise, generate new hypotheses, and reveal meaningful biological insights, thereby establishing a foundation for future research and underscoring the importance of public data in exploratory biomarker discovery.
Cheryl L. Sesler, Lukasz S. Wylezinski, Guzel I. Shaginurova et al.· Journal of Molecular Diagnos...· 0 citations
This work introduces a two-stage statistical framework for interpretable patient-level disease classification from single-cell data, and recovered biologically coherent, cell-type specific gene signatures consistent with known disease mechanisms, demonstrating improved interpretability without sacrificing predictive accuracy.
Zhi-Wei Xiao, William Torous, Jeffrey B. Cheng et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.