Skip to content
Open access

Pan-cancer benchmarking reveals complementary copy number signatures with distinct multi-omic predictability

Jul 2026 · bioRxiv · 0 citations
Biology

TL;DR

A matched-sample pan-cancer benchmark of three major copy number signature compendia is established, showing that compendium choice is an analytical decision rather than an interchangeable preprocessing step and provides a reproducible framework for selecting and interpreting copy number signature representations in pan-cancer studies.

Read PDF

Similar papers

Open access Aug 2026

HopRatio: Profiling single-sample transcriptomic dysregulation using stable gene-ordering relationships

Summary Single-sample transcriptomic analysis can provide gene-level views of how individual tumors deviate from matched reference states, complementing cohort-level differential expression. Here, we present HopRatio, a rank-based framework that quantifies, for each gene in each sample, the fraction of stable reference gene-ordering relationships inverted relative to a context-matched reference cohort. By relying on within-sample ranks rather than cross-sample expression magnitudes, HopRatio enables sample-specific scoring without cross-sample normalization for score calculation. Across 16 cancer types from The Cancer Genome Atlas with tissue-matched Genotype-Tissue Expression references, HopRatio generated individualized dysregulation profiles that supported tumor-normal discrimination using single genes and compact multi-gene panels. Recurrent high-performing features defined a 246-gene set enriched for developmental, membrane-associated, and ion-transport programs and associated with poor survival across cancers. Benchmarking in the Sequencing Quality Control dataset supported its robustness relative to commonly used differential expression methods, highlighting a scalable route for interpretable individualized transcriptomic analysis.

Yue Zhao, Bo Gao, Rui-Zhong Chen · 0 citations
Preprint Aug 2026

REDE: A Quantitative Framework for Differential-Expression Reproducibility and Diagnostic Transfer Across Nine Cohorts and Three Cancers

Differential-expression analyses often turn cohort-specific significance into claims of stable gene signatures or diagnostic biomarkers. We evaluated which layers of evidence reproduce across independent datasets and whether discovery-derived panels retain locked tumor-versus-non-tumor classification performance. Nine public microarray cohorts were organized into fixed discovery, validation, and external-test experiments for pancreatic ductal adenocarcinoma, breast cancer, and lung cancer. Reproducibility was assessed for DEG burden, exact membership, top-rank overlap, signed effects, prespecified gene confirmation, and Hallmark pathways. We also introduced REDE-2Fold, in which each discovery cohort is split once at patient level, differential expression is performed independently in both folds, and only same-direction genes selected in both are retained. Four training-only panels were then evaluated with locked logistic models and thresholds: all discovery DEGs, the top 19 discovery DEGs, all REDE-2Fold genes, and the top 19 REDE-2Fold genes. Broad DEG-list confirmation in both independent cohorts ranged from 15.5% to 39.5%, rising to 50.1% to 84.3% for large effects. Pathway replication ranged from 52.2% to 88.9%. Compact 19-gene panels retained high external ROC-AUC, but locked operating points were often unstable: some models with ROC-AUC near 1.0 showed zero specificity or very low sensitivity. These results define a hierarchy from thresholded membership through effect, pathway, discrimination, and operating-point transfer. REDE provides a seven-level quantitative framework for matching transcriptomic claims to the evidence actually tested, while REDE-2Fold offers a minimum internal feature-stability procedure that strengthens but does not replace independent validation.

Athanasios Angelakis · 0 citations
Open access Aug 2026

Identifying multi-omics biomarkers for ovarian cancer survival estimation

An interpretable three-stage machine learning framework integrating mRNA, microRNA, DNA methylation, copy number variation, and protein expression data from The Cancer Genome Atlas that couples improved prognostic estimation with biological interpretability supporting multi-omics biomarker discovery in ovarian cancer.

Kosar Fateh, S. Sathipati · 0 citations
Open access Aug 2026

Prognostic stratification by LGR5 expression identifies surface-accessible, structurally ligandable and condensate-forming targets in colorectal cancer

Background LGR5 marks colorectal cancer stem cells and is associated with poor outcome, but its expression on normal intestinal stem cells has constrained direct therapeutic targeting, and the molecular landscape of LGR5-high tumors remains incompletely defined. A transcriptional signature is not itself a set of drug targets: its constituent genes differ in whether and how they can be engaged pharmacologically, a distinction rarely applied systematically to a tumor-defined gene set. Methods We stratified 396 colorectal tumors from The Cancer Genome Atlas by LGR5 expression and compared transcriptional, somatic mutation, and copy number profiles between LGR5-high and LGR5-low groups using non-parametric testing with combined significance and effect-size thresholds. Genome-wide CRISPR knockout data were interrogated to test genetic dependency. Each signature gene was then triaged by pharmacological tractability rather than essentiality, along three axes: surface accessibility, from surfaceome annotation and membrane topology; cavity ligandability, from pocket detection on predicted structures using three independent algorithms; and condensate propensity, from saturation concentration prediction and coarse-grained molecular dynamics simulation. Results LGR5-high tumors displayed a coordinated program spanning Wnt signaling, stemness, and matrix remodeling, arising on an APC-mutant background with co-occurring IGF2 amplification. No constituent gene scored as a selective dependency. The three axes partitioned the signature with minimal overlap and nominated three candidates engaged by orthogonal modalities: ENPP3, a single-pass ectoenzyme presenting an accessible ectodomain and carrying clinical antibody-drug conjugate precedent; PLCB4, combining a well-defined catalytic pocket with additional predicted ligandable sites; and NKD1, accessible by neither route but undergoing RNA-stabilized homotypic phase separation, unlike SATB1 and MEX3A. Simulations further indicated that NKD1 partitions into DVL2-containing condensates and reduces DVL2-Wnt contacts, suggesting a biophysical basis for its negative-feedback role. Conclusions LGR5 expression defines a colorectal cancer subset that is pharmacologically tractable despite the absence of genetic dependency. Triaging by modality rather than essentiality converts descriptive tumor signatures into stratified, experimentally testable therapeutic hypotheses, including condensate-directed modulation of NKD1 as a route to targets inaccessible by antibody- or pocket-based approaches.

Lucía Paniagua-Herranz, A. Feito, Cristian Privat et al. · 0 citations
Aug 2026

A topology-based framework for robust cancer-associated gene signature identification from scRNA-seq data.

Cancer transcriptomics faces a fundamental challenge: conventional gene selection methods capture statistical variance but fail to decode the intrinsic geometric architecture of high-dimensional single-cell RNA sequencing data, leaving critical cancer-specific signals obscured by noise and biological heterogeneity. We present a topology-guided framework that harnesses persistent homology to extract structurally invariant, cancer-associated gene signatures from scRNA-seq data - moving beyond gene-level statistics toward shape-aware biological discovery. The framework integrates highly variable feature selection and dimensionality reduction with Vietoris-Rips filtration-based gene correlation topology, followed by a dual-stage stability-driven classification strategy that identifies samples exhibiting reproducible cancer-specific topological patterns. Topologically significant genes are rigorously validated through differential expression analysis, ROC/AUC evaluation, KEGG pathway enrichment, protein-protein interaction network analysis, and literature evidence. Against conventional HVF+PCA-based selection, the TDA framework delivers markedly superior discriminative power, substantially higher literature-supported biological relevance, and dramatically more focused cancer-specific pathway enrichment - while converging to compact, functionally coherent gene sets that conventional approaches cannot achieve. In breast cancer, the framework reveals a dominant mitotic regulatory module centered on cell cycle dysregulation, while colorectal cancer is characterized by extracellular matrix remodeling and tumor microenvironment-driven mechanisms - demonstrating cancer-type-specific biological fidelity. Critically, the framework identifies computationally prioritized novel candidate biomarkers absent from standard pathway databases yet exhibiting topological and statistical significance. This work establishes persistent homology as a transformative paradigm for transcriptomic biomarker discovery, offering a principled, structure-aware foundation for precision oncology and next-generation cancer diagnostics.

Sudarshana Gogoi, S. Bandyopadhyay, S. Bera et al. · 0 citations
Open access Aug 2026

Metagenomic profiling of blood-associated microbial DNA signatures in leukemia-associated febrile neutropenia

Febrile neutropenia (FN) is a life-threatening complication of chemotherapy, but the low microbial biomass of blood makes shotgun metagenomic profiles highly sensitive to technical background. We reanalyzed 47 publicly available patient sequencing runs representing 43 unique patient-timepoint samples from 19 SRA-labeled patients, together with 23 no-template-control (NTC) runs spanning 21 sequencing batches. To distinguish reference-catalogue content from progressively stronger evidence of patient-associated signal, we applied batch-matched NTC correction together with nested abundance thresholds and a feature-specific global NTC envelope. CheckM2 evaluated 1,013 bins; 13 met completeness ≥50% and contamination <10%, and dereplication yielded 11 draft MAG representatives. Ten representatives showed positive patient-to-control abundance excess, but only four showed recurrent support above both threefold matched-control abundance and the global NTC envelope. Functional annotations were therefore interpreted as reference-genome homologs rather than evidence of expression, phenotype, viability or bloodstream origin. Matched-control correction retained 19 read-level ARG types, but only seven subjects contributed complete longitudinal ARG-profile contrasts, limiting reliable temporal inference. The resulting run-resolved, nested evidence framework identified a subset of microbial DNA and ARG signals that remained detectable under increasingly stringent control criteria while distinguishing them from catalogue-level or background-sensitive signals. These findings support cautious reporting of patient-enriched microbial DNA and ARG signals rather than inference of a resident blood microbiome or clinical resistance phenotype.

Ming Zhang, Lan He, Huaping Xin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.