Skip to content
Open access

Identifying multi-omics biomarkers for ovarian cancer survival estimation

Aug 2026 · bioRxiv · 0 citations · 70 references
Biology

TL;DR

An interpretable three-stage machine learning framework integrating mRNA, microRNA, DNA methylation, copy number variation, and protein expression data from The Cancer Genome Atlas that couples improved prognostic estimation with biological interpretability supporting multi-omics biomarker discovery in ovarian cancer.

Abstract

Ovarian cancer is among the deadliest gynecologic malignancies, and its molecular heterogeneity limits accurate prognostic stratification. Although multi-omics approaches have improved predictive modeling, many prioritize predictive performance over biological interpretability, limiting their clinical translation. We developed an interpretable three-stage machine learning framework integrating mRNA, microRNA, DNA methylation, copy number variation, and protein expression data from The Cancer Genome Atlas. Hierarchical feature selection was combined with a weighted ensemble of ElasticNet, ridge regression, support vector regression, XGBoost, and random forest models to estimate overall survival time in patients with ovarian cancer. Multi-omics integration outperformed every single-modality model, achieving a Pearson correlation of 0.752, a concordance index of 0.779, and a mean absolute error of 8.57 months between estimated and observed survival time, compared with 0.48 for the best single modality. The framework identified a 20-biomarker signature dominated by tumor-associated macrophage and complement genes. In an independent survival analysis, VSIG4 and CD163 remained significant after false discovery rate correction, and the signature raised the concordance index over clinical covariates alone from 0.615 to 0.686Enrichment analysis implicated PI3K-Akt, MAPK, focal adhesion, hypoxia, apoptosis, and p53 signaling pathways. This framework couples improved prognostic estimation with biological interpretability supporting multi-omics biomarker discovery in ovarian cancer.

Read PDF

Similar papers

Open access Aug 2026

A heterogeneous multimodal ensemble framework for multi-omics breast cancer prognosis

Accurate breast cancer prognosis remains a major challenge in precision oncology due to tumor heterogeneity and the complexity of integrating high-dimensional multi-omics data. Although multimodal learning approaches have improved predictive performance by combining clinical and molecular information, many existing methods rely on a single ensemble strategy that remains susceptible to prediction variance and limited robustness in high-dimensional, low-sample-size biomedical datasets. This study investigated whether integrating complementary ensemble strategies within a unified multimodal framework could improve the robustness and predictive performance of breast cancer prognosis. A heterogeneous multimodal ensemble framework was developed in which stacking was used to integrate complementary information from clinical, gene expression, and copy number variation (CNV) data through meta-learning, while bagging was incorporated to stabilize the meta-learning process via bootstrap aggregation. The outputs of the stacking and bagging branches were combined using weighted probability fusion. The framework was evaluated on the METABRIC breast cancer cohort and compared with unimodal models and a conventional stacking ensemble using an independent test set and stratified tenfold cross-validation. The proposed hybrid framework achieved a ROC-AUC of 0.936, outperforming unimodal clinical and molecular models (ROC-AUC = 0.8140.885) and the conventional stacking ensemble (ROC-AUC = 0.898). Stratified tenfold cross-validation further demonstrated consistent improvements in mean ROC-AUC, recall, F1-score, balanced accuracy, and Matthews correlation coefficient, indicating improved robustness and stable performance across the internal validation folds. On the independent test set, the hybrid framework reduced false-negative predictions and increased sensitivity relative to the stacking ensemble, demonstrating a more favorable balance between identifying high-risk patients and maintaining overall predictive performance. Rather than introducing a new ensemble algorithm, this study demonstrates that assigning complementary roles to stacking multimodal information integration and bagging for prediction stabilization provides an effective and robust framework for multi-omics breast cancer prognosis. The proposed hybrid strategy consistently improved predictive performance and robustness compared with conventional stacking while demonstrating stable performance across internal validation, supporting the use of complementary ensemble paradigms for multimodal prediction in precision oncology.

Reza Bozorgpour, Mohammadreza Soltany Sadrabadi · 0 citations
Aug 2026

Machine learning-enabled multi-omics discovery of prognostic biomarkers and signaling targets in pancreatic cancer.

Pancreatic ductal adenocarcinoma (PDAC) remains difficult to subtype using single omics layers. We conducted an exploratory investigation integrating reverse-phase protein array (RPPA) and DNA methylation data from the cancer genome atlas (TCGA)- pancreatic adenocarcinoma (PAAD) to assess the feasibility of multi-omics subtyping, alongside a supervised machine learning analysis of a small gene expression omnibus (GEO) transcriptomic cohort (n = 26) to identify candidate diagnostic genes. RPPA-based K-means clustering suggested a weak, possible two-subtype structure (silhouette ≈ 0.16) that remained unassociated with overall survival (log-rank p = 0.113) and lacked independent prognostic value. An independently performed similarity network fusion (SNF) analysis integrating RPPA and methylation data showed low concordance with RPPA-derived subtypes (Adjusted Rand Index (ARI) = 0.014), indicating limited convergence between molecular modalities. Supervised machine learning analysis of the GEO cohort using a fully nested leave-one-out cross-validation pipeline achieved a mean (area under the curve) AUC of 0.896 across four classifiers and identified four-fold-stable candidate genes (ESCO2, COL17A1, BCL2L14, and SOWAHB). However, this gene panel demonstrated limited external validity across two independent PDAC cohorts (log-rank p = 0.438 for both GSE62452 and GSE28735), indicating limited generalizability despite robust internal performance. Collectively, these findings provide limited evidence for a robust, prognostically significant multi-omics subtype or a validated diagnostic gene signature; instead, this study serves as a hypothesis-generating resource and highlights the importance of rigorous cross-validation and independent external validation in small-sample transcriptomic biomarker discovery.

Shafiul Haque, D. Mathkor, M. Wahid et al. · 0 citations
Open access Jul 2026

Multi-omics fusion with machine learning enables robust prediction of treatment response in ovarian cancer for precision population health.

Inter-patient heterogeneity complicates predicting treatment response in ovarian cancer (OC). We developed OMICS-FUSE, an early-fusion multi-omics predictive model integrating proteomic, transcriptomic, and methylomic data from OC patients, evaluated across five machine learning algorithms with SHapley Additive exPlanations (SHAP) and experimental validation. The early-fusion Random Forest model achieved excellent predictive accuracy (AUC = 0.939, accuracy = 0.896, F1 = 0.939), with performance comparable to or surpassing that of the best-performing single-omics models. Nevertheless, the multi-omics framework yielded superior balance across accuracy and F1 score. SHAP analysis identified key determinants of treatment response, including CLEC2A, MYH4, and methylation of SYT12_1, with functional enrichment implicating immune regulation, metabolic pathways, and drug resistance signaling. Experimental validation confirmed six hub genes (CASP8, AQP8, CAV1, FN1, CREB1, KDR), exhibiting expression patterns associated with drug resistance, immune regulation, and prognosis. This multi-omics machine learning model enables robust, interpretable prediction, uncovering molecular signatures for therapeutic stratification and precision oncology in OC.

Jie Chen, Tianshi Mao, Yu Yang et al. · 0 citations
Open access Aug 2026

Machine learning-based identification of hub genes and prognostic biomarkers in prostate cancer

Background Prostate cancer (PCa) is the second most prevalent malignancy in men worldwide, and accurate stratification of biochemical recurrence (BCR) risk remains challenging using conventional clinicopathological parameters alone. Identification of robust molecular biomarkers and integrated prognostic models is therefore of high clinical priority. Methods RNA-seq count data and clinical annotations for 554 TCGA-PRAD samples were obtained and normalized to log2(CPM+1). Weighted gene co-expression network analysis (WGCNA) identified co-expression modules correlated with Gleason score, PSA, and pathologic T stage. Protein-protein interaction (PPI) network analysis with CytoHubba topological scoring defined consensus hub genes. Four machine learning algorithms - LASSO Cox regression, random forest, SVM, and XGBoost - were applied to construct and validate a prognostic risk model. Immune cell infiltration was quantified and a prognostic nomogram was constructed and evaluated by decision curve analysis. Hub gene expression was experimentally validated by qRT-PCR and ELISA in prostate cancer and normal prostatic epithelial cell lines. Results Five hub genes - EZH2, CDK1, AURKA, TOP2A, and CCNB1 - were identified within the turquoise WGCNA module, which showed the strongest correlations with Gleason score (r = 0.78), PSA (r = 0.68), and pathologic T stage (r = 0.62). LASSO Cox regression and random forest consensus selected EZH2, CDK1, and AURKA for a three-gene risk score (Risk Score = 0.312xEZH2 + 0.285xCDK1 + 0.241xAURKA). High-risk patients demonstrated markedly inferior BCR-free survival (HR = 3.21, 95% CI: 2.05–5.03; log-rank P < 0.0001), with time-dependent AUCs of 0.821, 0.842, and 0.836 at 1, 3, and 5 years, respectively. Multivariate Cox regression confirmed the risk score as an independent prognostic factor (HR = 2.87; P < 0.001). A nomogram integrating the risk score with clinical parameters showed superior net benefit by decision curve analysis. Hub-high tumors exhibited reduced CD8+ T cell infiltration, elevated M2 macrophage abundance, and upregulated immune checkpoints (PD-L1, CTLA4, TIM-3, LAG3). All hub genes were confirmed overexpressed at both mRNA and protein levels in PCa cell lines by qRT-PCR and ELISA. Conclusion EZH2, CDK1, and AURKA constitute an internally validated prognostic risk signature in PCa that links cell cycle dysregulation to an immunosuppressive tumor microenvironment. This signature provides clinically actionable risk stratification and highlights candidate therapeutic targets in prostate cancer.

Gu-Quan Chen, Jie-Feng Zhang, Lin-Fu Zhao et al. · 0 citations
Open access Aug 2026

Integrated Transcriptomic Analyses Identify Four Prognosis-Associated Genes in Hepatocellular Carcinoma

Hepatocellular carcinoma (HCC) is one of the malignant tumors with high incidence and mortality rates worldwide. Given the poor prognosis of patients with HCC, it is crucial to explore the molecular mechanisms underlying HCC development and to evaluate prognostic markers. Differential expression analysis followed by univariate Cox, LASSO, and multivariate Cox regression identified four genes (EPO, SOCS2, IL18RAP, and KPNA2), and a Cox-based risk score was evaluated in the TCGA-LIHC cohort and externally in GSE14520 using Kaplan–Meier and time-dependent ROC analyses. Bulk, single-cell, and protein resources provided convergent expression context. Survival machine-learning analysis using observed overall-survival time and censoring status identified Cox–Ridge as the best-performing model in TCGA-LIHC, with more modest performance in GSE14520, and immune profiling revealed risk-group-associated differences in estimated immune and stromal components, immune-cell composition, and immune-checkpoint expression. The oncoPredict/GDSC2 screen highlighted five potential drug candidates for experimental prioritization. Because the drug screen is based on computationally predicted sensitivities, these findings should be regarded as hypothesis-generating and require validation in prospective cohorts and experimental systems before clinical translation.

Yu-Xian Liu, Xing-Jie Chen, Junyuan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.