Skip to content

Explainable AI for analyzing cancer outcomes using large-scale genome sequencing data

TL;DR

A multi-tier, explainable AI framework designed to risk-stratify patients and predict overall survival using clinical and genomic covariates is developed and demonstrates that explainable machine learning models can robustly predict survivability and highlight actionable features for oncology dashboards.

Abstract

Metastatic cancer remains a leading cause of global mortality, yet accurate prognosis is frequently hampered by high-dimensional molecular features and heterogeneous clinical presentations. While traditional staging systems and linear models provide a foundational risk assessment, they often fail to capture the complex, nonlinear interactions between metastatic topology, genomic burden, and functional sequence variation. To address this, recent advances in machine learning and genomic foundation models present a transformative opportunity to integrate diverse data types into an explainable predictive framework. Consequently, this research developed a multi-tier, explainable AI framework designed to risk-stratify patients and predict overall survival using clinical and genomic covariates. Additionally, the framework aimed to surface sequence-level disease drivers by implementing joint variant calling from RNA-seq data and leveraging transformer-based architectures. The study employed a two-track methodological approach encompassing populationscale modeling and sequence-level deep learning. For the population-scale aim, a retrospective analysis was conducted on the Memorial Sloan Kettering-Metastatic cohort, consisting of 25,775 patients. Five distinct classifiers XGBoost, Logistic Regression, Random Forest, Decision Tree, and Naive Bayes were trained on a balanced subset of 20,338 patients utilizing an 80/20 stratified split. Model explainability was established through Shapley Additive Explanations (SHAP), while survival dynamics were evaluated using Kaplan-Meier estimates, Cox proportional hazards models, and an XGBoost-Cox variant. Concurrently, a pilot study involving 60 individuals, comprising 30 breast cancer cases and 30 controls, investigated sequence-level drivers using RNAseq data. A joint variant calling pipeline generated a unified genomic variant call format for association testing, and three genomic foundation models DNABERT-2, HyenaDNA, and Nucleotide Transformer were fine-tuned for 50 epochs on variantcentered windows spanning 100 base pairs in either direction to classify case versus control status. The results revealed stark contrasts in performance between the clinical and genomic modeling tracks. In survivability predictions, XGBoost emerged as the superior classifier, achieving an accuracy of 0.74 and an AUC of 0.82, while the XGBoost-Cox model outperformed the traditional Cox model with a C-index of 0.70 compared to 0.66. Through explainability and hazard-based analyses, metastatic site count, tumor mutational burden, the fraction of the genome altered, and the presence of liver and bone metastases were identified as the most potent prognostic indicators across pan-cancer and cancer-specific models. Conversely, the sequence-level transformer models exhibited severe overfitting, with test performance remaining near stochastic levels between 49 percent and 51 percent accuracy. Although DNABERT-2 achieved the highest nominal accuracy at 50.63 percent and HyenaDNA showed superior computational efficiency, the pilot ultimately indicated that fine-tuning transformers on raw sequences in small cohorts is heavily limited by a high signal-to-noise ratio and the polygenic complexity of cancer. Ultimately, this research demonstrates that explainable machine learning models can robustly predict survivability and highlight actionable features for oncology dashboards. However, future sequence-level deep learning efforts must pivot toward using frozen transformer embEd. D.ings or larger, multi-center cohorts to ensure equitable and generalizable clinical adoption.

View source

Similar papers

Review Open access Aug 2026

Machine learning and AI for cancer research and care: a review of applications, limitations, and future directions

This review synthesizes key developments in ML for oncology, covering foundational algorithms alongside emerging approaches, and describes future directions, including federated learning, graph neural networks, longitudinal modeling, and integration of real-world and wearable data to support precision oncology.

Kanishk Yadav, Taneesha Gupta · 0 citations
Review Jul 2026

Abstract A025: Real-World Data–Driven Target Discovery in Advanced NSCLC

This work demonstrates the novel application and value of multimodal liquid-biopsy RWD for scalable, clinically anchored therapeutic Target ID in solid tumors and indicates a complementary discovery paradigm that accelerates hypothesis generation grounded in patient outcomes.

Aaron Hardin, Peili Zhang, Amar K. Das · 0 citations
Review Aug 2026

Artificial Intelligence-Driven Multiomics Integration in Lung Cancer: From Data Convergence to Precision Phenomics.

Precision phenomics is introduced as a unifying framework that links molecular, spatial, functional, and clinical characteristics of tumors to support personalized cancer management and has the potential to transform lung cancer research and improve patient outcomes.

Sanjukta Dasgupta, D. De · 0 citations
Preprint Aug 2026

A Multimodal Foundation Model for Longitudinal Patient Representation and Scalable Insight Generation in Oncology

The oFM is introduced, a foundation model developed on a real-world oncology cohort of 1.67 million cancer patients that integrates clinical trajectories with DNA, RNA, and H&E pathology and achieves a three-fold higher pooled and scale-normalized treatment-benefit AUTOC than baseline features.

E. Vorontsov, Yikan Wang, A. Bozkurt et al. · 0 citations
Jul 2026

Biomarkers for minimally invasive multigrade liver cancer diagnosis using integrative heterogeneous models and evolutionary algorithms.

Liver cancer, a leading global cause of death, requires early and accurate diagnosis for better treatment outcomes and reduced mortality rates. This retrospective cohort study presents a machine learning-based predictive framework utilizing a combined panel of minimally invasive circulating biomarkers for early diagnosis and multi-grade prediction of liver cancer. This curated dataset is composed of urine/blood biomarkers like β-hCG, PD-L1 and Alpha-Fetoprotein. A hybrid and heterogeneous ensemble learning-based prediction technique has been proposed for the thorough analysis of collected data about the subject and associated biomarkers from highly diverse and broad age groups. The model's heterogeneity involved applying multiple distinct learning algorithms to the same training dataset, with the outcomes of each classifier used to expand the feature space and guide decision-making in subsequent stages. The proposed ensemble model is characterized by an iterative methodology incorporating both bagging and boosting techniques, ensuring adaptability and sustained enhancement, thereby surmounting limitations inherent in conventional static ensembles. The best hyperparameters using evolutionary algorithm achieved an accuracy of 93.33% and an F1-score of 94.29% on the test samples. The developed model accurately predicts liver cancer risk and grade, outperforming conventional methods. This AI model minimizes invasive procedures and healthcare costs, enhancing overall public health. The proposed minimally invasive biomarker-based framework may be particularly useful in resource-limited clinical settings where advanced imaging infrastructure and invasive diagnostic procedures are not readily accessible.

R. Joshi, P. Srivastava, R. Mishra et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.