Skip to content

Leakage-controlled benchmarking of multi-omics patient-graph construction for pan-cancer tumor-type classification and prognosis analysis.

Sep 2026 · Computational biology and chemistry · Vol 126 Pt 1, pp. 109341 · 0 citations · 46 references
Medicine

Abstract

Pan-cancer multi-omics analysis requires models that integrate complementary molecular signals while preserving biologically meaningful relationships among patients. This study presents a leakage-controlled benchmarking framework for patient-graph learning in pan-cancer classification and prognosis analysis, focusing on how graph construction affects downstream performance. The benchmark explicitly separates fold-specific graph formation from downstream prediction. Using the TCGA Pan-Cancer cohort of 8204 primary tumors across 31 cancer types, RNA expression and copy-number variation data were used to compare early feature fusion, lightweight similarity network fusion (SNF-lite), and fused k-nearest neighbor similarity graphs under a common GATv2 encoder family with a matched attention-head search space and inner-validation selection procedure. A strict 5 × 3 nested cross-validation protocol ensured that imputation, gene selection, feature scaling, similarity computation, and neighbor search were fitted on training folds only. At G'=2000, graph-level fusion approaches achieved about 0.92 accuracy and 0.89 Macro-F1, outperforming early fusion at about 0.89 accuracy and 0.84 Macro-F1. Fused kNN graphs also showed higher neighborhood label purity than SNF-lite despite similar predictive performance. A weighted topology audit showed that local label agreement alone did not determine graph utility. Gene and omics ablations showed that RNA carried the dominant subtype-discriminative signal, while CNV and mutation contributed weaker but complementary information. A Cox auxiliary objective retained classification performance when used alone and enabled out-of-fold prognostic stratification. These findings show that patient-graph construction is a key design choice in pan-cancer multi-omics learning and that leakage-controlled evaluation is essential for reliable and biologically informative benchmarking in computational oncology.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.