Aug 2026· Molecular Biomedicine· Vol 7· 0 citations· 32 references
Medicine
TL;DR
Fragmentia-AI™ WGS, a mutation-calling-independent framework that uses a transformer-based multiple-instance learning architecture with sequential fine-tuning across tumor fraction (TF) strata to extract latent cancer-associated signals from ULP-WGS data, enables robust cancer detection and clinically meaningful risk stratification from highly sparse cfDNA sequencing data.
Abstract
Ultra-low-pass whole-genome sequencing (ULP-WGS) of cell-free DNA (cfDNA) offers a cost-efficient strategy for cancer detection, but its clinical application is limited by extreme data sparsity and poor model generalization. We developed Fragmentia-AI™ WGS, a mutation-calling-independent framework that uses a transformer-based multiple-instance learning architecture with sequential fine-tuning across tumor fraction (TF) strata to extract latent cancer-associated signals from ULP-WGS data. Model performance was evaluated in multiple independent cohorts, including a pan-cancer test set covering 17 cancer types, an external public dataset generated on a different sequencing platform, and a technical variability cohort with heterogeneous pre-analytical and experimental conditions. Clinical relevance was assessed by correlating model predictions with progression-free survival (PFS) in patients with advanced non-small cell lung cancer receiving chemoimmunotherapy. Sequential fine-tuning across TF strata significantly improved performance in low-TF samples, achieving a 35.6% relative increase in AUC compared with high-TF-only training (0.884 vs. 0.652). In the independent test cohort, the model achieved an overall AUC of 0.930, with consistent performance across TF strata and cancer types. External validation confirmed robust cross-platform generalizability (AUC: 0.929; sensitivity: 0.78; specificity: 0.92). The model maintained stable classification performance despite score fluctuations associated with pre-analytical and technical variables. Importantly, model-negative status, defined as a prediction score below the training-derived cutoff, remained significantly associated with improved PFS compared with model-positive status (HR = 0.49, 95% CI: 0.29–0.82) after multivariable adjustment. Collectively, this framework enables robust cancer detection and clinically meaningful risk stratification from highly sparse cfDNA sequencing data.
Cancer type classification is challenging due to tumor heterogeneity and undefined tissue of origin (TOO), particularly in cancers of unknown primary (CUP) and multiple primary cancers (MPC). Accurate TOO identification is critical for guiding treatment and prognosis. We developed a stacked ensemble machine learning classifier that integrates 11 multidimensional cfDNA features spanning genomic, fragmentomic, methylation/repeat, and microbial signals. Base models were constructed using five algorithms, including Deep Learning, Distributed Random Forest, Gradient Boosting Machine, Generalized Linear Model, and XGBoost, within a five-fold cross-validation framework, and their predictions were aggregated into a final ensemble optimized for top-1 accuracy. The classifier achieved robust performance across 17 cancer types, with top-1 and top-2 accuracies of 78% and 89% in the training cohort (n = 1,814), and 80% and 90% in an independent validation cohort (n = 1,221). Notably, predictive performance was retained in samples with low tumor fraction (71% top-1, 85% top-2). Sensitivity varied across tumor types, with the highest performance observed in head and neck and colorectal cancers. Among CUP cases, 11 of 15 (73.3%) predictions matched clinically inferred primary sites based on multimodal diagnostics. Feature importance analysis identified nucleosome positioning, fragment size distribution, and repeat elements as key contributors to model performance. Collectively, this cfDNA-based classifier provides a robust and non-invasive approach for accurate cancer type identification and has the potential to support clinical decision-making.
Yunjian Zhang, Liang Liu, H. Bao et al.· Molecular Biomedicine· 0 citations
Abstract Motivation Accurate survival prediction is crucial for personalized cancer treatment but remains challenging for rare cancers due to limited data. Most deep learning models require large training datasets, which are unavailable for rare cancer types, creating a significant clini-cal bottleneck. Results We propose MoESurv, a zero-sample survival prediction framework that leverages a mix-ture-of-experts architecture to extract generalizable prognostic patterns from pan-cancer data. The model integrates shared experts, cancer-specific experts, and routing experts within an autoencoder to disentangle common and type-specific survival features. Evaluated on seven rare TCGA cancer types, MoESurv achieved state-of-the-art performance, improving the average C-index by 4 percentage points over the best baseline. Further-more, external validation across diverse populations and independent cohorts—including a Chinese glioma cohort (CGGA mRNAseq_693, C-index = 0.7433), a rare GBM IDH-mutant subtype (C-index = 0.8064), and the pan-cancer PCAWG cohort (C-index = 0.7090)—demonstrated that MoESurv possesses the most robust predictive per-formance, highlighting its generalizability. MoESurv also effectively stratified high- and low-risk patient groups and identified potential survival-associated genes, demonstrating both clinical utility and biological interpretability. Availability The code is freely available at https://github.com/HuaYC666/MoESurv and https://zenodo.org/records/20785500.
Per-dataset analysis reveals three reproducible regimes: probabilistic variational autoencoder variants help on the smallest datasets, deep autoencoders win on mid-scale data with multi-batch or many-type structure, and classical PCA pipelines remain competitive when linear projection already captures the dominant variation.
Phong T. Nguyen, T. Vu, T. Nguyen et al.· arXiv.org· 0 citations
Early detection of oral cancer improves outcomes, but remains limited by the invasiveness and low sensitivity of current screening methods.
This study included 342 participants (200 in the training and 142 in the validation cohorts). Plasma cfDNA underwent low-depth whole-genome sequencing, with data normalized to a standardized 5× coverage via down-sampling to ensure consistent analysis. Three cfDNA-derived features, including fragment size ratio (FSR), copy number variation (CNV), and repetitive element profiles (REP), were extracted and integrated using a stacked machine learning ensemble to generate prediction scores. Model performance was evaluated by area under the receiver operating characteristic curve (AUC), sensitivity, and specificity with 95% confidence intervals.
In the training cohort, oral cancer patients had a median age of 53 years (91% male), while healthy individuals had a median age of 57 years (37% male). In the validation cohort, the median ages were 53 and 57 years, with male proportions of 83.1% and 38.0% in the cancer and healthy groups, respectively. Cancer samples exhibited shorter cfDNA fragments, recurrent CNV gains at 3q and 8q, and losses at 3p and 5q. The individual AUCs for FSR, CNV, and REP were 0.973, 0.983, and 0.979 in the training cohort, and 0.960, 0.982, and 0.972 in the validation cohort, respectively. The integrated stacked model achieved AUCs of 0.996 (95% CI: 0.991-1.000) and 0.988 (95% CI: 0.975–0.988) in the training and validation cohorts, respectively. To prioritize high specificity for large-scale screening, a decision threshold was established in the training cohort to target 98.0% specificity, which yielded a corresponding sensitivity of 98.0%. When this pre-fixed threshold was blindly applied to the independent validation cohort, a sensitivity of 94.4% and a specificity of 95.8% were achieved. Prediction scores were not significantly associated with age, sex, or tumor site.
The cfDNA fragmentomics-based stacked model distinguished oral cancer from healthy individuals, including early stage disease, supporting its potential for noninvasive approach for oral cancer detection.
Weiwei Wang, Tao Ding, Cuicui Liu et al.· BMC Cancer· 0 citations
A multi-tier, explainable AI framework designed to risk-stratify patients and predict overall survival using clinical and genomic covariates is developed and demonstrates that explainable machine learning models can robustly predict survivability and highlight actionable features for oncology dashboards.
P. Nalela· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.