Skip to content
Open access

Machine Learning approaches for the detection of disease-causing variants in whole-genome data need to address the expression of functional genes

Aug 2026 · PLoS ONE · Vol 21, pp. e0355557 · 0 citations · 120 references
Medicine

TL;DR

A benchmark is constructed that includes real data and synthetic data with known generating mechanisms and various dataset sizes and levels of noise, and shows that it is necessary to take into account the expression of functional genes in order to successfully predict disease.

Abstract

Gene-dosage combinations have been recognised as leading factors of disease. Given that those combinations may include dozens of genes, it is hypothesised that machine learning (ML) approaches may be useful in the classification of cases and controls and the identification of causative genes. We aimed to assess the validity of this hypothesis. Here, we have constructed a benchmark that includes real data (with ground truth knowledge) and synthetic data with known generating mechanisms and various dataset sizes and levels of noise. We trained standard statistical learning/ ML models on these datasets to classify disease phenotype. We present an analysis of how model performance varies across different synthetic genetic scenarios, and how it is impacted by dataset size. The logistic regression model was found to be the most reliable at causative gene identification across the synthetic datasets, despite not always performing the best in terms of classification performance and, in some cases, having a relatively low ROC AUC score. When our training attempts on the UK Biobank datasets failed, we performed an analysis into model performance vs dataset richness. Our results show that it is necessary to take into account the expression of functional genes in order to successfully predict disease.

Read PDF

Similar papers

Open access Aug 2026

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

Karthik V, S. Prejesh, Sumedh Deepak Kudale et al. · 0 citations
Review Open access Jul 2026

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review

This review describes the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS and presents 30 methods designed to leverage AI in GWAS.

S. D’Antona, Mawada Elmagboul Abdalla Abakar, Daniele Ramazzotti et al. · 0 citations
Review Open access Aug 2026

A machine learning framework for predictive interpretation of variants of uncertain significance in hereditary cancer

Introduction Variant interpretation remains a major bottleneck in clinical genomics, with variants of uncertain significance (VUS) representing a critical unresolved challenge due to insufficient evidence for definitive classification. Existing in silico tools exhibit variable and often inconsistent performance complicating clinical decision-making, particularly in the context of hereditary cancer genomics. Methods In this study, we developed a machine learning framework trained on 1,04,646 high-confidence ClinVar germline variants (3-star+ review status) annotated with Ensembl VEP (v114, GRCh38) and CADD v1.6 pathogenicity scores to classify variants as Pathogenic or Benign, subsequently applying the trained model to reclassify 40894 ClinVar VUS. Train/test partitioning was performed at the variant level (80/20 split) to prevent data leakage, with hyperparameter optimization via GridSearchCV and performance assessed by 10-fold cross-validation. Four classifiers were evaluated viz. Logistic Regression, Support Vector Machine, Random Forest and XGBoost, with Random Forest achieving the highest performance (AUC-ROC = 0.9995, 95% CI: 0.9993–0.9997; 10-fold CV AUC = 0.9992 ± 0.0004). Probability thresholds of P ≥ 0.80 (Pathogenic) and P <= 0.20 (Benign) were derived from Precision-Recall curve analysis, achieving empirically validated precision of 99.63% and 99.77% respectively on held-out test variants. Results and Discussion Applied to 40,894 ClinVar VUS, the model reclassified 19393 (47.4%) as Likely Pathogenic and 8,957 (21.9%) as Likely Benign, while 12,544 (30.7%) were conservatively retained as uncertain. External validation on 7,462 ENIGMA-classified BRCA1/BRCA2 variants from the BRCA Exchange database, completely independent of the ClinVar training data demonstrated an overall concordance of 98.83% (AUC = 1.0000). Further validation of VUS reclassification against 671 variants classified as VUS in ClinVar but definitively classified by ENIGMA yielded an overall concordance of 89.57% (Pathogenic: 96.4%, Benign: 87.4%). SHAP-based explainability analysis confirmed that predictions were predominantly driven by biologically interpretable features, including CADD Phred score, VEP functional impact tier, variant consequence class and population allele frequency, consistent with ACMG/AMP evidence criteria. This reproducible pipeline provides a clinically grounded computational approach to VUS triaging in precision oncology, with external validation supporting its generalizability to independent hereditary cancer gene datasets.

Nayeema Nizamuddin, Soham Biswas, Akshaykumar Zawar et al. · 0 citations
#explainable ai Open access Aug 2026

aiDIVA – hybrid AI for rare disease diagnostics using evidence-based, machine learning and language models

aiDIVA is presented, an ensemble-AI combining statistical and machine learning models trained on genomic and phenotypic data to identify causal variants among tens of thousands per patient, and applies a random forest model to classify pathogenicity and generates evidence-based scores for dominant and recessive diseases.

D. Boceck, L. Laugwitz, Marc Sturm et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.