Aug 2026· Genetics in Medicine· pp. 102692 - 102692· 0 citations
Medicine
TL;DR
Pathogenicity K-Nearest Neighbors (P-KNN) provides robust joint calibration for any set of pathogenicity prediction tools, thereby alleviating the constraint of pre-committing to a single predictor while enhancing statistical rigor and diagnostic yield.
Abstract
Purpose: Clinical guidelines for interpreting genetic variants in the context of Mendelian disease require converting the outputs of pathogenicity prediction tools into well-calibrated probabilities. However, the existing calibration method is only valid when pre-committing to one tool, preventing clinical laboratories from using multiple tools with complementary strengths. To lift this restriction, we introduce Pathogenicity K-Nearest Neighbors (P-KNN), a flexible method that jointly calibrates any set of tools. Methods: P-KNN represents each variant in a multidimensional space defined by tool scores and estimates the probability of pathogenicity based on the proportion of pathogenic neighbors. We compared P-KNN against standard single-tool calibration of multiple predictors and meta-predictors at four historical time points. Results: P-KNN outperforms standard calibration of single tools and meta-predictors in two aspects: i) overall evidence strength and ii) alignment of the calibrated probabilities with true pathogenicity frequencies. Additionally, the evidence from P-KNN keeps improving with the addition of newer tools. It also correctly integrates correlated computational and experimental evidence that is overestimated by existing protocols. Conclusion: P-KNN provides robust joint calibration for any set of pathogenicity prediction tools, thereby alleviating the constraint of pre-committing to a single predictor while enhancing statistical rigor and diagnostic yield. P-KNN is available via command line (https://github.com/Brandes-Lab/P-KNN) and precomputed scores (https://huggingface.co/datasets/brandeslab/P-KNN).
It is found that this test to judge whether a computer program predicts whether a genetic variant causes disease is easier to pass than it looks, and about half the advantage held by programs trained on clinical data disappears once the gene pattern is removed.
This reproducible pipeline provides a clinically grounded computational approach to VUS triaging in precision oncology, with external validation supporting its generalizability to independent hereditary cancer gene datasets.
Nayeema Nizamuddin, Soham Biswas, Akshaykumar Zawar et al.· Frontiers in Systems Biology· 0 citations
Background: Variants of uncertain significance (VUS) in the LDLR gene remain a major barrier to the molecular diagnosis of familial hypercholesterolemia. Although computational approaches offer scalable prioritization, their clinical utility is limited by predictor discordance, incomplete annotations, and inflated perf...
BalaSubramani Gattu Linga, Faisal E. Ibrahim, Nader I. Al-Dewik· Genes· 0 citations
A probabilistic gradient boosting model on variant pathogenicity prediction that applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making is presented.
Karthik V, S. Prejesh, Sumedh Deepak Kudale et al.· Frontiers in Digital Health· 0 citations
Machine learning-based annotation methods are increasingly used to assess the pathogenicity of genetic variants, but their performance at prioritizing variants for gene-level association testing remains poorly characterized. Here, to better understand and optimize for this use case, we assess variant annotations from...
Matthew Aguirre, Flaviyan Jerome Irudayanathan, M. Crow et al.· BMC Genomics· 0 citations
Abstract Motivation Predicting variant pathogenicity is crucial for clinical genetics. Existing approaches face two primary limitations. First, biologically, data for pathogenicity prediction often lacks explicit modeling of the gene-variant-feature association structure. A single gene can harbor multiple variants, and...
Hong-Dong Li, Chen-Lu Wang, Dongfang Yan et al.· Bioinformatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.