Skip to content
Open access

P-KNN: joint calibration of multiple pathogenicity prediction tools streamlines variant classification

Aug 2026 · Genetics in Medicine · pp. 102692 - 102692 · 0 citations
Medicine

TL;DR

Pathogenicity K-Nearest Neighbors (P-KNN) provides robust joint calibration for any set of pathogenicity prediction tools, thereby alleviating the constraint of pre-committing to a single predictor while enhancing statistical rigor and diagnostic yield.

Abstract

Purpose: Clinical guidelines for interpreting genetic variants in the context of Mendelian disease require converting the outputs of pathogenicity prediction tools into well-calibrated probabilities. However, the existing calibration method is only valid when pre-committing to one tool, preventing clinical laboratories from using multiple tools with complementary strengths. To lift this restriction, we introduce Pathogenicity K-Nearest Neighbors (P-KNN), a flexible method that jointly calibrates any set of tools. Methods: P-KNN represents each variant in a multidimensional space defined by tool scores and estimates the probability of pathogenicity based on the proportion of pathogenic neighbors. We compared P-KNN against standard single-tool calibration of multiple predictors and meta-predictors at four historical time points. Results: P-KNN outperforms standard calibration of single tools and meta-predictors in two aspects: i) overall evidence strength and ii) alignment of the calibrated probabilities with true pathogenicity frequencies. Additionally, the evidence from P-KNN keeps improving with the addition of newer tools. It also correctly integrates correlated computational and experimental evidence that is overestimated by existing protocols. Conclusion: P-KNN provides robust joint calibration for any set of pathogenicity prediction tools, thereby alleviating the constraint of pre-committing to a single predictor while enhancing statistical rigor and diagnostic yield. P-KNN is available via command line (https://github.com/Brandes-Lab/P-KNN) and precomputed scores (https://huggingface.co/datasets/brandeslab/P-KNN).

Read PDF

Similar papers

#protein folding Open access Aug 2026

Gene identity, not variant effect, dominates ClinVar benchmarks of missense pathogenicity predictors

It is found that this test to judge whether a computer program predicts whether a genetic variant causes disease is easier to pass than it looks, and about half the advantage held by programs trained on clinical data disappears once the gene pattern is removed.

Saad Harrizi, Imane Nait Irahal, Kabine Mostafa et al. · 0 citations
Review Open access Aug 2026

A machine learning framework for predictive interpretation of variants of uncertain significance in hereditary cancer

This reproducible pipeline provides a clinically grounded computational approach to VUS triaging in precision oncology, with external validation supporting its generalizability to independent hereditary cancer gene datasets.

Nayeema Nizamuddin, Soham Biswas, Akshaykumar Zawar et al. · 0 citations
Open access Aug 2026

Machine Learning Triage of LDLR Variants of Uncertain Significance Using Predictor Concordance and ACMG-Aligned Evidence Mapping

Background: Variants of uncertain significance (VUS) in the LDLR gene remain a major barrier to the molecular diagnosis of familial hypercholesterolemia. Although computational approaches offer scalable prioritization, their clinical utility is limited by predictor discordance, incomplete annotations, and inflated perf...

BalaSubramani Gattu Linga, Faisal E. Ibrahim, Nader I. Al-Dewik · 0 citations
Open access Aug 2026

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting

A probabilistic gradient boosting model on variant pathogenicity prediction that applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making is presented.

Karthik V, S. Prejesh, Sumedh Deepak Kudale et al. · 0 citations
Open access Sep 2026

Markedly divergent performance of variant annotation methods for gene-level association testing

Machine learning-based annotation methods are increasingly used to assess the pathogenicity of genetic variants, but their performance at prioritizing variants for gene-level association testing remains poorly characterized. Here, to better understand and optimize for this use case, we assess variant annotations from...

Matthew Aguirre, Flaviyan Jerome Irudayanathan, M. Crow et al. · 0 citations
Open access Aug 2026

SIMLINK enables accurate variant pathogenicity prediction through modeling the gene-variant-feature association structure

Abstract Motivation Predicting variant pathogenicity is crucial for clinical genetics. Existing approaches face two primary limitations. First, biologically, data for pathogenicity prediction often lacks explicit modeling of the gene-variant-feature association structure. A single gene can harbor multiple variants, and...

Hong-Dong Li, Chen-Lu Wang, Dongfang Yan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.