Abstract Motivation Copy Number Variations (CNVs) play pivotal roles in complex disease etiology, often requiring large sample sizes to analyze disease associations. While genotyping arrays offer a cost-effective approach for CNV detection using Log R Ratio (LRR) and B Allele Frequency (BAF) signals, existing independent array-based callers suffer from high false positive rates and noise susceptibility, burdening manual validation. Results We present CNV-Finder, a deep learning pipeline employing Long Short-Term Memory (LSTM) networks for large-scale CNV identification within user-defined genomic regions. Trained on expert-annotated samples from the Global Parkinson’s Genetics Program across four neurodegenerative disease-associated genes (PRKN, LINGO2, MAPT, SNCA), CNV-Finder integrates human feedback to iteratively improve performance. In benchmarking across 105 936 samples spanning 11 ancestries and nearly 150 cohorts, the model achieved 91% and 89% visual confirmation rates for PRKN deletions and duplications at high-confidence thresholds. In two validation cohorts, CNV-Finder nominated 83% fewer candidates than a popular Hidden Markov Model-based caller while maintaining higher confirmation rates. Validation through MLPA, short-read, and long-read sequencing demonstrated robust performance, generalizing to diverse signatures including homozygous deletions and SNCA triplications absent from training. Our findings highlight human expertise’s value in complex loci like 17q21.31. Availability and implementation CNV-Finder is freely available at https://github.com/nvk23/CNV-Finder.
Nicole Kuznetsov, Kensuke Daida, M. Makarious et al.· Bioinformatics Advances· 0 citations
Despite major progress in genomic risk loci identification, biological mechanisms underlying Parkinson's disease (PD) remain incompletely understood and the informativity of polygenic score (PGS)-based prediction remains modest. Polytranscriptomic scores (PTS) - the sum of an individual's observed gene expression weighted by transcriptome-wide association z-scores - combine the stability of genetics with the dynamic biology of gene expression and have the potential to identify associations not captured by genetics alone. We present the first large-scale cross-trait and multi-PTS analysis of PD, calculating ~550 PTS for 100 phenotypes using whole-blood RNA-seq data from three independent clinical cohorts in the Accelerating Medicines Partnership Parkinson's Disease programme (AMP-PD) (N = 2,741; NCASES = 1,644). We identify 26 Bonferroni significant cross-trait PTS associations with PD (p<9x10-5) involving 18 phenotypes and 11 trait categories, including neurodegenerative diseases, respiratory function, sleep and cardiovascular traits. Only one of these associations was observed using corresponding PGS, highlighting the added value of integrating directly measured transcriptomic data. Combining multiple PTS within machine learning multi-PTS models improved prediction of PD case/control status beyond age, sex and PD-PGS in external validation, with a sparse model including an additional seven PTS achieving an AUC of 0.74 [0.70-0.78], representing a 0.09-point improvement. These findings reveal transcriptomic overlap between PD and a range of clinically relevant traits, providing novel insights into disease biology with potential to improve disease prediction.
L. Gilchrist, O. Pain, S. Calhas et al.· medRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.