This work developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals and demonstrated that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.
Abstract
Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.
The results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.
F. Mazzarotto, Özem Kalay, E. Arslan et al.· Genome Medicine· 0 citations
A strategy is presented that calculates the median and median absolute deviation of gene-level fold changes across all samples within each sequencing batch and incorporates these measures into the result interpretation, providing batch-level reference metrics and supporting more reliable interpretation in comprehensive genomic profiling.
Chung Lee, Sejoon Lee, Hyun-Hee Koh et al.· Journal of Molecular Diagnos...· 0 citations
RIDDLER stands out as a scalable, generalizable multi-modal method for accurate CNV detection, empowering studies aiming to link CNV dynamics to epigenetic alterations within the same cell, empowering studies aiming to link CNV dynamics to epigenetic alterations within the same cell.
Travis W Moore, H. Mohammed, Andrew C. Adey et al.· Nucleic Acids Research· 0 citations
Background. Pharmacogenetic (PGx) testing can guide drug prescribing but remains limited by the genomic assay used. Genotyping arrays are widely implemented yet limited to predefined variants, whereas low-pass whole-genome sequencing (LP-WGS) is not constrained by fixed probe design and may provide broader PGx variant availability after imputation. Methods. We compared Illumina Global Screening Array (GSA) v3 with ~1x LP-WGS for PGx profiling in 500 hospital biobank participants with electronic health record evidence of exposure to pharmacogenetically actionable drugs and reported adverse drug reactions. Concordance was evaluated genome-wide, at 20 actionable pharmacogenes for PharmCAT-derived star alleles and metabolizer phenotypes, and for HLA alleles. Results. Genome-wide concordance between imputed array and LP-WGS data was high (median 99.63%; interquartile range, 99.59%-99.64%). For pharmacogenetically relevant variants, LP-WGS captured a larger fraction, particularly rare alleles absent from the array data, whilst maintaining high concordance at shared sites. Predicted phenotype concordance exceeded 98% for most genes, although gene-specific differences in phenotype classification were observed. LP-WGS reduced missing phenotype assignments for selected loci, particularly CYP2C19 and NAT2, by improving resolution of star-allele structure. However, in structurally complex or incompletely characterized genes such as CYP2C9 and CYP2D6, broader variant recovery increased indeterminate classifications rather than consistently improving clinical interpretability. For HLA loci, concordance varied by imputation strategy, with SNP2HLA performing marginally better utilizing the GSA array compared to the LP-WGS approach. Conclusions. Overall, LP-WGS provides broader variant coverage and improved resolution for selected pharmacogenes but did not resolve all clinically important loci. These findings support further evaluation of LP-WGS as a scalable PGx screening approach, especially where long-term genomic data reuse is a priority.
F. Hodel, C. Thorball, D. Haefliger et al.· medRxiv· 0 citations
Abstract Motivation Copy Number Variations (CNVs) play pivotal roles in complex disease etiology, often requiring large sample sizes to analyze disease associations. While genotyping arrays offer a cost-effective approach for CNV detection using Log R Ratio (LRR) and B Allele Frequency (BAF) signals, existing independent array-based callers suffer from high false positive rates and noise susceptibility, burdening manual validation. Results We present CNV-Finder, a deep learning pipeline employing Long Short-Term Memory (LSTM) networks for large-scale CNV identification within user-defined genomic regions. Trained on expert-annotated samples from the Global Parkinson’s Genetics Program across four neurodegenerative disease-associated genes (PRKN, LINGO2, MAPT, SNCA), CNV-Finder integrates human feedback to iteratively improve performance. In benchmarking across 105 936 samples spanning 11 ancestries and nearly 150 cohorts, the model achieved 91% and 89% visual confirmation rates for PRKN deletions and duplications at high-confidence thresholds. In two validation cohorts, CNV-Finder nominated 83% fewer candidates than a popular Hidden Markov Model-based caller while maintaining higher confirmation rates. Validation through MLPA, short-read, and long-read sequencing demonstrated robust performance, generalizing to diverse signatures including homozygous deletions and SNCA triplications absent from training. Our findings highlight human expertise’s value in complex loci like 17q21.31. Availability and implementation CNV-Finder is freely available at https://github.com/nvk23/CNV-Finder.
Nicole Kuznetsov, Kensuke Daida, M. Makarious et al.· Bioinformatics Advances· 0 citations
Personalized pharmacotherapy requires systematic consideration of genetic factors influencing drug efficacy and safety. The accumulation of large-scale whole-exome sequencing (WES) resources provides an opportunity to assess population frequencies of clinically significant pharmacogenetic variants; however, the applicability of exome-based pharmacogenomics across populations, particularly those underrepresented in existing reference datasets, requires further evaluation.
A retrospective analysis of 6102 anonymized sequencing datasets obtained between 2020 and 2025 was performed using the DNBSEQ-G400 (MGI) platform and Agilent SureSelect Human All Exon v6/v7/v8 enrichment kits. SNV and indel detection, CNV analysis, high-resolution
HLA
typing, and diplotype assignment for key pharmacogenes were conducted. Pharmacogenomic annotations were derived from ClinPGx (formerly PharmGKB) (levels of evidence 1 A–2B), CPIC, and PharmVar. Haplotype phasing (SHAPEIT5), genotype imputation (IMPUTE5), and phased linkage disequilibrium analysis were performed using a reference panel of 814 Russian whole-genome sequences generated with the same bioinformatic workflow to assess the feasibility of reconstructing clinically relevant non-coding pharmacogenetic variants not captured by WES.
WES reliably detected 33 of 34 Very Important Pharmacogenes (VIPs), allowing determination of allele frequencies, metabolizer statuses for 13 VIPs, and
HLA
diversity. The highest allelic and phenotypic variability was observed in
CYP2D6
,
CYP2C19
, and
CYP2B6
. A total of 663 ClinPGx annotations were identified, predominantly related to drug metabolism (50.38%) and toxicity (29.56%), involving psychotropic drugs, anticoagulants, statins, opioid analgesics, antineoplastic agents, and immunosuppressants. However, at least 32 drugs require assessment of non-coding pharmacogenetic variants or accurate
CYP2D6
copy number determination – both beyond standard WES capability. Phased haplotype analysis demonstrated that haplotype-based imputation was feasible for only a limited subset of clinically relevant non-coding pharmacogenetic variants, whereas the majority lacked informative exonic proxies required for accurate inference from WES data.
These findings represent one of the largest reference resource to date of pharmacogenetically significant variant and
HLA
allele frequencies in the Russian population. The results confirm WES as a robust and scalable platform for population-level pharmacogenomic screening and broader clinical genomic applications, while demonstrating that haplotype-based imputation cannot fully overcome the intrinsic limitations of exome sequencing for comprehensive pharmacogenomic profiling. Selection of genomic testing strategies should balance diagnostic completeness, analytical complexity, cost, scalability, and long-term potential for reanalysis.
A. Buianova, V. V. Cheranev, Anna O. Shmitko et al.· Human Genomics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.