Skip to content
Open access

Robust and generalizable CNV detection for single-cell sequencing assays

Jul 2026 · Nucleic Acids Research · Vol 54 · 0 citations · 81 references
Medicine

TL;DR

RIDDLER stands out as a scalable, generalizable multi-modal method for accurate CNV detection, empowering studies aiming to link CNV dynamics to epigenetic alterations within the same cell, empowering studies aiming to link CNV dynamics to epigenetic alterations within the same cell.

Abstract

Abstract Copy number variations (CNVs) are genomic structural variants that are strongly linked to cancer progression and genetic disorders. CNVs can be highly heterogeneous at population and tissue scale; thus, single-cell resolution detection holds great promise for studying clonal evolution and CNV-driven changes. Despite advanced sc-RNA-seq CNV detection methods, accurate methods for epigenomic single-cell modalities lag behind. We developed RIDDLER; a robust, unsupervised method that uses outlier-aware statistical modeling to detect CNVs across multiple single-cell modalities and assays. RIDDLER utilizes a robust regression framework to model the expected distribution of reads genome-wide by accounting for assay-specific biases, identifying CNVs as outliers from that distribution. This versatile framing allows deployment of RIDDLER in multiple modalities with appropriate bias features. We demonstrate the accuracy of RIDDLER in calling single-cell CNVs and dissecting clonal heterogeneity in sc-ATAC-seq and sc-methylation. RIDDLER is more accurate and more robust to data sparsity than competing methods. We illustrate useful applications of RIDDLER for dissection of clonal structure, identification of subclonal accessibility peaks, and multimodal integration from CNV structure. RIDDLER stands out as a scalable, generalizable multi-modal method for accurate CNV detection, empowering studies aiming to link CNV dynamics to epigenetic alterations within the same cell.

Read PDF

Similar papers

Open access Jul 2026

CNV-Finder: streamlining copy number variation discovery

Abstract Motivation Copy Number Variations (CNVs) play pivotal roles in complex disease etiology, often requiring large sample sizes to analyze disease associations. While genotyping arrays offer a cost-effective approach for CNV detection using Log R Ratio (LRR) and B Allele Frequency (BAF) signals, existing independent array-based callers suffer from high false positive rates and noise susceptibility, burdening manual validation. Results We present CNV-Finder, a deep learning pipeline employing Long Short-Term Memory (LSTM) networks for large-scale CNV identification within user-defined genomic regions. Trained on expert-annotated samples from the Global Parkinson’s Genetics Program across four neurodegenerative disease-associated genes (PRKN, LINGO2, MAPT, SNCA), CNV-Finder integrates human feedback to iteratively improve performance. In benchmarking across 105 936 samples spanning 11 ancestries and nearly 150 cohorts, the model achieved 91% and 89% visual confirmation rates for PRKN deletions and duplications at high-confidence thresholds. In two validation cohorts, CNV-Finder nominated 83% fewer candidates than a popular Hidden Markov Model-based caller while maintaining higher confirmation rates. Validation through MLPA, short-read, and long-read sequencing demonstrated robust performance, generalizing to diverse signatures including homozygous deletions and SNCA triplications absent from training. Our findings highlight human expertise’s value in complex loci like 17q21.31. Availability and implementation CNV-Finder is freely available at https://github.com/nvk23/CNV-Finder.

Nicole Kuznetsov, Kensuke Daida, M. Makarious et al. · 0 citations
Open access Jul 2026

CNVeil resolves haplotype-specific copy number and uncovers subclonal architecture hidden from total copy number profiling in single-cell cancer genomes

Single-cell DNA sequencing (scDNA-seq) resolves copy number variation (CNV) at single-cell resolution, revealing tumor heterogeneity and subclonal structure. Most existing methods, however, infer only total copy number. Haplotype-resolved copy number, which captures allelic imbalance and clonal evolution, remains far less developed, largely because low coverage, allelic dropout, and technical noise in scDNA-seq make phased allelic inference substantially harder than total copy number estimation. We present CNVeil, a haplotype-aware framework that infers total, allele-specific, and chromosome-scale haplotype-resolved copy number from scDNA-seq data. CNVeil first builds robust total copy number profiles through highly variable bin selection, hierarchical clustering, subclone-aware ploidy estimation, and cross-cell consensus segmentation. Using this profile as a stable scaffold, it infers allele-specific copy number with an expectation-maximization algorithm applied to heterozygous SNP allele counts, then reconstructs haplotype-specific copy number by enforcing coherent haplotype orientation across adjacent segments via dynamic programming. We benchmarked CNVeil against 12 state-of-the-art methods, including eight total copy number callers, two allele-specific callers, and two haplotype-resolved callers, across 20 simulated and real datasets spanning six experimental settings, including high-multiplexed single-nucleus sequencing, Acoustic Cell Tagmentation (ACT), and 10x Chromium. This constitutes the largest comparative evaluation of single-cell copy number inference methods to date. CNVeil consistently outperformed existing tools in segmentation accuracy, ploidy inference, subclone identification, and allele-specific copy number estimation. In a breast cancer multi-omics (wellDR-seq) cohort, CNVeil uncovered haplotype-specific subclonal diversification invisible to total copy number analysis alone and linked allele-specific copy number states to transcriptional variation. By transforming sparse single-cell allelic signals into chromosome-scale haplotype-resolved profiles, CNVeil closes a major methodological gap and provides a scalable framework for studying tumor evolution and functional genomic heterogeneity at single-cell resolution.

Weiman Yuan, Can Luo, Yunfei Hu et al. · 0 citations
Open access Aug 2026

Detecting CYP2C19 deletions from genotyping array signals using neural networks

This work developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals and demonstrated that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.

Burak Yelmen, R. Hofmeister, Viido Kaur Lutsar et al. · 0 citations
Review Jul 2026

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Seungseok Kang, Yujin Oh, Sangwoo Kim · 0 citations
Open access Jul 2026

Integrative Modeling of Read Depth and B-Allele Frequency Improves Single-Cell Copy Number Calling from Targeted DNA Sequencing Panels

Copy number variations (CNVs) drive cancer initiation and progression, but resolving them at single-cell resolution from targeted DNA sequencing panels remains challenging. The Mission Bio Tapestri platform generates 2 complementary signals for CNV inference: sequencing depth and B-allele frequency (BAF) from heterozygous variants; however, existing methods such as karyotapR rely primarily on read depth, leaving allele-specific events unused. Here, we introduce scPloidyR, a hidden Markov model (HMM) that jointly models read depth and BAF at amplicon resolution for single-cell copy number calling from Tapestri data. scPloidyR fits per-chromosome Markov chains with copy number as the hidden state, factorizes emissions into depth and BAF likelihoods, and learns parameters by Baum–Welch expectation-maximization with Viterbi decoding. We compared scPloidyR with the established karyotapR Gaussian mixture model (GMM) in 2 simulation studies spanning BAF noise, variant density, amplicon density, sample size, and heterozygosity rate, and on a public Tapestri 5-cell-line mixture dataset. In simulations, scPloidyR substantially outperformed karyotapR on class-balanced metrics (macro-F1: 0.477 versus 0.273; alteration F1: 0.903 versus 0.381 in simulation study 1) when allelic information was available. Adding just one heterozygous variant per amplicon increased scPloidyR accuracy from 0.556 to 0.897 for gains. However, when BAF information was absent, karyotapR outperformed scPloidyR, and high BAF noise sharply degraded joint-model performance. On real data, scPloidyR produced more spatially coherent and biologically plausible copy number profiles. These results show that joint depth-BAF modeling benefits single-cell CNV calling when allelic information is available, while depth-only methods remain preferable when it is absent.

D. Pei, Rachel Griffard-Smith, Brahian Cano Urrego et al. · 0 citations
Open access Aug 2026

High-Specificity Detection of Chromosomal Mosaicism Reveals Cell-Type-Specific Genomic Alteration Patterns in Aging Tissues

Mosaic chromosomal alterations (mCAs) increase with age and are associated with multiple diseases, yet the cell types and states that harbor these alterations remain largely unknown. Because mCAs arise in individual cells prior to clonal expansion, they are typically rare and obscured in bulk data. We develop CHASM, a method for detecting chromosomal copy number alterations (CNA) from single-cell chromatin accessibility (scATAC-seq) data, a scalable modality that captures both cell state and chromosomal alterations. CHASM estimates a CNA-null background for each cell, providing an individualized expectation for chromosomal accessibility, which is critical in non-neoplastic tissues where alteration-carrying cells are not readily distinguishable from normal. By comparing each cell against its expected background, CHASM distinguishes chromosomal alterations from background variation and achieves more stringent control of false positives. We validate CHASM using in silico spike-in experiments, cross-modality comparisons with matched single-cell DNA and RNA data, and established genome-instability contrasts, including p53 deficiency and chromosome Y loss. Applied to multiple aging data sets, CHASM consistently recovers mCA burden in age-susceptible cell populations and reveals aging-associated signatures not detected by existing methods. In a cohort of 99 human kidney samples spanning age and disease conditions, CHASM identifies enrichment of mCAs in injury-associated cell states (VCAM1-high proximal tubule cells). Notably, CHASM detects the age-associated emergence of mCAs in cancer-relevant genomic regions, including chromosomes 3 gains and losses and chromosome 7 gain, in ostensibly normal cell populations. Cells harboring mCAs exhibit activation of injury-response regulatory programs and reduced epithelial identity programs, while elevated mCA burden in specific epithelial populations are associated with increased immune and stromal infiltration. Overall, we develop CHASM for high-specificity detection of CNA at single cell resolution. Applied across tissues, CHASM reveals aging-patterns of genome instability within cell types and implicates mCAs in early, pre-disease cellular states.

Xinyi E. Chen, Hanzhi Wang, Yilin Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.