This work compared the performance of long-read sequencing against short-read sequencing and microarrays, identified two novel non-synonymous variants, including a unique InDel detected exclusively by long-read sequencing, and demonstrated superior detection power, particularly for InDels.
Abstract
While long-read sequencing technologies (e.g., PacBio Revio, ONT) have revolutionized high-quality genome assembly for the human pangenome, mitochondrial genome (mtDNA) analysis still largely relies on short-read and Sanger sequencing. However, short-read sequencing often lacks the resolution required to resolve complex variations due to the unique features of mtDNA, such as high mutation rates and repetitive homopolymeric regions, which frequently lead to alignment artifacts and mapping ambiguities. To address this, we evaluated whether applying long-read sequencing to mtDNA improves analytical quality in empirical data. Through comprehensive bioinformatics analyses, we compared the performance of long-read sequencing against short-read sequencing and microarrays. Our results revealed that long-read sequencing detected the highest number of variants (n = 533), significantly outperforming both short-read sequencing (n = 525) and microarrays (n = 49). Notably, both sequencing methods provided significantly higher resolution in haplogroup assignment compared to microarrays in terms of phylogenetic depth (p < 0.05). Long-read sequencing demonstrated superior detection power, particularly for InDels. We identified two novel non-synonymous variants, including a unique InDel detected exclusively by long-read sequencing. Protein modeling and stability analysis validated that this InDel causes structural instability (RMSD > 2.0 Å, ΔΔG = -45.21 kcal/mol). Furthermore, we confirmed that this novel InDel is shared among haplogroup A samples in both the 1000 Genomes Project ONT dataset and the Korean population, highlighting the practical implications of long-read sequencing for molecular biology and population genetics.
Fusion transcripts contribute to cancer, inherited diseases, developmental disorders, and evolution. Long-read RNA sequencing enables direct sequencing of full-length transcripts, creating new opportunities to detect complex fusion architectures, including previously inaccessible multi-segmented fusion transcripts. However, accurate transcriptome-wide fusion detection remains challenging because existing methods struggle to distinguish genuine fusion events from technical artefacts. Here, we present a comprehensive benchmark of transcriptome-wide fusion detection using simulated datasets and transcriptomes from three cancer cell lines across Oxford Nanopore Technologies (ONT) cDNA, PCR-cDNA, and direct RNA sequencing, Pacific Biosciences (PacBio) Kinnex sequencing, Illumina short-read RNA sequencing, six long-read fusion callers, and multiple analysis strategies. False-positive fusion calls remained the dominant limitation across sequencing platforms and algorithms. Increasing sequencing depth improved recall but also amplified spurious fusion calls, whereas higher read-support thresholds improved precision at the expense of sensitivity. ONT PCR-cDNA sequencing combined with CTAT-LR-Fusion achieved the best overall balance between precision and recall, whereas JAFFAL was the only caller to reliably identify simulated tri-gene fusions. Consensus calling reduced false positives but markedly reduced sensitivity, with only one of 400 simulated fusions detected by all six callers. Breakpoint localisation emerged as a major limitation across all methods. Long-read sequencing consistently recovered more validated fusion transcripts than short-read sequencing, enabled detection of complex tri-gene fusions, and produced more biologically plausible fusion landscapes with fewer promiscuous gene partners. Collectively, our results establish the first comprehensive benchmarking framework for transcriptome-wide fusion detection, using long-read RNA sequencing, and provide practical guidance for selecting sequencing workflows and computational strategies, while identifying key priorities for future algorithm development.
Ryley Dorney, Siyuan Wu, J. Y. Hung et al.· bioRxiv· 0 citations
A single genomic assay that delivers complete information across variant classes remains an aspirational goal. Currently, researchers and clinicians rely on an inefficient, expensive combination of short-read sequencing for single-nucleotide variants (SNVs) and small indels, comparative genomic hybridization (CGH) arrays for copy number variants (CNVs), and optical mapping and long-read sequencing for complex rearrangements, limiting the full potential of genomic discovery. To address these issues, TruPath Genome provides a one-test-for-all solution. By combining PCR-free whole-genome sequencing (WGS) with proximity-mapped read technology, it achieves high-resolution detection of SNVs and indels alongside long-range phasing for CNVs and structural variant (SV) refinement. We applied TruPath Genome on six clinical samples that were previously resolved by conventional methods. Across the cohort, TruPath Genome delivered coverage and variant-calling performance comparable to conventional WGS while achieving superior long-range phasing and enabling precise breakpoint resolution for clinically relevant structural events. This highlights TruPath Genome’s potential to consolidate genomic testing pipelines, accelerate diagnosis, and expand access to advanced genomic insights. Furthermore, its ultra-long-range data facilitates telomere-to-telomere assemblies and pangenome development, advancing our understanding of genome biology at an unprecedented scale.
Information about segregating haplotypes and structural variation (SV) can be extremely rich for a variety of applications in population genomics but remains largely inaccessible for many non-model species. Of the available methods, linked-read sequencing is especially promising for its low cost and scalability, but its adoption remains limited. One existing linked-read method is Haplotagging, which barcodes sequencing reads to reconstruct long molecules that encode haplotype information, with the potential to generate phased whole-genome data and detect structural variants. In this study, we present BLink-seq, a novel Haplotagging method that is compatible with standard short-read next-generation sequencing platforms, is locally reproducible with low-cost reagents, and is scalable for high-throughput sample processing. We optimized library preparation parameters, explored their relationship to linked-read library metrics, and validated phasing performance and structural variant detection in two evolutionary extremes: an experimental Drosophila melanogaster cross of inbred lines carrying known inversions, and four Atlantic silverside (Menidia menidia) parent-offspring trios sourced from highly outbred, wild-caught populations. We then applied our protocol to a cohort of 376 silversides to demonstrate its scalability and potential for SV detection and genotype imputation. Using BLink-seq, we generated chromosome-scale phased blocks and identified known inversions in both validation datasets. We discovered previously uncharacterized structural complexity within a known adaptive inversion on silverside chromosome 11, demonstrating that linked-read data can refine our understanding of SV architecture beyond what short reads alone can resolve. Finally, we provide a user guide for researchers interested in using BLink-seq.
Azwad R Iqbal, Pavel V. Dimens, J. Rick et al.· bioRxiv· 0 citations
Highly accurate sequencing approaches have substantially improved the detection of low-frequency variants by reducing technical artifacts and enhancing signal-to-noise ratios.
Farzaneh Darbeheshti, Azeet Narayan, Hayet Radia Zeggar et al.· Clinical Chemistry· 0 citations
SUMMARY Human genome sequencing typically relies on mapping reads to a reference genome to call variants, but this approach introduces technical biases, excluding duplicated and structurally polymorphic regions of the genome. To overcome this, we present a telomere-to-telomere genome benchmark with near-perfect accuracy across 99.4% of the diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), which were absent from prior benchmarks. We annotated genes and repeats on both haplotypes, including 19,956 protein-coding genes on the maternal haplotype and 19,190 on the paternal haplotype, and developed new methods to measure the accuracy of reads, phased variant call sets, and assemblies against a diploid reference. Genome-wide analyses show that de novo assembly resolves 2%–7% more sequence and outperforms variant calling accuracy by an order of magnitude, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.
Nancy F. Hansen, Nathan Dwarshuis, Hyun Joo Ji et al.· Cell· 6 citations
Long-read sequencing (LRS) has driven a transition in microbial genomics, overcoming the assembly fragmentation inherent to short-read sequencing. This review elucidates the impact of LRS across isolate genomics, metagenomics, and multi-omics domains. By spanning extensive repetitive regions, LRS facilitates the reconstruction of circular chromosomes and precisely resolves mobile genetic elements (MGEs). In metagenomics, LRS enables strain-level resolution, the recovery of circular metagenome-assembled genomes, and the precise localization of MGEs within host replicons. Furthermore, the single-molecule, amplification-free properties of LRS provide enhanced resolution of native epigenetic modifications and full-length transcriptomes. Despite these advancements, widespread implementation remains constrained by multidimensional challenges, including stringent high-molecular-weight DNA requirements, depth deficits, and computational overhead. Nevertheless, LRS is increasingly becoming the method of choice for isolate genomics and metagenomics. As detection technologies and algorithms progress, LRS will further improve our ability to decipher the structural and functional diversity of microbial ecosystems.
Xing Rao, Yu-He Gu, Gabriella et al.· GigaScience· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.