This work describes the full spectrum of genetic variation and shows that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions.
Abstract
Long-read sequencing (LRS) and diploid genome assembly have enabled nearly complete structural variant (SV) discovery. Using 293 nearly complete genomes, we characterize the full spectrum of genetic variation and show that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions. We identify 24 gene-rich regions subject to megabase-scale variation, 2,293 potentially unstable tandem repeats, and 890 novel expression quantitative trait loci associated with SVs in humans. Expanding to 1,218 LRS samples from the 1000 Genomes Project and applying a newly developed cross-platform breakpoint evaluation tool, BoostSV, we construct a nonredundant callset comprising 614,522 SVs. We demonstrate the utility of this population-level SV reference callset by filtering >99% of the common variation from 44 unsolved LRS probands from the Undiagnosed Diseases Network to discover likely disease-causing SVs. Second, we genotype 1,053 high-impact biallelic SVs from the pangenome callset in 232,090 samples from All of Us and discover 105 SVs with significant associations, including 26% where the SV is the lead variant. This publicly available pangenome SV resource will drive new disease associations and further our understanding of the missing heritability of human genetic disease.
SUMMARY Human genome sequencing typically relies on mapping reads to a reference genome to call variants, but this approach introduces technical biases, excluding duplicated and structurally polymorphic regions of the genome. To overcome this, we present a telomere-to-telomere genome benchmark with near-perfect accuracy across 99.4% of the diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), which were absent from prior benchmarks. We annotated genes and repeats on both haplotypes, including 19,956 protein-coding genes on the maternal haplotype and 19,190 on the paternal haplotype, and developed new methods to measure the accuracy of reads, phased variant call sets, and assemblies against a diploid reference. Genome-wide analyses show that de novo assembly resolves 2%–7% more sequence and outperforms variant calling accuracy by an order of magnitude, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.
Nancy F. Hansen, Nathan Dwarshuis, Hyun Joo Ji et al.· Cell· 6 citations
Copy number variants (CNVs) are a major source of genetic diversity and could contain some of the missing heritability for mouse models of human disease. However, mouse CNVs have not been comprehensively characterized because they are difficult to resolve in repeat-rich, segmentally duplicated or reference sequence-absent regions of the genome. Here we analyzed long range sequence (LRS) data for 40 inbred mouse strains and characterized CNVs using pangenome graph-based (and other) methods and a C57BL/6J telomere to telomere (T2T) genome reference sequence. We resolved 1,594 high-confidence CNVs that often overlap tandem repeats (60.3%), segmental duplications (44.8%) or pericentromeric regions (11.5%); and 131 CNVs were T2T sequence-specific. CNVs affected 384 protein-coding genes, which spanned a range of important functional classes. The 40-strain pangenome map expanded the genome sequence from 2.29 to 3.32 Gb, with the wild-derived strains accounting for the largest sequence increments. Two different AIs were sequentially used to analyze this database and identify a 29-kb deletion CNV within the Nlrp1b locus of KK mice that contributed to the metabolic syndrome they develop. Human NLRP1 alleles also were associated with metabolic syndrome features in human populations. Hence, AI analyses of this comprehensive T2T pangenome-based resource could uncover some of the missing heritability for mouse models of human diseases and biomedical traits.
Wenlong Ren, Zhuanfen Cheng, Gary Peltz· bioRxiv· 0 citations
A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium’s (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.
Julian K. Lucas, Prajna Hebbar, Wen-Wei Liao et al.· bioRxiv· 1 citation
A single genomic assay that delivers complete information across variant classes remains an aspirational goal. Currently, researchers and clinicians rely on an inefficient, expensive combination of short-read sequencing for single-nucleotide variants (SNVs) and small indels, comparative genomic hybridization (CGH) arrays for copy number variants (CNVs), and optical mapping and long-read sequencing for complex rearrangements, limiting the full potential of genomic discovery. To address these issues, TruPath Genome provides a one-test-for-all solution. By combining PCR-free whole-genome sequencing (WGS) with proximity-mapped read technology, it achieves high-resolution detection of SNVs and indels alongside long-range phasing for CNVs and structural variant (SV) refinement. We applied TruPath Genome on six clinical samples that were previously resolved by conventional methods. Across the cohort, TruPath Genome delivered coverage and variant-calling performance comparable to conventional WGS while achieving superior long-range phasing and enabling precise breakpoint resolution for clinically relevant structural events. This highlights TruPath Genome’s potential to consolidate genomic testing pipelines, accelerate diagnosis, and expand access to advanced genomic insights. Furthermore, its ultra-long-range data facilitates telomere-to-telomere assemblies and pangenome development, advancing our understanding of genome biology at an unprecedented scale.
East Asian populations, representing over 20% of the global population, remain critically underrepresented in human genomic studies, limiting our understanding of population-stratified genetic variation and its implications for health and disease. Here we present the first phase of the Asian Pan-Genome project (APG), comprising 320 nearly complete, fully phased haploid genome assemblies from 160 East Asian individuals. These assemblies achieve unprecedented quality, with an average contig N50 of 144.3 megabase pairs and an average quality value of 64.5. Leveraging these superior assemblies, we reveal previously uncharacterized diversity in human repeatome, including population-stratified patterns in centromere satellites and rDNA arrays. Compared to existing global human genome assemblies, the newly generated genomes supplement 152 million base pairs of novel sequences, 355 gene gains, 18,300 structural variation loci and 26 large euchromatic inversions missing from current human pangenomes. We perform population stratification analyses of structural variations, and further resolve the structural haplotypes of complex genomic regions such as Major Histocompatibility Complex and Survival Motor Neuron loci across global pangenomes, exemplifying tandem-duplicate and inversion-rich complex locus architectures in the human genome, respectively. This resource provides a critical foundation for human genetic studies, especially for East Asian populations, promoting more accurate variant discovery, reducing bias, and ultimately advancing the equity and efficacy of genomic medicine.
The results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.
F. Mazzarotto, Özem Kalay, E. Arslan et al.· Genome Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.