A graph-based method for associating sequence variation with traits using short-read genome sequencing data, and shows that it captures complex forms of genetic variation missed by other methods, suggest that integration of pangenomic methods into human genetic studies will improve trait association and genomic prediction at a meaningful subset of genes.
Abstract
The Human Pangenome Reference Consortium has generated 462 open-access reference genomes and a variation graph that represents differences among them, providing a substrate for pangenome-based analysis methods that overcome the longstanding limitation of comparing all genomic data to a single linear reference. A key unresolved question is the extent to which these approaches can improve trait mapping. We investigate this using the genetics of gene expression variation as a model. We developed a graph-based method (EdgeDepth) for associating sequence variation with traits using short-read genome sequencing data, and show that it captures complex forms of genetic variation missed by other methods. We evaluated trait mapping performance using 430 samples with deep RNA-seq data, and found that pangenomic methods enable the detection of expression quantitative trait loci involving multiallelic indels and structural variants, leading to increased power at a subset of genes. These include 812 genes (7.9% of total) with ≥20% improvement in statistical significance relative to the 1000 Genomes Project callset, and 185 (1.8%) with a 50% improvement, 10 of which are candidates to explain prior GWAS results. Notably, these analyses implicate GBAP1 pseudogene copy number as a causal factor in Crohn’s disease, likely via miRNA-mediated regulation of GBA1, which explains prior GWAS results based on flanking SNPs. The inclusion of pangenome-specific variation also improved the performance of gene expression prediction models, with median variance explained increasing from 10.1% to 12.5%, and 14.6% of genes showing significant improvement (Δr2>0.05). Taken together, these results suggest that integration of pangenomic methods into human genetic studies will improve trait association and genomic prediction at a meaningful subset of genes.
Structural variations (SVs) represent a significant source of genomic diversity, with demonstrated roles in livestock gene expression and traits. However, a comprehensive understanding of the SV landscape across large sample sets and its impact on gene regulation in cattle remains incomplete. This study aimed to construct high-fidelity pangenome graphs by integrating both assembly-based and whole-genome sequencing (WGS) derived SV catalogs. We evaluated the efficacy of pangenome graphs for SV genotyping and identified 80,328 high-quality SVs from a cohort of 2929 samples. We systematically characterized these SVs, including their linkage disequilibrium with single nucleotide polymorphisms (SNPs), functional annotations, formation mechanisms, and genomic distributions. Furthermore, we generated paired WGS (24.4 ×) and blood RNA-seq data in 170 Simmental cattle. Utilizing our pangenome graphs, we identified 637 SV-expression quantitative trait loci (SV-eQTL), which accounted for 10.81% of expression heritability of target genes, with 38.09% of the effects linked to promoter/enhancer regions. Forty-six of these SV-eQTL were replicated using CattleGTEx results through SV imputation using a joint SNP-SV reference panel. Notably, insertions in the GHSR gene were significantly associated with its expression levels, likely linked to Bos indicus cattle adaptation to heat tolerance. Our findings provide novel insights into the SV landscape and its contribution to gene regulation, underscoring its importance in cattle genetics and genomics.
This work describes the full spectrum of genetic variation and shows that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions.
J. Lin, J. Gustafson, J. Wertz et al.· medRxiv· 0 citations
A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium’s (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.
Julian K. Lucas, Prajna Hebbar, Wen-Wei Liao et al.· bioRxiv· 1 citation
Purpose We introduce GraNPA, standing for Graph Node-Phenotype Association, a method performing a GWAS-like analysis on a pangenome variation graph (PVG) built using a small number of individual genome sequences, without the need for additional population materials or kinship information for qualitative phenotypes. This method reduces the number of individuals required for association studies and prevents reference bias from variant calling in these types of analyses. Background A PVG represents the multiple alignment of a set of complete genomes. It contains all variations, from single nucleotide polymorphisms (SNPs) to large structural variations (SVs), which are represented as nodes in the graph. By integrating phenotype information within nodes, we can assign a Phenotype Score (PS) to each node in the PVG and identify phenotype-related regions directly within it. These regions represent statistically significant shifts in PS distribution, highlighting their implication in the phenotype. Finally, GraNPA provides their positions and scores for further analysis. Results This method was tested using simulated data and two publicly available datasets: the Sub1A gene locus for Oryza sativa in a 13 individuals PVG, and the insertion responsible for the white-headed cattle with a PVG of 24 individuals. Source code of GraNPA is available here https://forge.ird.fr/diade/graphgwas/granpa under GNU GPLv3. Conclusion GraNPA was able to identify the expected area in two simulated datasets and the responsible loci for these two known traits using only a few dozen complete genomes in these PVGs. While currently limited to qualitative phenotypes, this method opens the way to more efficient ones relying on PVGs and few individuals.
Camille Carrette, François Sabot, Cédric Muller· bioRxiv· 0 citations
Panvar is a tool developed to integrate existing software and resources to perform GWAS and fine-mapping in one seamless step and seeks to bridge the gap between GWAS and gene speeding up an important step of quantitative genetic studies.
Collin Luebbert, Rijan R. Dhakal, Phillip Ozersky et al.· bioRxiv· 0 citations
This work provides a foundation for applications that link epigenome variation to gene expression in human cells, by benchmarking methods on a per-gene basis, illustrating their use in a disease context and making trained models available to the community.
Fatemeh Behjati Ardakani, Shamim Ashrafiyan, Laura Rumpf et al.· Genome Biology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.