Panvar is a tool developed to integrate existing software and resources to perform GWAS and fine-mapping in one seamless step and seeks to bridge the gap between GWAS and gene speeding up an important step of quantitative genetic studies.
Abstract
Genome-wide association studies (GWAS) use statistical models to correlate single nucleotide polymorphisms (SNPs) to a phenotype of interest. This scan of the entire genome identifies regions of association with a phenotype, but due to linkage disequilibrium (LD), GWAS on their own cannot identify single genes responsible for phenotypic variation. Rather, fine-mapping of GWAS regions is required, necessitating the use of additional tools and software. With the introduction of more pangenomic resources in a number of crops (Guo et al. 2025; Hufford et al. 2021), the fidelity of these fine-mapping efforts is growing, presenting the opportunity to leverage new information about allelic variation towards gene discovery (Shi et al. 2023; Della Coletta et al. 2021). Panvar is a tool developed to integrate existing software and resources to perform GWAS and fine-mapping in one seamless step. For each identified GWAS peak, panvaR outputs information about LD and SNP effect prediction for each SNP and by layering locations of nearby genes, creates a refined list of possible candidate genes. We have implemented Panvar as an R package, “panvaR”, which runs the analysis functions, creates interactive and static visualizations, and outputs results tables. This tool seeks to bridge the gap between GWAS and gene speeding up an important step of quantitative genetic studies.
This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets.
Christina Y. Feng, P. Sugier, Nan Zou et al.· Statistics in Medicine· 0 citations
The current challenges for using GWAS to prioritize variants for functional follow-up experiments are described and a multi-modal approach for resolving GWAS loci to a focused set of high-confidence variants for functional exploration is suggested.
Omar Y. Ahmed, N. Saravanan, A. B. Rovsing et al.· bioRxiv· 0 citations
Background Genome-wide association studies (GWAS) have identified thousands of loci associated with complex traits and diseases, yet translating these signals into biological insight remains challenging. Most associated variants are non-coding and reside in linkage disequilibrium (LD) blocks, where multiple correlated variants jointly contribute to association signals. These clusters, or haplotypes, may capture shared regulatory and functional contexts. Interpreting GWAS signals thus requires approaches that integrate regulatory, functional, and cross-trait evidence, while preserving the broader haplotypic context of disease-associated loci. At the same time, the rapid growth of publicly available GWAS summary statistics has enabled large-scale cross-trait analyses, but also introduced redundancy across closely related phenotypes. Efficient interpretation of GWAS data therefore requires tools that integrate heterogeneous data sources while preserving genomic and biological contexts. Results We present snpXplorer, an interactive web platform for haplotype-aware exploration and annotation of GWAS data. The platform incorporates >10,000 GWAS datasets from OpenGWAS and enables multi-scale analysis across variants, haplotypes, genes, and traits. Key features include (i) a haplotype-based representation of association signals derived from LD structure, (ii) a unified variant annotation framework integrating clinical annotations (ClinVar), allele frequencies (gnomAD), functional predictions (CADD, AlphaGenome), quantitative trait loci (GTEx), structural variation, and GWAS associations, and (iii) cross-trait exploration using semantic similarity-based clustering of phenotypes. Use cases centered on Alzheimer’s disease illustrate this utility: for example, at the TMEM106B locus, snpXplorer identified a haplotype linked to eleven distinct traits, revealing synergistic pleiotropy across neurological and behavioral phenotypes alongside antagonistic pleiotropy with height. Conclusions snpXplorer allows users to browse, filter, and inspect variant-, haplotype-, gene- and trait-level evidence, lowering the barrier to biological interpretation of GWAS results. Compared with existing tools that focus on specific aspects of GWAS interpretation, the strength of snpXplorer is that it reduces the need for fragmented queries across databases.
N. Tesi, G. Green, A. Salazar et al.· bioRxiv· 0 citations
Abstract Summary Genome-wide association studies (GWASs) have identified thousands of genetic variants associated with complex traits and diseases. However, explaining the mechanisms underlying phenotypic variation remains challenging. Here, we introduce SNPannotator, an automated post-GWAS analysis software package designed to streamline the interpretation of GWAS findings. Our pipeline implements a multi-step process that identifies proxy variants in high linkage disequilibrium (LD) with associated lead variants, then queries comprehensive resources (including Ensembl, the GTEx Portal, the eQTL Catalog, and STRING DB) for genomic position, deleteriousness, regulatory annotations, clinical significance, trait associations, expression (eQTLs) and splicing quantitative trait loci (sQTLs), and functional enrichment analyses and compiles the results into user-friendly reports. This package is implemented in the R programming language and includes auxiliary functions for variant lookup and LD exploration. SNPannotator provides a practical framework for efficiently deriving biologically meaningful insights from GWAS data and for assisting researchers in prioritizing candidate variants for functional validation. Availability and implementation The SNPannotator package is available from the Comprehensive R Archive Network (CRAN) at https://cran.r-project.org/web/packages/SNPannotator. The development version and tutorial is available on GitHub (https://github.com/omicslaboratory/SNPannotator). The online version of the package is available at https://omicslab.org/snpannotator.
Alireza Ani, I. Nolte, Zoha Kamali et al.· Bioinformatics· 1 citation
In livestock populations, genome-wide association studies (GWAS) can produce strong, apparently localized associations even when no truly discrete nearby causal effect exists. This occurs because small effective population sizes, strong family structure, long-range linkage disequilibrium (LD), and diffuse polygenic architecture can cause the effects of many variants to accumulate and be captured jointly across broad genomic intervals, making variant-level associations difficult to interpret biologically. Using real pig genotypes, we constructed a benchmark in which phenotypes were simulated under diffuse polygenic architecture across a genome partitioned into alternating effect and null windows, with central-null regions (at least 1 Mb away from effect-containing regions) positioned to detect long-range LD-driven signal propagation. We evaluated nine configurations of six GWAS methods (BOLT-LMM, REGENIE, fastGWA, FarmCPU, BLINK, and SLEMM) under this architecture. The central finding is that strong associations, of the kind normally read as evidence of nearby moderate- or large-effect variants, are produced by many of these methods even though the simulated signal is distributed across many tiny effects and cannot be localized to any single variant. The methods differed sharply in the extent of locus-level spillover: several produced large numbers of genome-wide significant loci within central-null regions, whereas the full-GRM mixed-model benchmark (SLEMM) produced no genome-wide significant loci in central-null regions. These results show that, under a highly polygenic architecture with livestock-like LD, GWAS tool choice has major consequences for biological interpretation. When the goal is to localize biologically meaningful signals rather than to flag association peaks that may merely reflect tiny effects accumulated through LD across a broad block, methods that control long-range LD spillover should be prioritized.
Xuesong Wang, Junji Wang, F. Tiezzi et al.· bioRxiv· 0 citations
Genome-wide association studies (GWAS) have revealed extensive polygenic signals and overlapping genetic architectures across human traits, creating a need for resources that connect trait-level genetic relationships with gene-level functional evidence. Here, we developed the Multi-Omics Causal Resource Database (MOCR-DB), an interactive platform that integrates large-scale GWAS summary statistics from UK Biobank, FinnGen, and the COVID-19 Host Genetics Initiative with molecular quantitative trait locus (QTL) datasets. In total, 613 traits with significant heritability were retained and harmonized using the Unified Medical Language System. MOCR-DB integrates phenotype-to-phenotype analyses, including genetic correlation and Mendelian randomization, with phenotype-to-gene analyses based on QTL-informed summary-data-based Mendelian randomization analysis (SMR) within a single searchable and interactive framework. The platform supports exploration of cross-trait genetic correlations, putative causal relationships, and candidate functional gene associations. An AI-assisted module provides concise plain-language summaries to help contextualize statistical findings. As a case study, we examined obesity and COVID-19 severity, where genetically predicted obesity showed a stronger association with critical COVID-19 and lung eQTL-based SMR analyses revealed distinct immune- and neuronal-related molecular patterns across severity groups. MOCR-DB thus provides a unified and accessible resource for investigating shared genetic architectures and prioritized functional gene candidates across complex traits, supporting the generation of reproducible and biologically interpretable hypotheses. The database is publicly available at https://chenhongwei.net/public/MOCRdb/.
Graphical Abstract
Data resources, analytical framework, and interpretation in MOCR-DB
The Multi-Omics Causal Resource Database (MOCR-DB) integrates large-scale GWAS summary statistics and molecular QTL datasets to provide a unified framework for genetic correlation, causal inference, and functional mediation. Data resources include GWAS summary statistics from UK Biobank, FinnGen, and the COVID-19 Host Genetics Initiative, together with 53 xQTL datasets across 49 tissues (eQTL, mQTL, sQTL, and caQTL). The analytical framework combines linkage disequilibrium score regression (LDSC) for estimating heritability and cross-trait genetic correlation, Mendelian randomization (MR) to infer potential causal relationships between traits, and summary-data-based Mendelian randomization (SMR) to identify tissue-specific functional genes. Results are presented through interactive genetic network searches that link diseases, biomarkers, lifestyle factors, and molecular traits via correlation, causality, and functional annotation. An AI-assisted module further facilitates causal and functional interpretation by summarizing complex results from LDSC, MR, and SMR analyses into accessible biological insights. Together, MOCR-DB provides systematic exploration of shared genetic architectures and functional mediators across complex human traits.