Skip to content
Review Open access

A Guide for Exploring Pleiotropic Associations in Genome‐Wide Association Studies Using Summary Statistics

Aug 2026 · Statistics in Medicine · Vol 45 · 0 citations · 71 references
Medicine

TL;DR

This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets.

Abstract

ABSTRACT Genome‐wide association studies (GWAS) have shown that pleiotropy, whereby a single genetic variant or gene influences multiple traits, is common in complex human diseases. Detecting cross‐phenotype associations from GWAS summary statistics remains challenging because of small effect sizes, extensive multiple testing, heterogeneous effects, and possible differences in effect direction across traits. Methods that jointly analyze multiple traits can improve the ability to detect pleiotropic signals while retaining the practical advantages of summary statistic‐based analyses. Although a range of statistical approaches has been developed for this purpose, practical guidance on their application, assumptions, and interpretation remains limited. This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets. We also highlight the importance of accounting for effect heterogeneity, correlation, and biological group structure at the gene and pathway levels in the detection and interpretation of pleiotropic association signals.

Read PDF

Similar papers

Open access Jul 2026

snpXplorer: an interactive platform for haplotype-aware exploration and integrated annotation of GWAS data

Background Genome-wide association studies (GWAS) have identified thousands of loci associated with complex traits and diseases, yet translating these signals into biological insight remains challenging. Most associated variants are non-coding and reside in linkage disequilibrium (LD) blocks, where multiple correlated variants jointly contribute to association signals. These clusters, or haplotypes, may capture shared regulatory and functional contexts. Interpreting GWAS signals thus requires approaches that integrate regulatory, functional, and cross-trait evidence, while preserving the broader haplotypic context of disease-associated loci. At the same time, the rapid growth of publicly available GWAS summary statistics has enabled large-scale cross-trait analyses, but also introduced redundancy across closely related phenotypes. Efficient interpretation of GWAS data therefore requires tools that integrate heterogeneous data sources while preserving genomic and biological contexts. Results We present snpXplorer, an interactive web platform for haplotype-aware exploration and annotation of GWAS data. The platform incorporates >10,000 GWAS datasets from OpenGWAS and enables multi-scale analysis across variants, haplotypes, genes, and traits. Key features include (i) a haplotype-based representation of association signals derived from LD structure, (ii) a unified variant annotation framework integrating clinical annotations (ClinVar), allele frequencies (gnomAD), functional predictions (CADD, AlphaGenome), quantitative trait loci (GTEx), structural variation, and GWAS associations, and (iii) cross-trait exploration using semantic similarity-based clustering of phenotypes. Use cases centered on Alzheimer’s disease illustrate this utility: for example, at the TMEM106B locus, snpXplorer identified a haplotype linked to eleven distinct traits, revealing synergistic pleiotropy across neurological and behavioral phenotypes alongside antagonistic pleiotropy with height. Conclusions snpXplorer allows users to browse, filter, and inspect variant-, haplotype-, gene- and trait-level evidence, lowering the barrier to biological interpretation of GWAS results. Compared with existing tools that focus on specific aspects of GWAS interpretation, the strength of snpXplorer is that it reduces the need for fragmented queries across databases.

N. Tesi, G. Green, A. Salazar et al. · 0 citations
Open access Aug 2026

Robust Inference With Ghostknockoffs in Genome‐Wide Association Studies With Sample Relatedness

Genome‐wide association studies (GWASs) have been extensively adopted to depict the underlying genetic architecture of complex traits. Recent studies show that knockoff‐based methods can identify variants with unique, potentially causal effects on phenotypes. However, their statistical validity and effectiveness in studies with related individuals, such as the UK Biobank, remain unexplored. In this paper, we extensively evaluate a simple and effective analytical strategy that integrates GhostKnockoffs and state‐of‐the‐art marginal association tests. We show that this approach is robust to arbitrary relatedness structure as long as the input Z‐scores are derived from valid generalized linear mixed models. This robustness also extends GhostKnockoffs to other GWASs settings, including meta‐analysis of studies with sample overlap when the input score test Z‐scores are properly calibrated, and association test statistics beyond score tests in independent sample settings. We demonstrate the method's validity and practical utility using simulation studies and a meta‐analysis of nine European ancestral genome‐wide association studies and whole exome/genome sequencing studies for the Alzheimer's disease.

Xinran Qi, M. Belloy, Jiaqi Gu et al. · 0 citations
Open access Jul 2026

PanvaR: An R package for fine-mapping and visualizing results from genome-wide association studies

Panvar is a tool developed to integrate existing software and resources to perform GWAS and fine-mapping in one seamless step and seeks to bridge the gap between GWAS and gene speeding up an important step of quantitative genetic studies.

Collin Luebbert, Rijan R. Dhakal, Phillip Ozersky et al. · 0 citations
Open access Aug 2026

Widely used GWAS methods can be poorly suited to SNP-level localization under diffuse polygenic architecture in livestock

In livestock populations, genome-wide association studies (GWAS) can produce strong, apparently localized associations even when no truly discrete nearby causal effect exists. This occurs because small effective population sizes, strong family structure, long-range linkage disequilibrium (LD), and diffuse polygenic architecture can cause the effects of many variants to accumulate and be captured jointly across broad genomic intervals, making variant-level associations difficult to interpret biologically. Using real pig genotypes, we constructed a benchmark in which phenotypes were simulated under diffuse polygenic architecture across a genome partitioned into alternating effect and null windows, with central-null regions (at least 1 Mb away from effect-containing regions) positioned to detect long-range LD-driven signal propagation. We evaluated nine configurations of six GWAS methods (BOLT-LMM, REGENIE, fastGWA, FarmCPU, BLINK, and SLEMM) under this architecture. The central finding is that strong associations, of the kind normally read as evidence of nearby moderate- or large-effect variants, are produced by many of these methods even though the simulated signal is distributed across many tiny effects and cannot be localized to any single variant. The methods differed sharply in the extent of locus-level spillover: several produced large numbers of genome-wide significant loci within central-null regions, whereas the full-GRM mixed-model benchmark (SLEMM) produced no genome-wide significant loci in central-null regions. These results show that, under a highly polygenic architecture with livestock-like LD, GWAS tool choice has major consequences for biological interpretation. When the goal is to localize biologically meaningful signals rather than to flag association peaks that may merely reflect tiny effects accumulated through LD across a broad block, methods that control long-range LD spillover should be prioritized.

Xuesong Wang, Junji Wang, F. Tiezzi et al. · 0 citations
Open access Sep 2026

Large-scale pleiotropic analysis across cancers reveals shared genetic mechanisms and identifies novel functional genes

Abstract Pleiotropic genetic loci have been increasingly reported in cancer, and identifying genetic variants with pleiotropic associations can reveal shared biological pathways influencing multiple cancers. Using summary statistics from genome-wide association studies for 37 cancer types (N = 433 836), we identified extensive genome-wide and local genetic correlations among cancers. Through pairwise pleiotropic analysis, we identified 75 243 significant pleiotropic single nucleotide polymorphisms (SNPs) across 372 cancer pairs, among which 3472 were lead SNPs with potential regulatory functions. Using FUMA and MAGMA, we identified 2527 pleiotropic risk loci and 4272 candidate pleiotropic genes. Notably, genes such as TERT (5p15.33), POU5F1B (8q24.21), and FANCA (16q24.3) exhibited widespread pleiotropy across multiple cancer types. Pathway enrichment analysis highlighted the critical roles of pigment synthesis, metabolism, and apoptosis in skin-related cancers, while cross-cancer enrichment analysis emphasized pathways related to apoptosis, chromatin structure, and intermediate filaments. We also identified 33 novel functional genes harboring previously unreported cancer risk variants. Drug-gene interaction analysis revealed several repositionable FDA-approved drugs. Importantly, drug sensitivity assays demonstrated that bosutinib and cobimetinib exhibited promising therapeutic potential in breast cancer cell lines. Finally, we developed the PleioCancer database (https://gonglab.hzau.edu.cn/PleioCancer/), providing a comprehensive resource for cancer pleiotropy research. These findings have important implications for carcinogenesis cancer, prevention and treatment.

Unknown authors · 0 citations
Jul 2026

A novel support vector regression approach for detecting gene-environment interactions and predicting trait values.

A novel statistical framework, iSVR, is presented that incorporates gene-environment interaction terms into a support vector regression model, enabling both modeling of interaction effects and their statistical testing, and provides a powerful approach to characterize the gene-environment interaction landscapes underlying complex traits.

Xuewei Li, Wanqiu Xie, Liang Tong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.