Skip to content
Open access

When the Background Matters: Topic-Dependent reference lists in GWAS and Exome Analyses

Aug 2026 · bioRxiv · 0 citations
Biology

TL;DR

To support reproducible best practice, a simple command set is provided for selecting and documenting study-appropriate backgrounds and for assessing sensitivity of GO Biological Process results to the chosen universe.

Read PDF

Similar papers

Open access Aug 2026

Robust Inference With Ghostknockoffs in Genome‐Wide Association Studies With Sample Relatedness

Genome‐wide association studies (GWASs) have been extensively adopted to depict the underlying genetic architecture of complex traits. Recent studies show that knockoff‐based methods can identify variants with unique, potentially causal effects on phenotypes. However, their statistical validity and effectiveness in studies with related individuals, such as the UK Biobank, remain unexplored. In this paper, we extensively evaluate a simple and effective analytical strategy that integrates GhostKnockoffs and state‐of‐the‐art marginal association tests. We show that this approach is robust to arbitrary relatedness structure as long as the input Z‐scores are derived from valid generalized linear mixed models. This robustness also extends GhostKnockoffs to other GWASs settings, including meta‐analysis of studies with sample overlap when the input score test Z‐scores are properly calibrated, and association test statistics beyond score tests in independent sample settings. We demonstrate the method's validity and practical utility using simulation studies and a meta‐analysis of nine European ancestral genome‐wide association studies and whole exome/genome sequencing studies for the Alzheimer's disease.

Xinran Qi, M. Belloy, Jiaqi Gu et al. · 0 citations
Open access Aug 2026

Scalable context-dependent single-cell eQTL mapping reveals disease-relevant regulatory variation beyond static models

Many disease-associated variants are thought to act through gene regulation, yet conventional eQTL mapping explains only a fraction of GWAS loci, potentially because regulatory effects vary across cellular states and environments. We present CASTIE, a scalable Poisson mixed-model framework that directly models sparse single-cell read counts and enables genome-wide testing of genotype-by-context interactions without pre-screening for static effects. Applying CASTIE to 1.2 million peripheral blood mononuclear cells from 982 OneK1K donors identified 3,155 context-dependent eQTL associations, including 2,022 eGenes without detectable static effects. These associations yielded 374 colocalizations across 94 traits, representing 270 unique loci, of which 197 were not recovered using the corresponding static eQTLs. The colocalizations linked trait associations to specific cellular contexts and genes including GCHFR, RNASET2 and ATP1A3. In adipose-derived mesenchymal stem cells exposed to metabolic stimulations, CASTIE increased eGene discovery by 36-92% across cell populations and identified stimulation-dependent regulatory effects at metabolic trait loci. Thus, modeling cellular context reveals disease-relevant regulatory variation beyond static eQTL mapping.

Y. C. Liu, A. Cuomo, Y. Huang et al. · 0 citations
Review Open access Aug 2026

A Guide for Exploring Pleiotropic Associations in Genome‐Wide Association Studies Using Summary Statistics

This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets.

Christina Y. Feng, P. Sugier, Nan Zou et al. · 0 citations
Open access Jul 2026

Cross-cohort analysis of expression and splicing quantitative trait loci in TOPMed.

Most genetic variants associated with complex traits are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterized 14,324 RNA-sequencing samples from the Trans-Omics for Precision Medicine program and performed expression and splicing quantitative trait locus (e/sQTL) analyses in six tissues and cell types, including whole blood (n = 6454) and lung (n = 1291). We detected tens of thousands of secondary cis-e/sQTLs, showing that secondary cis-e/sQTL discovery remains unsaturated. We fine-mapped UK Biobank-derived genome-wide association study (GWAS) signals from 164 traits and identified e/sQTL colocalizations for 10,611 GWAS signals, including 7096 that colocalize with secondary e/sQTLs. Our results suggest that even larger e/sQTL analyses will uncover additional secondary e/sQTLs, further benefiting GWAS interpretation.

Peter Orchard, T. Blackwell, L. Kachuri et al. · 0 citations
Review Open access Aug 2026

Revisiting differential expression analysis: An updated six-dimensional comparative study

Differential expression (DE) analysis is probably the most prevalent task for transcriptomic studies. However, recent technological advances have seen a revival of methodological interest in DE algorithms. In this study, we performed a comprehensive updated comparative study of 12 representative DE methods using 80 simulated and real datasets. We assessed the adaptability of these methods across varying sample sizes and diverse data scenarios. This evaluation compiled a six-dimensional overview of key properties: detection accuracy, sensitivity at a low false discovery rate, false positives, stability, robustness to outliers, and robustness under noisy conditions. Strikingly, no single methods outperformed others across all evaluation criteria and sample sizes, emphasizing data-specific and scenario-specific method choice. At the widely adopted small-sample size of n = 3, ABSSeq generally outperformed other methods. As sample size increased to n = 5, the sensitivity of DESeq2 and two edgeR v4 algorithms (QLF slightly better than LRT) also raise up under a stringent false-positive control. DESeq had even fewer false positives than DESeq2, at the price of reduced sensitivity. In terms of robustness, Wilcoxon and ROTS are robust to noises for small sample sizes. Moreover, Wilcoxon is also robust to outliers, together with several other methods (ABSSeq, voom, and T.test). NBPSeq and most methods had a good stability even at small sample sizes, except three methods (ROTS, DSS, and T.test). For larger sample sizes (n > 30), all methods performed much better. Finally, we provided a “BaGua (eight trigrams)” map summarizing the multi-dimensional performances of methods, as well as a tree diagram guiding practical method selection. Together, this study outlines a systematic and updated benchmarking framework for DE analysis, emphasizing a balance between accuracy and consistency.

Jian-Xiong Wu, Shaolei Lu, Hui Yao et al. · 0 citations
Open access Jul 2026

snpXplorer: an interactive platform for haplotype-aware exploration and integrated annotation of GWAS data

Background Genome-wide association studies (GWAS) have identified thousands of loci associated with complex traits and diseases, yet translating these signals into biological insight remains challenging. Most associated variants are non-coding and reside in linkage disequilibrium (LD) blocks, where multiple correlated variants jointly contribute to association signals. These clusters, or haplotypes, may capture shared regulatory and functional contexts. Interpreting GWAS signals thus requires approaches that integrate regulatory, functional, and cross-trait evidence, while preserving the broader haplotypic context of disease-associated loci. At the same time, the rapid growth of publicly available GWAS summary statistics has enabled large-scale cross-trait analyses, but also introduced redundancy across closely related phenotypes. Efficient interpretation of GWAS data therefore requires tools that integrate heterogeneous data sources while preserving genomic and biological contexts. Results We present snpXplorer, an interactive web platform for haplotype-aware exploration and annotation of GWAS data. The platform incorporates >10,000 GWAS datasets from OpenGWAS and enables multi-scale analysis across variants, haplotypes, genes, and traits. Key features include (i) a haplotype-based representation of association signals derived from LD structure, (ii) a unified variant annotation framework integrating clinical annotations (ClinVar), allele frequencies (gnomAD), functional predictions (CADD, AlphaGenome), quantitative trait loci (GTEx), structural variation, and GWAS associations, and (iii) cross-trait exploration using semantic similarity-based clustering of phenotypes. Use cases centered on Alzheimer’s disease illustrate this utility: for example, at the TMEM106B locus, snpXplorer identified a haplotype linked to eleven distinct traits, revealing synergistic pleiotropy across neurological and behavioral phenotypes alongside antagonistic pleiotropy with height. Conclusions snpXplorer allows users to browse, filter, and inspect variant-, haplotype-, gene- and trait-level evidence, lowering the barrier to biological interpretation of GWAS results. Compared with existing tools that focus on specific aspects of GWAS interpretation, the strength of snpXplorer is that it reduces the need for fragmented queries across databases.

N. Tesi, G. Green, A. Salazar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.