Skip to content

A Dual-Level Sparsity Bayesian Framework for Rare-Variant Association Analysis Using Integrated Nested Laplace Approximation

Aug 2026 · Journal of Bioinformatics and Computational Biology · 0 citations

Abstract

Background. Rare genetic variants (minor allele frequency, MAF < 1%) carry a substantial share of the unexplained heritability of complex human traits, yet set-based association tests lose power precisely in the regime in which rare variation is most informative — when only a small minority of the variants inside a gene is causal and the aggregate contribution to phenotypic variance is well below one per cent. Burden tests dilute a genuine signal across neutral polymorphisms because they impose a common effect direction, whereas variance-component tests such as the optimal unified sequence kernel association test (SKAT-O) estimate a single dispersion parameter shared by every variant in the set and therefore cannot isolate the few variants that actually drive the association. Methods. We present BIM (Bayesian INLA Model), a hierarchical Bayesian framework that couples the Integrated Nested Laplace Approximation (INLA) with a dual-level sparsity prior. At the gene level, a bimodal mixture prior on the log-precision of the gene-specific variance component performs explicit model selection between an associated and a null state. At the variant level, a horseshoe prior supplies adaptive, variant-specific shrinkage whose posterior shrinkage factor admits a closed-form characterisation, so that individual pathogenic variants escape penalisation while neutral effects are driven towards zero. Functional annotations enter as hierarchical covariates that modulate both the location and the scale of the variant-effect prior. Gene-level evidence is summarised by marginal-likelihood Bayes factors and posterior inclusion probabilities, and discoveries are declared by a Bayesian false-discovery-rate (FDR) rule that controls the average local FDR of the selected set. Posterior uncertainty is decomposed into parametric, structural, internal (genotype uncertainty) and external (population structure) components through an explicit application of the law of total variance. Results. Across 100 replicated simulations calibrated on 1000 Genomes Project Phase 3 European haplotypes, BIM attained 75.2% power (95% confidence interval [CI] 69.5–81.0%) to detect causal genes in the most demanding scenario of 0.5% variance explained, against 58.0% for BATI, 38.0% for MiST, 22.0% (95% CI 16.9–27.1%) for SKAT-O and 9.0% (95% CI 5.4–12.6%) for the burden test (paired t-test P < 0.001 for every pairwise comparison). The realised FDR was 3.8%, below the nominal 5% level, and gene-level posterior inclusion probabilities were well calibrated against the empirical frequency of true association. In a whole-exome sequencing analysis of 500 chronic lymphocytic leukaemia (CLL) cases and 1,300 ancestry-matched controls, BIM returned a Bayesian FDR 5% discovery set of 12 genes, including the established susceptibility genes BRCA2 (Bayes factor, BF = 50.3) and CHEK2 (BF = 30.1) and the novel candidate ABCD3 (BF = 22.4), a peroxisomal ABC transporter with emerging links to cancer metabolism. The genomic inflation factor was λ GC = 1.02, and INLA posteriors agreed with gold-standard Hamiltonian Monte Carlo to a Pearson correlation of 0.9987 at a small fraction of the computational cost. Conclusion. Placing sparsity at two biological scales simultaneously, and solving the resulting latent Gaussian model with INLA rather than Markov chain Monte Carlo, converts a computationally prohibitive Bayesian formulation into a practical genome-scale tool. BIM delivers three- to four-fold power gains over SKAT-O in the sparse architectures characteristic of complex-trait genetics while retaining variant-level interpretability and calibrated error control. The open-source implementation is available at https://github.com/meibujun/BIM-INLA .

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.