A novel statistical framework, iSVR, is presented that incorporates gene-environment interaction terms into a support vector regression model, enabling both modeling of interaction effects and their statistical testing, and provides a powerful approach to characterize the gene-environment interaction landscapes underlying complex traits.
Abstract
Genome-wide association studies and genomic prediction are fundamental for investigating complex traits, but they have different objectives and are rarely unified within a shared analytical framework. Although machine learning has broadened the applicability of both lines of research, their combined use in detecting gene-environment interactions remains underexplored. This study presents a novel statistical framework, iSVR, that incorporates gene-environment interaction terms into a support vector regression model, enabling both modeling of interaction effects and their statistical testing. By formulating a score test based on M-estimation theory within this framework, the iSVR facilitates robust detection of gene-environment interactions while accommodating complex genotype-phenotype relationships. Extensive simulations demonstrate that the iSVR effectively controls the type I error rate and attains competitive or improved statistical power relative to existing methods under the investigated scenarios. Application to soybean and GAW19 datasets further highlights the iSVR's ability to accurately predict trait values and identify significant gene-environment interactions. Collectively, these findings illustrate that the unification of association testing and predictive modeling within a common statistical framework provides a powerful approach to characterize the gene-environment interaction landscapes underlying complex traits.
Motivation Identifying the mechanisms by which genetic variants affect the molecular response to an applied treatment is important across multiple biological fields, and an effective approach to this end is interaction molecular QTL mapping. However, the statistical models commonly used to detect such gene-by-treatment interactions (G×T) are non-trivially misspecified, and this can lead to decreased power. Results We developed an R software package, DetectGxT, that uses nonlinear regression to more accurately model the relationship between the genotype and the transformed molecular count phenotypes. It also optionally models donor or polygenic random effects. Simulations show that nonlinear regression can increase the power to detect interactions. In existing interaction expression QTL mapping data from primary human neural progenitor cells, nonlinear and linear regression approaches identified overlapping but distinct sets of gene-SNP pairs with significant G×T interactions. Overall, our results suggest an advantage of nonlinear regression over linear regression in detecting G×T interactions on molecular phenotypes. Availability The DetectGxT software is available at https://github.com/yharigaya/detectgxt. Contact milove@email.unc.edu, william.valdar@unc.edu
Yuriko Harigaya, Michael I. Love, William Valdar· bioRxiv· 0 citations
This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets.
Christina Y. Feng, P. Sugier, Nan Zou et al.· Statistics in Medicine· 0 citations
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
J. Zhu, A. Baousi, A. P. Morris et al.· medRxiv· 0 citations
Genetic contributions to complex traits are often mediated through coordinated gene–gene interaction networks, yet most existing association frameworks focus on marginal single-gene effects and overlook higher-order dependency structures. Direct modeling of interactions remains challenging due to combinatorial complexity and statistical instability. We introduce Interaction-Bridged Association Study (IBAS), a general framework that incorporates pathway-level interaction patterns into genotype–phenotype association analysis without explicitly enumerating interactions. IBAS leverages transcriptomic reference data to construct low-dimensional representations of pathway activity, which guide SNP-weighting and gene-level association testing within a kernel-based framework. In perturbation-based simulations, IBAS demonstrates improved stability and reproducibility compared to conventional TWAS and gene-based methods, while maintaining well-calibrated Type I error under phenotype permutation. Application to the WTCCC datasets identifies both known and novel genes across multiple complex diseases, including candidates with modest marginal effects missed by standard approaches. These findings are supported by replication in an independent cohort, and analyses across multiple reference tissues revealing both shared and tissue-specific signals. Overall, IBAS provides a statistically robust and computationally tractable framework for incorporating interaction effects into association mapping, extending beyond the single-gene paradigm and enabling more comprehensive characterization of complex trait. IBAS is available on GitHub at: https://github.com/QingrunZhangLab/IBAS
Genotype-environment association (GEA) analyses are widely used to identify loci underlying local adaptation by examining correlations between allele frequencies and environmental variables across a species’ range. A major challenge for this approach is distinguishing true adaptive signals from spurious associations arising from population structure. Several methods have been developed to account for population structure, but these methods can suffer from reduced statistical power or increased false positives under some conditions. To address this, we introduce a new GEA method, termed SimGEA. In essence, SimGEA infers a neutral evolutionary model that reproduces the population structure observed in empirical data and uses this model to simulate neutral alleles. By applying the same GEA statistic to both the empirical and simulated data, SimGEA evaluates the significance of observed associations against neutral expectations that account for population structure. We compared the performance of SimGEA with that of existing GEA methods, including LFMM2 and BayPass, using simulations of local adaptation in two-dimensional space. We found that SimGEA consistently controlled the false discovery rate without substantially sacrificing statistical power across the scenarios examined. These results suggest that calibrating statistics using neutral simulations provides a robust and flexible approach for accounting for population structure in GEA analyses.
Takahiro Sakamoto, Sam Yeaman· bioRxiv· 0 citations
Background Although many complex phenotypes and diseases are influenced by shared genetic and environmental factors, risk prediction methods typically rely on genetic information from a single trait, leaving a rich source of predictive information largely unexploited. Phenotypic correlations can potentially be used to improve the accuracy of polygenic scores (PGS), but the conditions under which correlated traits meaningfully enhance prediction remain poorly understood. Here, we develop a general theoretical and simulation framework that quantifies the extent to which correlated “helper” traits improve predictive accuracy and identifies the factors that determine the magnitude of these gains. Results We show that helper traits can substantially improve predictive accuracy, with the magnitude of these gains governed by baseline model performance, genetic and environmental correlations, and the heritability of the target and helper traits, providing principled guidance for helper-trait selection. Paradoxically, when the target trait is itself weakly heritable, helper traits need not be highly heritable to substantially improve the accuracy of PGS, because low-heritability traits can still capture non-redundant environmental factors shared with the target trait. We empirically evaluated the use of helper traits by developing PGS models to predict type 2 diabetes using data from the UK Biobank. Helper traits substantially improved predictive accuracy relative to a single-trait PGS (AUC-ROC = 0.907 versus 0.677) and achieved performance comparable to models that use HbA1c (AUC-ROC = 0.889), the current clinical gold-standard biomarker. Conclusions Our results establish a general theoretical and practical framework for exploiting correlated traits to improve polygenic prediction, provide principled guidance for selecting informative helper traits, and demonstrate how shared genetic and environmental architecture can be leveraged to substantially increase predictive accuracy. Furthermore, we developed an interactive web application to estimate the expected gain in accuracy from candidate helper traits using empirically measurable quantities.
Kaiqian Zhang, Rob Bierman, Joshua M. Akey· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.