A novel tree-based feature engineering framework that uses Classification and Regression Trees (CART) to explicitly encode high-order interaction decision paths as dummy variables that significantly improves classification accuracy and model interpretability compared to using the original feature space alone is proposed.
Genetic contributions to complex traits are often mediated through coordinated gene–gene interaction networks, yet most existing association frameworks focus on marginal single-gene effects and overlook higher-order dependency structures. Direct modeling of interactions remains challenging due to combinatorial complexity and statistical instability. We introduce Interaction-Bridged Association Study (IBAS), a general framework that incorporates pathway-level interaction patterns into genotype–phenotype association analysis without explicitly enumerating interactions. IBAS leverages transcriptomic reference data to construct low-dimensional representations of pathway activity, which guide SNP-weighting and gene-level association testing within a kernel-based framework. In perturbation-based simulations, IBAS demonstrates improved stability and reproducibility compared to conventional TWAS and gene-based methods, while maintaining well-calibrated Type I error under phenotype permutation. Application to the WTCCC datasets identifies both known and novel genes across multiple complex diseases, including candidates with modest marginal effects missed by standard approaches. These findings are supported by replication in an independent cohort, and analyses across multiple reference tissues revealing both shared and tissue-specific signals. Overall, IBAS provides a statistically robust and computationally tractable framework for incorporating interaction effects into association mapping, extending beyond the single-gene paradigm and enabling more comprehensive characterization of complex trait. IBAS is available on GitHub at: https://github.com/QingrunZhangLab/IBAS
This review describes the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS and presents 30 methods designed to leverage AI in GWAS.
S. D’Antona, Mawada Elmagboul Abdalla Abakar, Daniele Ramazzotti et al.· BioData Mining· 0 citations
A novel statistical framework, iSVR, is presented that incorporates gene-environment interaction terms into a support vector regression model, enabling both modeling of interaction effects and their statistical testing, and provides a powerful approach to characterize the gene-environment interaction landscapes underlying complex traits.
It is increasingly recognized that genetic effects on complex traits and diseases are shaped by environmental context. Biobanks that measure diverse environmental exposures alongside genotypes and phenotypes at scale enable systematic study of gene-environment (G×E) interactions. Existing approaches, however, are limited in their ability to accurately model polygenic G×E involving many exposures across genome-wide genetic variants. It is unclear which exposure combinations are relevant for a given trait while distinguishing true interactions from environment-dependent heteroskedastic noise. To address these challenges, we develop Efficient multi-eNvironmental Gene-environment Interaction iNference Estimator (ENGINE), a supervised variance-component framework that learns an embedding that combines multiple environmental exposures while jointly estimating additive, G×E, and heteroskedastic noise components. To enable biobank-scale inference, ENGINE makes a single pass over the genotype matrix to cache genotype-dependent summaries, then assembles normal-equation components and gradients at each iteration. In simulations, ENGINE controls type I error rates, achieves high power, and accurately recovers the environmental embedding while remaining efficient at biobank-scale. It is roughly five-fold faster than the state-of-the-art method at biobank scale, making polygenic G×E analysis tractable when both the number of individuals and the number of SNPs reach the millions. Applied to five complex traits paired with lifestyle exposures in N = 291,273 unrelated white British individuals and M = 454,207 common SNPs (MAF>0.01) from the UK Biobank, ENGINE recovered G×E variance that was on average 1.4-fold larger than that captured by a single exposure and 5.5-fold larger than that captured by the first principal component of the exposures.
Genome-wide association studies (GWAS) have identified numerous variant-trait associations; yet, assigning effector genes to GWAS loci remains challenging. Similarity-based machine-learning methods, such as PoPS, prioritize effector genes from shared functional profiles among trait-relevant genes. These models assign a prioritization score for each gene and nominate a single effector gene within a GWAS locus. However, the scores provide limited insight into why a gene was prioritized or whether the nomination is biologically plausible. To address this gap, we introduce Kernelized Polygenic Priority Score, K-PoPS, a kernelized reformulation of PoPS that enables gene-centric explanations by decomposing each prediction into contributions from training genes. For each prioritized gene, K-PoPS reports top contributor genes and an anchor score that quantifies support from a user-defined set of trait-relevant genes. Across 38 Pan-UK Biobank traits, the full-feature OLS implementation underlying K-PoPS improved closest-gene enrichment relative to default PoPS for 26 of 37 evaluable traits. Across 25 traits with curated anchor sets, predictions supported by anchor scores were more enriched for closest-gene proxies than unsupported predictions. When applying to blood level apolipoprotein B, K-PoPS nominated SCARB1 over UBC gene, and further provided convincing explanations that support this prediction. Using explanation evidence, K-PoPS identified multiple plausible effector genes within a dilated cardiomyopathy locus, contrary to the parsimonious assumption. In summary, K-PoPS provides a post hoc framework for examining and interpreting GWAS effector-gene nominations.
Taotao Tan, Md. Abul Hassan Samee· bioRxiv· 0 citations
Machine learning has become an important tool in plant genomic prediction for modeling complex genotype–phenotype relationships and improving breeding decisions. However, many high-performing models, particularly ensemble and deep learning approaches, remain difficult to interpret, limiting their biological applicability. This review summarizes major machine learning methods and explainable artificial intelligence (XAI) approaches used in plant genomics, including SHAP (SHapley Additive exPlanations), LIME (Local Interpretable Model-Agnostic Explanations), attention mechanisms, permutation importance, tree-based feature importance, and gradient-based attribution methods. XAI can help identify influential SNPs, genomic regions, candidate genes, regulatory elements, and omics features associated with complex traits. For example, SHAP analysis in an almond germplasm collection identified a genomic region associated with shelling fraction, illustrating how XAI can generate testable hypotheses for further validation. The review further discusses applications in trait prediction, breeding, functional genomics, and multi-omics integration. Importantly, we emphasize major limitations, including data bias, model instability, correlated genomic markers, limited model transferability, and the common misconception that feature importance implies biological causality. We recommend integrating XAI with linkage disequilibrium pruning, stability assessment, biological annotation, and experimental validation before prioritizing candidate genes. Overall, XAI should be considered a framework for model interpretation, feature prioritization, and hypothesis generation rather than a replacement for experimental validation in plant genomics and breeding.
Agata Głuchowska, Muhammad Hafeez Ullah Khan, M. Pawełkowicz· Applied Sciences· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.