Jun 2026· IEEE transactions on computational biology and bioinformatics· Vol PP· 0 citations
Medicine
Abstract
Identifying gene-disease associations (GDAs) remains a fundamental challenge in biomedical research due to the enormous combinatorial space of candidate gene-disease pairs and the limited scalability of experimental validation. As wet lab studies cannot keep pace with rapidly expanding omics data, computational approaches have become essential for prioritizing plausible GDAs and accelerating biological discovery. Recent advances in artificial intelligence (AI), particularly graph neural networks (GNNs) and large language models (LLMs), are trans forming this field by enabling richer biological representations and more accurate predictive modeling. In this survey, we provide a unified and up-to-date overview of AI-driven GDA prediction. We first summarize major public resources containing gene, dis ease, and auxiliary biological information that underpin computational studies. We then review methodological developments ranging from traditional network-based methods to machine learning, deep learning, and the emerging integration of GNNs and LLMs, which has received limited attention in previous GDA-focused surveys. Representative applications in gene prioritization, drug repurposing, and clinical research are also discussed to demonstrate the practical impact of these approaches. Finally, we outline current challenges and promising future directions. By integrating data resources, methodological advances, and translational applications, this survey provides a comprehensive overview of modern AI techniques for GDA prediction and aims to support the development of more robust, interpretable, and clinically actionable computational tools. All curated resources and re viewed literature are publicly available in our GitHub repository (last updated September 2025; including peer-reviewed publications and preprints on AI-driven GDA prediction published through September 2025): https://github.com/linyaoyang/gene disease-association-prediction-papers.
Introduction Identifying potential miRNA–disease associations is essential for clarifying the molecular basis of complex diseases and accelerating the discovery of diagnostic biomarkers and therapeutic targets. However, the performance of existing computational methods is often limited by sparse biological interaction networks, highly imbalanced disease distributions, and the small number of experimentally validated associations. Methods To address these challenges, we propose MMAG, a novel framework that formulates miRNA–disease association prediction as a meta-conditional distribution alignment problem on multi-scale biological graphs. MMAG integrates three complementary components. First, a multi-scale representation learning module captures hierarchical biological information from local topological connectivity, mesoscopic functional organization, and global spectral structure. Second, a meta-learning strategy models each disease as an individual task, enabling the model to learn disease-specific prototype representations from support samples and adapt effectively to few-shot settings. Third, a conditional adversarial alignment mechanism reduces feature distribution discrepancies across diseases with different data scales, thereby enhancing cross-task knowledge transfer and generalization. Results Extensive experiments demonstrate that MMAG consistently outperforms several state-of-the-art methods under few-shot, long-tailed, and cross-dataset transfer scenarios. Discussion These results indicate that MMAG provides an effective and scalable solution for miRNA–disease association prediction and offers a promising strategy for broader biological network inference tasks.
Zhang Yu, Zuo Xuan, Tang Ying et al.· Frontiers in Bioinformatics· 0 citations
The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.
S. D’Antona, Mawada Elmagboul Abdalla Abakar, Daniele Ramazzotti et al.· BioData Mining· 0 citations
Abstract The rapid expansion of high-throughput omics has created molecular datasets of unprecedented scale and complexity. These data are rich in biological information yet inherently sparse and high-dimensional, often limiting the effectiveness of conventional machine learning techniques. Foundation models (FMs), built on large-scale self-supervised pretraining, offer a robust alternative by learning generalizable representations directly from raw biological data. This review systematically analyzes the emerging landscape of FMs in omics research, spanning sequence modeling, cell state characterization, and multimodal integration. We organize the current literature into three distinct paradigms—sequence-centric, cell-centric, and multi-omics—to clarify a field currently fragmented by diverse tokenization strategies and architectural choices. Beyond methodology, we evaluate the practical utility of these models in tasks ranging from biomarker discovery to perturbation response prediction. We also identify critical barriers to adoption, including high computational costs, interpretability challenges, and the lack of standardized benchmarks. To support reproducible research, we provide a curated catalog of essential datasets and evaluation frameworks. Finally, we propose a roadmap for the next generation of FMs, advocating for architectures that move beyond statistical correlation to incorporate causal reasoning, temporal dynamics, and autonomous experimental validation.
Haozhe Liu, Wenhao Cai, Yizheng Sun et al.· Briefings in Bioinformatics· 0 citations
Cancer is a complex and heterogeneous disease that is characterized by multi-level biological variability. Advances in high-throughput technologies have led to large-scale, high-dimensional data sets in cancer research, creating a pressing need for powerful computational techniques for successful data analysis. Current techniques may be inadequate for this purpose, thus underscoring the potential of artificial intelligence (AI) and machine learning (ML) for successful data analysis. This review provides a comprehensive pipeline for artificial intelligence/machine learning in cancer research, including preclinical research, clinical decision support, and real-world implementation. It emphasizes several important technologies, data integration, and implementation challenges. The review critically examines multi-omics fusion architectures, regularization-based machine learning, batch-effect harmonization, explainable AI, and federated learning, while addressing translational barriers including algorithmic bias, covariate drift, and regulatory asynchrony across Indian, US, and EU frameworks. Anchored by Decision Curve Analysis as a clinical utility benchmark, this narrative framework establishes that meaningful progress in precision oncology, early detection, and patient outcomes demands not only predictive accuracy but also externally validated, population-representative, and governance-compliant AI systems capable of sustained real-world oncology impact.
Shalini Saha, Md Saif Ali, A. Tengli et al.· Journal of Translational Med...· 0 citations
The rapid expansion of next-generation sequencing technologies has generated unprecedented volumes of genomic data; however, translating these data into reliable and clinically actionable insights remains a major challenge in precision medicine. Artificial intelligence (AI) has emerged as a key enabling technology across the genomic medicine pipeline, supporting variant detection, variant interpretation, polygenic risk prediction, disease subtyping, biomarker discovery and treatment–response modelling. This review provides a clinically oriented, pipeline-based synthesis of contemporary AI applications in genomic medicine. Major computational paradigms, including machine learning, deep learning, ensemble methods, multimodal AI, explainable AI frameworks and emerging foundation models, are discussed in the context of their contribution to genomic analysis and clinical decision support. Particular emphasis is placed on the factors that determine model robustness and clinical utility, including dataset composition, class imbalance, label noise, calibration, ancestry representation, distributional shift and external validation. Evidence from rare genetic disorders, cardiovascular genetics and precision oncology is examined to illustrate both successful translational applications and persistent barriers to implementation. The review further analyses common sources of failure in real-world genomic AI systems, including overfitting, limited transportability across populations and sequencing environments, inadequate interpretability, and insufficient prospective validation. Ethical and regulatory challenges are discussed in relation to clinical accountability, genomic privacy, algorithmic bias and equitable implementation. Ultimately, the successful clinical translation of genomic AI will depend not only on methodological innovation, but also on rigorous validation, transparent reporting, continuous calibration, robust governance and sustained expert oversight.
Alexandra-Maria Blaga, Răzvan-Octavian Mihuț, A. Treteanu et al.· International Journal of Mol...· 0 citations
Recent literature on the application of artificial intelligence (AI) and data science within bioinformatics-driven cancer drug discovery is synthesized, examining how these tools are reshaping target identification, molecular design, biomarker discovery, and treatment personalization.
Yejide Eniola Dabiri· Magna Scientia Advanced Rese...· 0 citations