Jan 2024· Briefings in Bioinformatics· Vol 27· 41 citations· ⚡ 4 influential· 161 references
BiologyComputer ScienceMedicine
TL;DR
This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, drug discovery, and single-cell analysis, and highlights major challenges that remain insufficiently addressed in prior reviews.
Abstract
Abstract Large language models (LLMs) are deep learning-based artificial intelligence models that have achieved remarkable success in natural language processing. Typically composed of neural networks with billions of parameters, they are trained on massive unlabeled datasets using self-supervised or semi-supervised learning. Beyond language, LLMs hold immense potential for addressing complex bioinformatics challenges. This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. We discuss critical components, including tokenization strategies for diverse biological data, transformer architectures, attention mechanisms, and pretraining approaches. We also survey currently available foundation models and their downstream applications across bioinformatics domains. Finally, we highlight major challenges that remain insufficiently addressed in prior reviews and outline future perspectives and design principles for next-generation biological language models, offering practical guidance for both users and developers.
Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining.
Mahinaz A. Mashhour, Manal Abdel Wahed, Mai S. Mabrouk· Biochemical and Biophysical...· 0 citations
An engineering-oriented, end-to-end roadmap that structures the full lifecycle of clinical language model systems—from model design and domain adaptation to optimization and real-world evaluation is introduced.
This survey reviews the evolution of language models from early statistical approaches to modern Transformer-based architectures and summarizes key developments, including attention mechanisms, scaling laws, alignment techniques, and efficient inference methods.
P. Peykani, V. Charles, Ali Emrouznejad et al.· Archives of Computational Me...· 0 citations
This work investigates three key questions: do transformer-based models provide sufficient improvements on fine-tuning tasks upon heavy pretraining, what is the actual contribution of pretraining in this setting, and how does BPE tokenization impact performance on genomics-related tasks.
R. Karpinsky, J. Mozziconacci, Mickaël Delcey· arXiv.org· 0 citations
Identifying gene-disease associations (GDAs) remains a fundamental challenge in biomedical research due to the enormous combinatorial space of candidate gene-disease pairs and the limited scalability of experimental validation. As wet lab studies cannot keep pace with rapidly expanding omics data, computational approaches have become essential for prioritizing plausible GDAs and accelerating biological discovery. Recent advances in artificial intelligence (AI), particularly graph neural networks (GNNs) and large language models (LLMs), are trans forming this field by enabling richer biological representations and more accurate predictive modeling. In this survey, we provide a unified and up-to-date overview of AI-driven GDA prediction. We first summarize major public resources containing gene, dis ease, and auxiliary biological information that underpin computational studies. We then review methodological developments ranging from traditional network-based methods to machine learning, deep learning, and the emerging integration of GNNs and LLMs, which has received limited attention in previous GDA-focused surveys. Representative applications in gene prioritization, drug repurposing, and clinical research are also discussed to demonstrate the practical impact of these approaches. Finally, we outline current challenges and promising future directions. By integrating data resources, methodological advances, and translational applications, this survey provides a comprehensive overview of modern AI techniques for GDA prediction and aims to support the development of more robust, interpretable, and clinically actionable computational tools. All curated resources and re viewed literature are publicly available in our GitHub repository (last updated September 2025; including peer-reviewed publications and preprints on AI-driven GDA prediction published through September 2025): https://github.com/linyaoyang/gene disease-association-prediction-papers.
Yupeng Zhai, Mengjing Li, Lei Yu et al.· IEEE transactions on computa...· 0 citations
A context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML, which establishes the model as a powerful framework to extract useful information from bulk transcriptomics data and has potential applications in precision medicine by connecting computational predictions with biological insights.
Yi-Bo Chai, Yang Li, Jianbiao Zhou et al.· Cancer Research· 0 citations