Jan 2024· Briefings in Bioinformatics· Vol 27· 41 citations· ⚡ 4 influential· 161 references
BiologyComputer ScienceMedicine
TL;DR
This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, drug discovery, and single-cell analysis, and highlights major challenges that remain insufficiently addressed in prior reviews.
Abstract
Abstract Large language models (LLMs) are deep learning-based artificial intelligence models that have achieved remarkable success in natural language processing. Typically composed of neural networks with billions of parameters, they are trained on massive unlabeled datasets using self-supervised or semi-supervised learning. Beyond language, LLMs hold immense potential for addressing complex bioinformatics challenges. This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. We discuss critical components, including tokenization strategies for diverse biological data, transformer architectures, attention mechanisms, and pretraining approaches. We also survey currently available foundation models and their downstream applications across bioinformatics domains. Finally, we highlight major challenges that remain insufficiently addressed in prior reviews and outline future perspectives and design principles for next-generation biological language models, offering practical guidance for both users and developers.
Genomic language models (gLMs) are rapidly becoming important tools for learning biological information directly from sequence data. By adapting concepts from natural language processing, these models aim to capture contextual dependencies, regulatory grammar, evolutionary constraint, and sequence-level functional patterns that may be difficult to detect using alignment-based, motif-based, or conventional supervised methods alone. This systematic review evaluates recent model-development studies of genomic, RNA, nucleotide, codon-level, and regulatory DNA language models, with emphasis on model architecture, tokenization, training objective, biological task, benchmarking strategy, and reported limitations. A structured search of PubMed, Scopus, and Web of Science identified 469 records. After duplicate removal, screening, and full-text eligibility assessment, 58 studies met the strict inclusion criteria for primary model development or substantial model adaptation. The included studies covered diverse applications, including regulatory sequence prediction, variant-effect modeling, genome annotation, microbial and viral genome analysis, RNA splicing and regulation, codon optimization, mRNA design, and generative design of regulatory or RNA sequences. Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining. However, the evidence was heterogeneous and did not support a general claim of superiority over established bioinformatics tools or specialized supervised models. In several regulatory genomics tasks, specialized supervised models, k-mer-based approaches, or conventional deep-learning baselines remained competitive or superior to pretrained language-model representations. Generative models showed growing promise for RNA, codon, viral genome, and cis-regulatory element design, although many were evaluated mainly in silico. Overall, the field is advancing quickly, but broader impact will require standardized benchmarks, clearer reporting, stronger external validation, improved interpretability, and experimental confirmation of predicted or generated biological functions.
Mahinaz A. Mashhour, Manal Abdel Wahed, Mai S. Mabrouk· Biochemical and Biophysical...· 0 citations
An engineering-oriented, end-to-end roadmap that structures the full lifecycle of clinical language model systems—from model design and domain adaptation to optimization and real-world evaluation is introduced.
Large Language Models (LLMs) have emerged as a transformative technology in artificial intelligence, significantly advancing natural language understanding, generation, and reasoning capabilities. This survey reviews the evolution of language models from early statistical approaches to modern Transformer-based architectures and summarizes key developments, including attention mechanisms, scaling laws, alignment techniques, and efficient inference methods. The paper further explores the growing impact of LLMs on everyday life and a wide range of application domains, including healthcare, finance, education, agriculture, marketing, software engineering, and scientific research. To provide a systematic perspective, LLM applications are categorized according to their maturity level and integration across major artificial intelligence subfields, such as natural language processing, multimodal learning, intelligent decision support, autonomous agents, and knowledge-based systems. The survey highlights how these models enhance automation, data-driven decision-making, personalized services, and human–AI interaction across both consumer and industrial environments. Despite their remarkable capabilities, LLMs face several critical challenges, including high computational costs, limited interpretability, hallucinations, privacy and security risks, ethical concerns, and environmental sustainability issues. Existing mitigation approaches and recent advancements are reviewed to assess their effectiveness and limitations. Finally, the paper outlines key future research directions, including trustworthy and explainable AI, efficient model architectures, domain-specific adaptation, multimodal intelligence, and human-centered alignment. This survey provides a comprehensive overview of the current landscape, challenges, and future prospects of LLMs, serving as a valuable reference for researchers and practitioners.
P. Peykani, V. Charles, Ali Emrouznejad et al.· Archives of Computational Me...· 0 citations
Recent breakthroughs in foundation models and Large Language Models (LLMs) have introduced new opportunities for studying and decoding genomic sequences. Several state-of-the-art approaches, such as DNABERT2, rely on transformer-based architectures, while others, such as ConvNova, still build upon more conventional convolutional models. However, systematic benchmark comparisons across these methods remain scarce. Given that transformer-based models require extensive and costly pretraining, it is crucial to evaluate whether their performance gains justify this overhead. Moreover, LLMs such as DNABERT2 typically rely on Byte Pair Encoding (BPE) tokenization, whose relevance for DNA sequence representation is still debated within the genomics community. In this work, we investigate three key questions: (i) do transformer-based models provide sufficient improvements on fine-tuning tasks upon heavy pretraining, (ii) what is the actual contribution of pretraining in this setting, and (iii) how does BPE tokenization impact performance on genomics-related tasks?
R. Karpinsky, J. Mozziconacci, Mickaël Delcey· arXiv.org· 0 citations
Identifying gene-disease associations (GDAs) remains a fundamental challenge in biomedical research due to the enormous combinatorial space of candidate gene-disease pairs and the limited scalability of experimental validation. As wet lab studies cannot keep pace with rapidly expanding omics data, computational approaches have become essential for prioritizing plausible GDAs and accelerating biological discovery. Recent advances in artificial intelligence (AI), particularly graph neural networks (GNNs) and large language models (LLMs), are trans forming this field by enabling richer biological representations and more accurate predictive modeling. In this survey, we provide a unified and up-to-date overview of AI-driven GDA prediction. We first summarize major public resources containing gene, dis ease, and auxiliary biological information that underpin computational studies. We then review methodological developments ranging from traditional network-based methods to machine learning, deep learning, and the emerging integration of GNNs and LLMs, which has received limited attention in previous GDA-focused surveys. Representative applications in gene prioritization, drug repurposing, and clinical research are also discussed to demonstrate the practical impact of these approaches. Finally, we outline current challenges and promising future directions. By integrating data resources, methodological advances, and translational applications, this survey provides a comprehensive overview of modern AI techniques for GDA prediction and aims to support the development of more robust, interpretable, and clinically actionable computational tools. All curated resources and re viewed literature are publicly available in our GitHub repository (last updated September 2025; including peer-reviewed publications and preprints on AI-driven GDA prediction published through September 2025): https://github.com/linyaoyang/gene disease-association-prediction-papers.
Yupeng Zhai, Mengjing Li, Lei Yu et al.· IEEE transactions on computa...· 0 citations
While language models extract linguistic structures from text, similar approaches can uncover biological rules from genetic patterns. Though these methods have shown promise in single-cell analysis, bulk transcriptomics remains underexplored despite offering distinct clinical advantages including preserved tissue-level information, higher sequencing depth, and cost-effectiveness. Here, we present a transformer-based foundation model leveraging transcriptomic profiles from over 30,000 diverse bulk RNA samples, including normal tissues and various cancer types. Unlike conventional language models, our model incorporates specialized modules for modelling pairwise gene interactions through a dual representation system that captures both gene-level features and their higher-order relationships. Our model shows robust performance across multiple downstream applications. It achieves zero-shot accuracy of 78.81% in cancer classification without fine-tuning and outperforms existing approaches in cancer stages prediction through simple fine-tuning. Notably, it can extract critical gene interaction networks without relying on prior biological knowledge. More importantly, we leverage it to introduce dynamic interpretations to static bulk transcriptomic data, successfully modelling logical gene regulation rules with 91.07% overall accuracy—reaching 100% for rules related to key genes like GATA2 and SCL. With the context-specific modelling ability, it also identifies, for example, transcriptional dynamics in normal haematopoiesis and dysregulated circuits during transition to leukemic states. We further demonstrate clinical utility in predicting patient response to first induction chemotherapy (AUROC=0.75) in acute myeloid leukemia, a challenging task due to patient and mechanism heterogeneity. Through our novel response-directed feature-space gradient ascent approach, we identify patient-specific gene expression modifications that could computationally redirect resistant phenotypes toward responsive ones, revealing potential therapeutic targets aligned with individual patients' clinical features. These results establish our model as a powerful framework to extract useful information from bulk transcriptomics data and has potential applications in precision medicine by connecting computational predictions with biological insights.
Yi Chai, Yang Li, Jianbiao Zhou, Wee Joo Chng, Yang Zhang. Context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML [abstract]. In: Proceedings of Frontiers in Cancer Science 2025; 2025 Nov 5-7; Singapore. Philadelphia (PA): AACR; Cancer Res 2026;86(13_Suppl):Abstract nr P15.
Yi-Bo Chai, Yang Li, Jianbiao Zhou et al.· Cancer Research· 0 citations