Skip to content
Review Open access

Advancing bioinformatics with language models: components, applications, and perspectives

Jan 2024 · Briefings in Bioinformatics · Vol 27 · 41 citations · ⚡ 4 influential · 161 references
Biology Computer Science Medicine

TL;DR

This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, drug discovery, and single-cell analysis, and highlights major challenges that remain insufficiently addressed in prior reviews.

Abstract

Abstract Large language models (LLMs) are deep learning-based artificial intelligence models that have achieved remarkable success in natural language processing. Typically composed of neural networks with billions of parameters, they are trained on massive unlabeled datasets using self-supervised or semi-supervised learning. Beyond language, LLMs hold immense potential for addressing complex bioinformatics challenges. This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. We discuss critical components, including tokenization strategies for diverse biological data, transformer architectures, attention mechanisms, and pretraining approaches. We also survey currently available foundation models and their downstream applications across bioinformatics domains. Finally, we highlight major challenges that remain insufficiently addressed in prior reviews and outline future perspectives and design principles for next-generation biological language models, offering practical guidance for both users and developers.

Read PDF

Similar papers

Review Jul 2026

Genomic language models (gLMs): Emerging applications, challenges, and future directions in computational genomics.

Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining.

Mahinaz A. Mashhour, Manal Abdel Wahed, Mai S. Mabrouk · 0 citations
Review Open access Aug 2026

The Versatility of Large Language Models: A Comprehensive Review and Structured Survey of Architectures, Applications, Challenges, and Future Trajectories

This survey reviews the evolution of language models from early statistical approaches to modern Transformer-based architectures and summarizes key developments, including attention mechanisms, scaling laws, alignment techniques, and efficient inference methods.

P. Peykani, V. Charles, Ali Emrouznejad et al. · 0 citations
Jun 2026

DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks

This work investigates three key questions: do transformer-based models provide sufficient improvements on fine-tuning tasks upon heavy pretraining, what is the actual contribution of pretraining in this setting, and how does BPE tokenization impact performance on genomics-related tasks.

R. Karpinsky, J. Mozziconacci, Mickaël Delcey · 0 citations
Review Jun 2026

Decoding Gene-Disease Associations with Computational Methods: A Survey.

Identifying gene-disease associations (GDAs) remains a fundamental challenge in biomedical research due to the enormous combinatorial space of candidate gene-disease pairs and the limited scalability of experimental validation. As wet lab studies cannot keep pace with rapidly expanding omics data, computational approaches have become essential for prioritizing plausible GDAs and accelerating biological discovery. Recent advances in artificial intelligence (AI), particularly graph neural networks (GNNs) and large language models (LLMs), are trans forming this field by enabling richer biological representations and more accurate predictive modeling. In this survey, we provide a unified and up-to-date overview of AI-driven GDA prediction. We first summarize major public resources containing gene, dis ease, and auxiliary biological information that underpin computational studies. We then review methodological developments ranging from traditional network-based methods to machine learning, deep learning, and the emerging integration of GNNs and LLMs, which has received limited attention in previous GDA-focused surveys. Representative applications in gene prioritization, drug repurposing, and clinical research are also discussed to demonstrate the practical impact of these approaches. Finally, we outline current challenges and promising future directions. By integrating data resources, methodological advances, and translational applications, this survey provides a comprehensive overview of modern AI techniques for GDA prediction and aims to support the development of more robust, interpretable, and clinically actionable computational tools. All curated resources and re viewed literature are publicly available in our GitHub repository (last updated September 2025; including peer-reviewed publications and preprints on AI-driven GDA prediction published through September 2025): https://github.com/linyaoyang/gene disease-association-prediction-papers.

Yupeng Zhai, Mengjing Li, Lei Yu et al. · 0 citations
Jul 2026

Abstract P15: Context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML

A context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML, which establishes the model as a powerful framework to extract useful information from bulk transcriptomics data and has potential applications in precision medicine by connecting computational predictions with biological insights.

Yi-Bo Chai, Yang Li, Jianbiao Zhou et al. · 0 citations