Skip to content
Review

Genomic language models (gLMs): Emerging applications, challenges, and future directions in computational genomics.

Jul 2026 · Biochemical and Biophysical Research Communications - BBRC · Vol 831, pp. 154321 · 0 citations · 69 references
Medicine

TL;DR

Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining.

Abstract

Genomic language models (gLMs) are rapidly becoming important tools for learning biological information directly from sequence data. By adapting concepts from natural language processing, these models aim to capture contextual dependencies, regulatory grammar, evolutionary constraint, and sequence-level functional patterns that may be difficult to detect using alignment-based, motif-based, or conventional supervised methods alone. This systematic review evaluates recent model-development studies of genomic, RNA, nucleotide, codon-level, and regulatory DNA language models, with emphasis on model architecture, tokenization, training objective, biological task, benchmarking strategy, and reported limitations. A structured search of PubMed, Scopus, and Web of Science identified 469 records. After duplicate removal, screening, and full-text eligibility assessment, 58 studies met the strict inclusion criteria for primary model development or substantial model adaptation. The included studies covered diverse applications, including regulatory sequence prediction, variant-effect modeling, genome annotation, microbial and viral genome analysis, RNA splicing and regulation, codon optimization, mRNA design, and generative design of regulatory or RNA sequences. Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining. However, the evidence was heterogeneous and did not support a general claim of superiority over established bioinformatics tools or specialized supervised models. In several regulatory genomics tasks, specialized supervised models, k-mer-based approaches, or conventional deep-learning baselines remained competitive or superior to pretrained language-model representations. Generative models showed growing promise for RNA, codon, viral genome, and cis-regulatory element design, although many were evaluated mainly in silico. Overall, the field is advancing quickly, but broader impact will require standardized benchmarks, clearer reporting, stronger external validation, improved interpretability, and experimental confirmation of predicted or generated biological functions.

View source

Similar papers

Review Open access Jul 2026

Decoding viral protein sequences by large language models

This mini-review summarizes recent developments in devising and applying protein language models for biological sequences, emphasizing viral protein analysis, and outlines a road map for the potential application of LLMs in empowering virology research and pathogen surveillance.

Tianyi Fei, Siqi Li, Ziyue Yang et al. · 0 citations
Review Open access Jan 2024

Advancing bioinformatics with language models: components, applications, and perspectives

This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, drug discovery, and single-cell analysis, and highlights major challenges that remain insufficiently addressed in prior reviews.

Jiajia Liu, Mengyuan Yang, Yankai Yu et al. · 41 citations · ⚡4
Review Open access Jul 2026

SNP Detection Strategies in Genomic Research: A Comparative Review of Major Tools, Algorithms, Challenges and Applications

Single nucleotide polymorphisms (SNPs), representing the most frequent form of genetic variation, serve as essential biomarkers for mapping complex traits, tracing evolutionary lineages, and identifying the genetic basis of disease susceptibility. Next-generation sequencing (NGS) enables large-scale SNP discovery across diverse organisms, yet accurate detection remains challenging due to sequencing errors, genome complexity, reference bias, and coverage depth. This narrative/comparative review synthesizes primary tool publications, benchmarking studies, and recent literature identified from PubMed, Google Scholar, Web of Science, and official software documentation. This review compares SNP detection programs such as GATK, BCFtools, FreeBayes, SAMtools, and DeepVariant and their algorithmic structures, namely pileup- based, haplotype-based, and machine-learning approaches. In general, the comparison suggests that no single tool is the best, as the performance of these tools largely depends on the organism type and genome complexity, sequencing platform, sequencing depth, type of variant, and computational resources. GATK is sensitive and precise among the detection tools reviewed, yet computationally intensive; BCFtools is fast and versatile for non-human datasets; FreeBayes is a high-precision tool for haplotype-based and multiallelic variants; SAMtools is a powerful and stable tool for low-coverage data; and DeepVariant achieves high precision through deep learning architectures but at a high computational cost. We address important hurdles in SNP detection, such as polyploidy, heterozygosity, and repetitive regions, alongside the emerging necessity of Pangenomic Equity, defined here as the use of diverse pangenome references to reduce ancestry-related reference bias in SNP discovery. Lastly, we discuss emerging developments, including artificial intelligence-based variant calling, transformer-based models, graph-based reference genomes, and pangenome-aware methods that help to overcome the limitations of linear genomic models. Therefore, this review summarizes the key considerations for tool selection and outlines future directions for robust and inclusive SNP discovery in genomic research.

Shikhi Baruri, Sunita Khanal · 0 citations
Open access Jul 2026

An embedding-based framework enables statistical testing of gene-set function hypotheses inferred by large language models

Emerging large language models (LLMs) can infer gene functions directly from gene lists, enabling hypothesis generation without predefined gene sets. However, these LLM-derived predictions are qualitative, and principled statistical validation is lacking. Here, we develop an embedding-based statistical framework that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs. We benchmark seven state-of-the-art embedding models using curated and retrieval-augmented literature-derived gene descriptions across diverse biological contexts. OpenAI’s text-embedding-3-large and Google’s gemini-embedding-001 perform best, capturing gene-gene functional relationships in 88.7-92.5% of Gene Ontology biological processes and approximately 98.6% of canonical pathways. In gene-function association analyses, these models achieve high sensitivity (95.2-98.4%) and specificity (72.7-84.3%). Through contamination analysis and evaluation using experimentally informed protein assembly gene sets, our framework distinguishes biologically meaningful LLM-inferred hypotheses from noise, outperforming confidence-based inference and conventional enrichment analysis. We further develop the open-source R package DEGEmbedR and demonstrate its utility for interpreting a drug perturbation-derived differentially expressed gene (DEG) signature lacking significant conventional enrichment results. Together, these results establish LLM-derived embeddings as a quantitative foundation for functional genomics and the statistical validation of LLM-based gene function inference.

Yanhao Tan, Li-Ju Wang, Tianyuzhou Liang et al. · 0 citations
Open access Aug 2026

Protein language models and the long tail of functional diversity

Protein language model performance on downstream tasks depends on the pretraining data, motivating recent efforts to combine genomic- and metagenomic-derived protein sequences into large-scale atlases. Because these datasets are highly redundant, sequences are typically clustered by similarity and sampled during training. Sequences that do not belong to any cluster, known as “singletons”, are typically excluded from training and evaluation because they are considered to be artifacts. However, singletons represent the long tail of functional diversity and are abundant in many large-scale atlases: nearly 43% of the 3.34 billion sequences in the joint genomic-metagenomic dataset GigaRef are singletons. Here, we characterize singletons derived from UniRef and GigaRef by assessing whether clustering missed homologs, how much their exclusion affects protein language model (PLM) training, and which biological domains they contain. We find that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations. We also show that singletons share mutual information with clustered sequences, making them learnable by PLMs and useful for training. Finally, metagenomic singletons carry denser, more diverse domain content than clustered sequences, including domain-level homology that sequence-identity clustering misses. Together, these results support including singletons in PLM training and call for closer examination of data curation in large-scale integrated sequence atlases.

R. Vinod, Samir Char, Ava A. Amini et al. · 0 citations
Preprint Aug 2026

Frozen but Not Always Accessible: A Representation Analysis of Genomic Language Models

Genomic foundation models are increasingly reused as frozen feature extractors for downstream sequence prediction, offering a compute-efficient alternative to full fine-tuning. However, it remains unclear when biological information encoded by these models is accessible without task-specific adaptation. We present a representation-accessibility analysis of frozen genomic language models across regulatory, epigenetic, promoter, splice-site, and variant-effect prediction tasks. We evaluate DNABERT-2, Nucleotide Transformer, HyenaDNA, GENERATOR-v2, and Omni-DNA under unified frozen-probing protocols, while separating diagnostic readout analyses from validation-selected checks. Our results reveal a consistent task-dependent pattern: frozen probes recover 95-100 % of fine-tuned performance on promoter tasks, but average splice-site recovery drops to 60-88 %. Frozen embeddings are also competitive on broad Genomic Benchmark tasks such as coding-region and species-discrimination classification, but show larger gaps on some regulatory and OCR tasks. Layer-wise probing, in-silico mutagenesis, variant-effect prediction, and embedding geometry show that local biological signal is partially present in frozen representations, but is not always accessible through final pooled embeddings.

Nirjhor Datta, Swakkhar Shatabda, M. S. Rahman · 0 citations