This mini-review summarizes recent developments in devising and applying protein language models for biological sequences, emphasizing viral protein analysis, and outlines a road map for the potential application of LLMs in empowering virology research and pathogen surveillance.
Abstract
Abstract Large language models (LLMs) for biological sequences are transforming computational biology, enabling a nuanced understanding of protein and nucleotide sequence data. Recent models, including ESM2, ESM3, AlphaGenome, Evo-1, and Evo-2, adapt natural language processing principles to the biological domain by learning high-dimensional hidden representations that capture evolutionary constraints, structural patterns, and functional motifs. This mini-review summarizes recent developments in devising and applying such models, emphasizing viral protein analysis. We highlight studies that have leveraged sequence-based LLMs in the protein domain (i.e. protein language models, or PLMs) for important application tasks such as viral protein annotation, variant effect prediction, and immune escape characterization. Additionally, we present a benchmark evaluation of these state-of-the-art protein language models to evaluate their core ability to capture evolutionary relationships between viral protein sequences. By discussing the opportunities and challenges of PLMs, the review outlines a road map for the potential application of LLMs in empowering virology research and pathogen surveillance.
AbstractAa Background The transformer architecture in deep learning has revolutionized protein sequence analysis. Recent advancements in protein language models have paved the way for significant progress across various domains, including protein function and structure prediction, multiple sequence alignments, and mutation effect prediction. A protein language model is commonly trained on individual proteins, ignoring the interdependencies between sequences within a genome. However, biological understanding reveals that protein–protein interactions span entire genomic regions, underscoring the limitations of focusing solely on individual proteins.Ab Findings To address these limitations, we propose a novel approach that extends the context size of transformer models across the entire viral genome. By training on large genomic fragments, our method captures putative long-range dependencies consistent with inter-protein relationships and encodes protein sequences with integrated information from distant proteins within the same genome, offering benefits across downstream tasks. Viruses, with their densely packed genomes, minimal intergenic regions, and protein annotation challenges, are ideal candidates for genome-wide learning. We introduce a long-context protein language model, trained on entire viral genomes, leveraging a biologically informed sparse attention mechanism in which inter-protein links are inferred computationally and used as sparsity priors. Our semi-supervised approach supports long sequences of up to 61,000 amino acids (aa).Ac Conclusion Our evaluations show improved prediction of masked aa and improved downstream discrimination relative to single-protein models and long-context baselines, with additional validation that our inferred links correlate with independently curated interaction resources.
T. Dejean, Barbra D. Ferrell, Zachary D. Schreiber et al.· GigaScience· 0 citations
Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining.
Mahinaz A. Mashhour, Manal Abdel Wahed, Mai S. Mabrouk· Biochemical and Biophysical...· 0 citations
Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.
Yo Akiyama, Zhidian Zhang, Olivia Tang et al.· Cell· 2 citations
What are the fundamental units of protein sequences? Most protein language models treat amino acids as tokens, yet biological functions are not encoded at the single-residue level. Instead, they emerge from combinations of residues that form functional units that corresponds to conserved sequence motifs. Just like how modern language models work at the level of learned sub-word units instead of characters, we argue that explicitly modeling at the functional motif level provides both mechanistic insight into sequence-function relationships and interpretable control over protein generation. We demonstrate this framework on metalloproteins, where similar coordination chemistry is shared at the structure level yet how sequence controls metal specificity remains underexplored. We develop a three-step workflow: (1) construction of a dictionary of functional motifs; (2) prediction of the next motif and inter-motif gaps; and (3) sequence infilling given the predicted anchor motifs. Motif analysis confirms that the extraction captures known metal-coordination chemistry. Compared to random masking, motif-guided generation improves Conserved Domain Database annotation rates with greater target-family enrichment. Compared to generation approaches that use BPE tokenization, our approach achieves more specific hits on metalloprotein families and substantially reduces off-target annotations. Generation output directly mirrors dictionary composition, demonstrating that motif vocabularies provide explicit control over generation scope. AlphaFold 3 structure prediction with explicit metal ions confirms plausible coordination geometry, validating functional binding sites in generated sequences. Together, functional motif modeling enables interpretable, controllable protein generation, an important step toward compositional design of novel protein functions.
Boon How Low, W. Goh, Boyang Li et al.· Proceedings of the 32nd ACM...· 0 citations
RNA-binding proteins (RBPs) are essential across biology, from viruses to complex multicellular organisms. They regulate gene expression and cellular responses, making RNA recognition central to understanding health and disease. Biochemical, biophysical, and structural studies have defined core principles of RNA binding, but recent RNA interactome surveys have expanded the RBP repertoire and revealed many noncanonical RNA-binding regions. This diversity demands highly scalable predictive methods. Here, we review machine learning predictors built on protein language models and structure-aware representations. These approaches improve generalisability, reduce reliance on deep evolutionary information, and enable proteome-scale prediction of RNA-binding residues, providing a route to map and interpret the molecular logic of protein-RNA interactions.
Rozeena Arif, Alfredo Castello· Current Opinion in Structura...· 0 citations