Aug 2026· Frontiers in Genetics· 0 citations· 86 references
TL;DR
This survey reviews the basic principles of LLMs and summarizes representative applications in gene and genome sequence analysis, protein structure and function prediction, and drug design, including virtual screening and personalized medicine.
Abstract
The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.
This study constructed a pH-responsive P-TN/SF@Fe-Cur composite coating that demonstrated significant anti-infective, anti-inflammatory, antioxidant, pro-angiogenic, and pro-osteogenic effects in rat subcutaneous infection and femoral defect models.
The results show that alternative transcript diversity extensively enters translation-supported proteoform space and establish a systematic link between transcript variation and protein functional diversification.
Felicia T. Jiang, Dengwang Chen, Ziwei Wang et al.· bioRxiv· 1 citation
Protein therapeutic design and property prediction are frequently hampered by data scarcity. Here we propose a model, DyAb, that addresses these issues by leveraging a pair-wise representation to predict differences in binding affinity, rather than absolute values. DyAb is built on top of a pre-trained protein language model and achieves a Spearman rank correlation of up to 0.85 on binding affinity prediction across monoclonal antibodies targeting three different antigens (EGFR, IL-6, and an internal target), given as few as 100 training data. We employ DyAb in two design contexts: as a ranking model to score combinations of known mutations, and combined with a genetic algorithm to generate new sequences. Our method consistently generates antibody variants with high binding rates, including designs that improve on the binding affinity of the lead molecule by more than ten-fold. DyAb represents a powerful tool for optimizing antibody binding affinity in low data regimes common in early-stage drug development.
Joshua Yao-Yu Lin, Jennifer L. Hofmann, Andrew Leaver‐Fay et al.· mAbs· 1 citation
Due to its importance and wide adoption, wheat cultivation is promptly required to shift towards sustainable practices, reducing the dependency on chemical components. Among bio-based solutions aimed at securing the sustainability of wheat cultivation, biostimulants offer a versatile platform of eco-friendly tools assuring sustainability and profitability. Microalgae present a concrete example of a biostimulant source due to their richness in metabolites and high value products. Therefore, this study evaluated the biostimulant potential of eleven eco-extracts prepared from soil-isolated microalgae strains. Eco-extracts applied via soil drench at low dose (0.1 g/L) were investigated for their biostimulant effects on wheat growth, physiology, yield, and quality under controlled conditions. Results demonstrated significant ameliorations in treated plants as compared to the control, with no phytoinhibitory effects. Remarkable enhancements were notable in growth parameters such as shoot and root lengths (+40-70%), physiological traits such as total chlorophyll and stomatal conductance (+7-52%), yield components in the example of grain number per spike and thousand grain weight (+17-103%), and grain quality namely protein and polyphenol content (+2-fold to 4-fold). Similarly, phosphorus accumulation and uptake were significantly improved, while soil physicochemical status was ameliorated, indicating enhanced fertility. Multivariate analysis and composite index ranking marked Chlorella sp. GA18, Chlorella sp. GA65, Scenedesmus sp. GA69, and Chlorococcum sp. GA63 as eco-extracts with consistent performances across all plant traits. These findings highlighted the promising potential of integrating microalgae-based eco-friendly extracts in sustainable wheat cultivation.
Amer Chabili, Z. Hakkoum, F. Minaoui et al.· Plant Science· 1 citation
ProteinReasoner is developed, a multimodal generative protein foundation model that sequentially connects amino acid sequence, evolutionary constraints and three-dimensional structure within a shared autoregressive architecture and suggests a general route towards reasoning across interdependent representations in other scientific domains.
Chaozhong Liu, Linlin Chao, Shaomin Ji et al.· bioRxiv· 1 citation
HydroGym is introduced, a solver-independent reinforcement learning platform providing more than 60 validated, openly available flow control environments spanning from canonical laminar flows to complex turbulent flows, with systematic progression in the Reynolds number up to Re = 4 × 105, and Mach number variations in two and three dimensions.
Christian Lagemann, Sajeda Mokbel, Miro Gondrum et al.· Nature· 1 citation