Back to feed
Review

AI-Driven Protein Research: From Prediction to Design.

2026 · Methods in molecular biology · Vol 3030, pp. 213-226 · 0 citations
Medicine

TL;DR

This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.

View source

Similar papers

Review Aug 2026

Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling

Accurate protein structure prediction is fundamental to structural biology because protein structure underlies molecular function and provides a basis for mechanistic interpretation. Recent advances in deep learning have transformed the field from multiple sequence alignment (MSA)-driven monomer folding into broader frameworks capable of modeling protein complexes and increasingly heterogeneous molecular systems. Existing reviews have summarized this progress from the perspectives of representative models, application domains, and protein design. Building on these efforts, this review focuses on the methodological evolution of the field itself. It examines recent developments through three closely related dimensions: representations and data, architectures and learning strategies, and confidence and evaluation. Within this perspective, the field is organized into four methodological phases and three cross-cutting transitions: from explicit evolutionary coupling features and early contact prediction to learned sequence representations in AlphaFold2, RoseTTAFold, and ESMFold; from protein-only monomer folding to increasingly integrated modeling of heterogeneous molecular systems in AlphaFold-Multimer, RoseTTAFoldNA, and AlphaFold3; and, more recently, from prediction-oriented structure inference to design-oriented generative modeling in RFdiffusion and related frameworks. This framework provides a clearer understanding of how methodological shifts have shaped the capabilities, limitations, and practical roles of recent models.

Wengan He, Yongsheng Luo, Lihong Jiang et al. · 0 citations
Jul 2026

Deep Learning for Proteins Notebook Series Teaches AI for Biomolecular Structure Prediction and Design

Computational methods for predicting and designing biomolecular structures are increasingly powerful. Although previous approaches relied on physics-based modeling, modern tools (e.g., AlphaFold2 in CASP14) leverage artificial intelligence (AI) to achieve significantly improved performance. The growing effect of AI-based tools in protein science necessitates enhanced educational materials that improve AI literacy among established scientists seeking to deepen their expertise and new researchers entering the field. To address this need, we developed Deep Learning for Proteins: a series of 10 interactive notebook modules that introduce fundamental machine-learning concepts, guide users through training machine-learning models for protein-related tasks, and ultimately present cutting-edge protein structure prediction and design pipelines. By using only a web browser, learners can access state-of-the-art computational tools used by professional protein engineers that range from all-atom protein design to fine-tuning protein language models for biophysically relevant functional tasks. By increasing accessibility, this notebook series broadens participation in AI-driven protein research. The complete notebook series is publicly available at https://github.com/Graylab/DL4Proteins-notebooks .

Michael Chungyoun, G. Au, Britnie Carpentier et al. · 0 citations
Jul 2026

Generalizable Protein Folding Pathway Exploration with DA2-GRASP: Extending Beyond Miniproteins.

Elucidating protein dynamics is crucial for deciphering fundamental biological processes, from enzyme catalysis to cellular signaling, as its dysregulation directly causes protein misfolding diseases such as Alzheimer's and Parkinson's. While artificial intelligence has revolutionized static protein structure prediction, capturing the high-dimensional dynamics of protein folding remains a formidable challenge that limits our ability to fully understand these vital biological phenomena. Here we present DA2-GRASP, a computational framework that overcomes this barrier by integrating deep learning with advanced sampling techniques to map protein folding pathways with unprecedented efficiency and accuracy. DA2-GRASP learns low-dimensional latent representations of protein conformations via a variational autoencoder and combines multidirectional generative sampling guided by local potential energy gradients to efficiently steer conformational transitions along energetically favorable paths, enabling accurate and efficient reconstruction of folding pathways. Our method achieves sublinear computational scaling with sequence length, contrasting the quadratical scaling of molecular dynamics-based conventional approaches, enabling tractable simulations. It maintains high precision in quantifying mutation-induced perturbations to folding thermodynamics, crucial for understanding disease mutations. It also enables atomistic characterization of the folding process of medium-sized proteins such as ubiquitin and small ubiquitin-like modifier (SUMO, ∼80 residues) on standard workstations, a task typically requiring specialized supercomputing platforms such as Anton. Analysis of these proteins provides new mechanistic insights into how structurally similar folds with low sequence identity navigate divergent folding pathways. DA2-GRASP thus establishes a versatile and powerful framework for exploring protein-folding dynamics and their functional consequences.

Yanbing Wen, Hao Dong · 0 citations
Review Jul 2026

The Advantages of AI for Computational Protein Studies and Looking Ahead at the Next Challenges: Single Structures Are Not Enough.

The ability to understand proteins and their behaviors has been drastically improved by major successes in structure prediction and the appearance of Large Protein Language Models. The speed with which Deep Learning and Artificial Intelligence are now affecting computational protein studies is remarkable, but there are now many opportunities for further rapid progress with applications of these methods. Rapid gains are likely to come from studies using the approaches identified in this perspective. Addressing and predicting ligand-binding sites in protein structures, as well as the prediction of reliable structures of proteins interacting with other proteins, will be pivotal for fully details of structural mechanisms and dynamics. The prediction of multi-state protein ensembles, conformational transitions, dynamics of large protein complexes, and integration with experimental data is likely to happen quickly.

Pradeep Bk, Shi-Jie Chen, R. Dima et al. · 0 citations
Jul 2026

Unveiling Large-Scale Kinase-Centric Protein-Protein Interactions through a Knowledge-Informed Workflow.

Protein phosphorylation regulates signaling, yet atomic-level substrate specificity remains elusive due to sparse structural data and phosphorylation-site-insensitive deep-learning predictors. Here we present a pipeline reformulating kinase-substrate modeling as a Bayesian inference problem. By integrating curated data sets and literature evidence parsed by Large Language Models, we converted diverse biological knowledge into structural restraints for the restraint-guided deep-learning model GRASP. For EGFR, BRAF and JNK1, we obtained 336 new phosphorylation-site-specific structure candidates refined by molecular dynamics. These models recapitulate known features, such as JNK1's hydrophobic docking groove, and enabled a Virtual Position Scanning Peptide Array (V-PSPA) to map recognition patches and derive sequence preferences. Cross-referencing predicted interfaces with AlphaMissense pathogenicity scores reveal that the interaction types and distances to the catalytic pocket significantly influence pathogenicity scores. A comparison with clinical mutation data sets further connects pathogenic mutations to the kinase-substrate interface. This high-resolution, high-throughput pipeline can be broadly applicable to kinase specificity studies and general drug discovery.

Jinyuan Hu, Shimian Li, Yue Xue et al. · 0 citations