Jul 2026· Proceedings of the National Academy of Sciences of the United States of America· Vol 123 29, pp.
e2609610123
· 0 citations· 55 references
Medicine
TL;DR
While AI-based tools perform exceptionally on natural sequences, they do not reliably encode the physico-chemical principles governing ionizable residue placement, and it is proposed to include brief molecular dynamics simulations as a vital validation step for AI-generated structures.
Abstract
A variant of the U1A protein containing four substitutions to ionizable residues was generated serendipitously due to a miscommunication. Biophysical measurements reveal this variant has twice the helical structure of wild-type U1A and is trimeric, unlike the monomeric wild type. In sharp contrast, structures predicted by deep-learning (AlphaFold2, RoseTTAFold2) and transformer-based tools (OmegaFold, ESMFold) are nearly identical to the wild-type (backbone RMSD < 1 Å). Surprisingly, these models predict ionizable residues buried within the nonpolar core, contradicting established physico-chemical principles. To explore this effect further, we generated sequences containing up to all twelve residues that make up the nonpolar core of U1A. Across thousands of sequences, and depending on the AI model used, the majority of predicted structures contained fully buried ionizable residues while still maintaining the overall U1A fold. We then examined two additional proteins of comparable size, acylphosphatase and the de novo designed TOP7 fold, and observed the same phenomenon: AI models frequently predicted structures with buried ionizable residues that nevertheless retained the parent fold. However, short (50 ns) molecular dynamics simulations with physics-based force fields (CHARMM/AMBER) rapidly relaxed these structures, exposing the ionizable residues. We conclude that while AI-based tools perform exceptionally on natural sequences, they do not reliably encode the physico-chemical principles governing ionizable residue placement. We propose including brief molecular dynamics simulations as a vital validation step for AI-generated structures.
Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.
Pascal Sturmfels, Naozumi Hiranuma, Milad Salem et al.· arXiv.org· 0 citations
UniStab is introduced, an end-to-end framework for predicting stability changes across all mutation types by leveraging the implicit geometric reasoning of a pre-trained folding model and demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels.
Hong Tan, Sheng-Geng Lin, Yi Xiong· Chemical Science· 0 citations
Deep learning-based protein structure prediction methods that leverage evolutionary information from multiple sequence alignments (MSAs), exemplified by AlphaFold2, have achieved remarkable accuracy. However, existing methods still struggle to predict challenging proteins, particularly those with novel folds or limited evolutionary information, and to recover alternative conformational states. Here we show that structure prediction models trained under different MSA-depth distributions corresponding to different levels of evolutionary information exhibit complementary generalization behaviors, and that a model trained on a mixture of these distributions can combine their complementary generalization strengths. Building on this insight, we developed ProtMonomer, a deep learning framework trained on MSA-depth distributions representing a broad range of evolutionary information levels to improve structure prediction. Across benchmarks comprising CASP15 targets, non-redundant experimentally determined structures, orphan proteins, and short peptides, ProtMonomer performed comparably to or better than leading methods, including AlphaFold2 and AlphaFold3, with particularly strong performance on challenging targets. For fold-switching proteins, ProtMonomer also recovered alternative conformational states more accurately than AlphaFold2 and AlphaFold3 across diverse homologous sequence sampling strategies. In addition to improving predictive accuracy, ProtMonomer substantially reduced inference cost through an efficient architecture, enabling high-throughput applications. Together, these findings provide insights into the generalization of evolution-informed structure prediction models and support ProtMonomer as an accurate and efficient framework for protein structure prediction.
Modelling sequences both “dry” and in the presence of explicit potassium cations are suggested as a simple, practical way to sample alternative conformations and to expose disordered regions that current predictors tend to over-fold.
It is found that, while AF3 can perform well in favourable settings, this performance is uneven across applications and its predictions and use of confidence metrics will depend strongly on the specific application area and must be interpreted with respect to training-set overlap.
O. Follonier, Yan Liu, Pablo Campomanes et al.· bioRxiv· 1 citation
A systematic NMR-characterized dataset of mutants of the GA/GB model fold-switching system is presented and it is found that this benchmark revealed variable and position-dependent performance across methods, with certain AlphaFold2-based algorithms able to predict mutant effects at individual sites, indicating some understanding of physical effects of residue substitutions.
Nathaniel R. Felbinger, K. Carillo, Yihong Chen et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.