Skip to content
Open access

The accuracy of electrostatic interactions captured by AI protein structure prediction models.

Jul 2026 · Proceedings of the National Academy of Sciences of the United States of America · Vol 123 29, pp. e2609610123 · 0 citations · 55 references
Medicine

TL;DR

While AI-based tools perform exceptionally on natural sequences, they do not reliably encode the physico-chemical principles governing ionizable residue placement, and it is proposed to include brief molecular dynamics simulations as a vital validation step for AI-generated structures.

Abstract

A variant of the U1A protein containing four substitutions to ionizable residues was generated serendipitously due to a miscommunication. Biophysical measurements reveal this variant has twice the helical structure of wild-type U1A and is trimeric, unlike the monomeric wild type. In sharp contrast, structures predicted by deep-learning (AlphaFold2, RoseTTAFold2) and transformer-based tools (OmegaFold, ESMFold) are nearly identical to the wild-type (backbone RMSD < 1 Å). Surprisingly, these models predict ionizable residues buried within the nonpolar core, contradicting established physico-chemical principles. To explore this effect further, we generated sequences containing up to all twelve residues that make up the nonpolar core of U1A. Across thousands of sequences, and depending on the AI model used, the majority of predicted structures contained fully buried ionizable residues while still maintaining the overall U1A fold. We then examined two additional proteins of comparable size, acylphosphatase and the de novo designed TOP7 fold, and observed the same phenomenon: AI models frequently predicted structures with buried ionizable residues that nevertheless retained the parent fold. However, short (50 ns) molecular dynamics simulations with physics-based force fields (CHARMM/AMBER) rapidly relaxed these structures, exposing the ionizable residues. We conclude that while AI-based tools perform exceptionally on natural sequences, they do not reliably encode the physico-chemical principles governing ionizable residue placement. We propose including brief molecular dynamics simulations as a vital validation step for AI-generated structures.

Read PDF

Similar papers

Jul 2026

Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.

Pascal Sturmfels, Naozumi Hiranuma, Milad Salem et al. · 0 citations
Open access Aug 2026

A unified predictor of protein stability changes across all mutation types via implicit structure learning

UniStab is introduced, an end-to-end framework for predicting stability changes across all mutation types by leveraging the implicit geometric reasoning of a pre-trained folding model and demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels.

Hong Tan, Sheng-Geng Lin, Yi Xiong · 0 citations
#protein folding Open access Aug 2026

Accurate and efficient prediction of protein conformations with ProtMonomer

Deep learning-based protein structure prediction methods that leverage evolutionary information from multiple sequence alignments (MSAs), exemplified by AlphaFold2, have achieved remarkable accuracy. However, existing methods still struggle to predict challenging proteins, particularly those with novel folds or limited evolutionary information, and to recover alternative conformational states. Here we show that structure prediction models trained under different MSA-depth distributions corresponding to different levels of evolutionary information exhibit complementary generalization behaviors, and that a model trained on a mixture of these distributions can combine their complementary generalization strengths. Building on this insight, we developed ProtMonomer, a deep learning framework trained on MSA-depth distributions representing a broad range of evolutionary information levels to improve structure prediction. Across benchmarks comprising CASP15 targets, non-redundant experimentally determined structures, orphan proteins, and short peptides, ProtMonomer performed comparably to or better than leading methods, including AlphaFold2 and AlphaFold3, with particularly strong performance on challenging targets. For fold-switching proteins, ProtMonomer also recovered alternative conformational states more accurately than AlphaFold2 and AlphaFold3 across diverse homologous sequence sampling strategies. In addition to improving predictive accuracy, ProtMonomer substantially reduced inference cost through an efficient architecture, enabling high-throughput applications. Together, these findings provide insights into the generalization of evolution-informed structure prediction models and support ProtMonomer as an accurate and efficient framework for protein structure prediction.

Yunda Si, Suqi Zhang, Luo-Nan Chen · 0 citations
Jul 2026

Deep-learning predictions of biomolecular structures : persistent limitations and new horizons extended by explicit ion addition

Modelling sequences both “dry” and in the presence of explicit potassium cations are suggested as a simple, practical way to sample alternative conformations and to expose disordered regions that current predictors tend to over-fold.

Jules Marien, Sujith Sritharan, Beatrice Caviglia et al. · 0 citations
Open access Jul 2026

Capabilities, specificity gaps and training-data dependence of AlphaFold3 across diverse application areas

It is found that, while AF3 can perform well in favourable settings, this performance is uneven across applications and its predictions and use of confidence metrics will depend strongly on the specific application area and must be interpreted with respect to training-set overlap.

O. Follonier, Yan Liu, Pablo Campomanes et al. · 1 citation
Open access Aug 2026

Benchmarking Deep Learning Predictions of Mutation-Induced Fold Switching

A systematic NMR-characterized dataset of mutants of the GA/GB model fold-switching system is presented and it is found that this benchmark revealed variable and position-dependent performance across methods, with certain AlphaFold2-based algorithms able to predict mutant effects at individual sites, indicating some understanding of physical effects of residue substitutions.

Nathaniel R. Felbinger, K. Carillo, Yihong Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.