Aug 2026· Bioinformatics· Vol 42· 0 citations· 60 references
Medicine
TL;DR
Synthetic MSAs are suggested as a generalizable framework for dissecting the conformational landscapes encoded by deep learning structure predictors, with direct implications for understanding model behavior and accessing biologically relevant hidden states.
Abstract
Abstract Motivation Proteins rely on conformational flexibility for biological function, yet predicting alternative states remains a major challenge in structural biology. Although deep learning models like AlphaFold2, AlphaFold3, and RoseTTAFold2 excel at static structure prediction, what these networks actually learn about the underlying conformational landscapes remains largely opaque. Results Here we introduce synthetic multiple sequence alignments (MSAs), designed by inverse folding to encode predefined structural constraints, as a programmable intervention for interrogating the internal logic of structure prediction systems. Synthetic MSAs systematically bias AlphaFold2, AlphaFold3, and RoseTTAFold2 toward distinct conformational states of fold-switching proteins, including alternative conformations inaccessible through natural sequence information alone. Adversarial experiments pairing query sequences with MSAs encoding competing folds reveal sequence-dependent responses, exposing how alignment-derived and sequence-derived signals are weighted within each system. Probing predictions initialized from molecular dynamics trajectories reveals a systematic bias toward compact, training-distribution-favored conformations. Hybrid alignments combining synthetic and natural MSA segments enable targeted steering toward specific conformational states. These results suggest synthetic MSAs as a generalizable framework for dissecting the conformational landscapes encoded by deep learning structure predictors, with direct implications for understanding model behavior and accessing biologically relevant hidden states. Availability and implementation Newly generated data can be found at https://zenodo.org/records/20916910. The code underlying this article is available on GitHub at https://github.com/ibmm-unibe-ch/msa-tests.
A systematic NMR-characterized dataset of mutants of the GA/GB model fold-switching system is presented and it is found that this benchmark revealed variable and position-dependent performance across methods, with certain AlphaFold2-based algorithms able to predict mutant effects at individual sites, indicating some understanding of physical effects of residue substitutions.
Nathaniel R. Felbinger, K. Carillo, Yihong Chen et al.· bioRxiv· 0 citations
Protein inverse folding aims to recover amino acid sequences for a given 3D protein structure, underpinning broad applications such as enzyme engineering and drug discovery.Current methods often follow a serial pipeline, in which a structure encoder predicts a coarse sequence, which is then refined by protein language models (PLMs). However, because PLMs only perform post-hoc sequence edits, the refinement is bounded by the quality of upstream predictions.Thanks to recent multimodal protein language models (MPLMs), we could directly encode structure to generate sequences with pretrained structural knowledge, but we observe that they are not effective for inverse folding. Therefore, we introduce a symmetric dual-path architecture that both leverages PLMs for pretrained sequence evolution knowledge and MPLMs for pretrained structural knowledge to iteratively guide protein sequence generation.Through extensive experiments across standard protein inverse folding benchmarks, our method achieves state-of-the-art performance, surpassing prior approaches, and ablation studies validate the rationale of our symmetric design, revealing a promising direction for the community.
Han-Dong Wang, Jiaxin Qi, Baisheng Lai et al.· 0 citations
Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states, provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.
Hassan Nadeem, D. Kleiman, Yu-Ming Zhou et al.· bioRxiv· 0 citations
We survey modern deep‐learning approaches to protein conformational modeling through the lens of architectural design. We organize the literature into three increasingly expressive paradigms: (I) single‐structure prediction, (II) prediction of molecular binding complexes, and (III) conformational ensemble generation. For each paradigm, we outline a representative set of models to sketch a practical taxonomy, and we summarize their key achievements, limitations, and common evaluation practices. Across the paradigms, we highlight recurring design choices that shape performance and generalization, including enforced SE(3) equivariance versus learned symmetry; MSA‐driven coevolution versus protein language model priors; deterministic prediction versus generative sampling; explicit energetic supervision versus implicit learning; and integrative modeling across heterogeneous data modalities. While single‐structure prediction is now relatively well established, comparable maturity has not yet been reached for binding‐complex prediction and, especially, for generating faithful thermodynamic ensembles with reliable population weights, which remains an open challenge. We discuss open challenges in building physically grounded and transferable models, including data availability and fidelity, the choice of inductive biases to pursue generalization, and the need for rigorous model evaluation. Ultimately, we indicate generative kinetics as an aspirational frontier.
Daniele Angioletti, Matteo Carli, Marco S. Nobile et al.· WIREs Computational Molecula...· 0 citations
This framework provides a clearer understanding of how methodological shifts have shaped the capabilities, limitations, and practical roles of recent models.
Wengan He, Yongsheng Luo, Lihong Jiang et al.· 0 citations
UniStab is introduced, an end-to-end framework for predicting stability changes across all mutation types by leveraging the implicit geometric reasoning of a pre-trained folding model and demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels.
Hong Tan, Sheng-Geng Lin, Yi Xiong· Chemical Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.