Skip to content
Open access

Accurate ΔTm Prediction Without Protein Structure Inputs for Biomolecular Stability

Jul 2026 · bioRxiv · 0 citations · 18 references
Biology

TL;DR

It is shown that accurate ΔTm prediction does not require structural input features, but can achieve state-of-the-art results with a careful training design for large sequence-based protein language models, and choices such as loss function, pooling strategy, auxiliary supervision, and finetuning regime materially affect performance.

Abstract

Predicting protein stability, like changes in melting temperature (ΔTm) caused by mutations, is a critical task in therapeutic protein engineering and drug discovery. This is reflected by a growing solution space, including both AI-based sequence and structure based methods. This paper demonstrates that accurate ΔTm prediction does not require structural input features, but can achieve state-of-the-art results with a careful training design for large sequence-based protein language models. We combine an autoresearch-inspired setup search with controlled ablation studies and show that a well-tuned sequence-only ESM2-650M model [6] outperforms structure-informed methods in our benchmark, achieving the lowest error (MAE/RMSE) and competitive Pearson correlation without pH or structural inputs. We further show that choices such as loss function, pooling strategy, auxiliary supervision, and finetuning regime materially affect performance.

Read PDF

Similar papers

Aug 2026

SA-MPNN: A Sequence-Aware ThermoMPNN for Accurate Prediction of Mutational Effects on Protein Thermodynamic Stability

Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding models (IFMs) upon fine-tuning, such as ThermoMPNN, achieved state-of-the-art performance in predicting thermostability changes in proteins caused by mutations. However, IFMs are limited in their capacity to capture protein deep evolutionary information, whereas protein language models (pLMs) excel. Here, we present SA-MPNN, a lightweight, end-to-end hybrid framework that dynamically integrates the protein sequence representations from a protein language model (ESM2) into the ThermoMPNN architecture to improve protein stability prediction by combining evolutionary representations with geometric structural embeddings. By evaluating various feature fusion strategies, we selected a self-attention-based integration mechanism to effectively combine the two modalities. Trained on the large-scale Megascale data set, SA-MPNN achieved modest but consistent gains over ThermoMPNN on various benchmark data sets, with particularly noticeable improvements in several correlation analysis and screening-oriented evaluations. Finally, wet-lab validation was performed on the top-ranking variants of Acetivibrio thermocellusβ-glucosidase (AtBgl1A) as a case study. The experimental results demonstrated that multiple designed mutants exhibited improved thermostability, and the optimal variant, GC20, achieved a melting temperature (Tm) of 76.98 °C, representing a 5.97 °C increase over the wild-type, thereby supporting the practical applicability of SA-MPNN in protein engineering.

Xin-Yue Zhang, Xiang Zheng, Ze-Yuan Dong et al. · 0 citations
Open access Aug 2026

A unified predictor of protein stability changes across all mutation types via implicit structure learning

UniStab is introduced, an end-to-end framework for predicting stability changes across all mutation types by leveraging the implicit geometric reasoning of a pre-trained folding model and demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels.

Hong Tan, Sheng-Geng Lin, Yi Xiong · 0 citations
#artificial intelligence Preprint Sep 2026

SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding

Protein inverse folding aims to recover amino acid sequences for a given 3D protein structure, underpinning broad applications such as enzyme engineering and drug discovery.Current methods often follow a serial pipeline, in which a structure encoder predicts a coarse sequence, which is then refined by protein language models (PLMs). However, because PLMs only perform post-hoc sequence edits, the refinement is bounded by the quality of upstream predictions.Thanks to recent multimodal protein language models (MPLMs), we could directly encode structure to generate sequences with pretrained structural knowledge, but we observe that they are not effective for inverse folding. Therefore, we introduce a symmetric dual-path architecture that both leverages PLMs for pretrained sequence evolution knowledge and MPLMs for pretrained structural knowledge to iteratively guide protein sequence generation.Through extensive experiments across standard protein inverse folding benchmarks, our method achieves state-of-the-art performance, surpassing prior approaches, and ablation studies validate the rationale of our symmetric design, revealing a promising direction for the community.

Han-Dong Wang, Jiaxin Qi, Baisheng Lai et al. · 0 citations
Open access Jul 2026

Capabilities, specificity gaps and training-data dependence of AlphaFold3 across diverse application areas

It is found that, while AF3 can perform well in favourable settings, this performance is uneven across applications and its predictions and use of confidence metrics will depend strongly on the specific application area and must be interpreted with respect to training-set overlap.

O. Follonier, Yan Liu, Pablo Campomanes et al. · 1 citation
Open access Jul 2026

Benchmarking AI Protein Structure Predictors Reveals a Persistent Bias in Multi-State Proteins

Protein structure predictors achieve high single-state accuracy, but it remains unclear whether they can recover functionally relevant conformational ensembles or account for the presence of ligands and/or binding partners. Here, we benchmark AlphaFold3, Boltz-2, Chai-1, and BioEmu on four canonical multi-state proteins (Pf-MATE, LAO, SecA, and β2AR), quantifying state bias and sampling breadth against experimental reference structures. Models frequently default to a dominant state represented in the PDB; small-molecule ligands have weak or inconsistent effects, while large protein partners drive clear conformational switching between states. Multiple sequence alignment (MSA)-based approaches (AF-Cluster and random subsampling) recapitulate similar biases, indicating that this behavior is not unique to newer architectures. These results underscore current limitations for multi-state protein structure prediction and structure-guided ligand discovery. TOC Graphic

Muhui Ye, Yu-Hong Wang, M. Brogi et al. · 0 citations
#protein folding Open access Aug 2026

Protein language model-generated enzyme sequences exhibit high stability in molecular dynamics simulations

Large Language Models (LLMs) have transformed protein engineering by capturing complex sequence patterns from large datasets, enabling applications such as structure prediction and functional annotation. Finenzyme applies conditional transfer learning to generate biologically plausible enzyme sequences conditioned on Enzyme Commission (EC) numbers. In this work, we extended Finenzyme with an in silico selection pipeline that first identifies generated sequences most likely to preserve or enhance the functional characteristics of specific EC categories and then evaluates them through molecular dynamics (MD) simulations to assess their structural stability and conformational dynamics. MD simulations of 236 Finenzyme-generated enzymes across four EC classes (59 µ s total simulation time) confirmed high structural stability. Across all enzyme classes, 74-95% of the models maintained stable tertiary structures and correct folding throughout the trajectories, with 195 out of 236 structures (82.6%) exhibiting sustained stability. By combining conditional pre-trained language model fine-tuning with dynamic structural evaluation, our framework advances beyond static sequence-based predictions to address the structural and functional dimensions of enzyme behavior, key aspects for both biomedical and industrial applications.

E. M. Fassi, M. Nicolini, Emanuele Saitto et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.