Skip to content

Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction

Mar 2026 · arXiv.org · Vol abs/2603.29529 · 0 citations · 35 references
Computer Science Physics Biology

TL;DR

It is found that, at variance with networks not based on the attention mechanism, the lack of a first--order--like transition in the loss of the transformer produces a range of intermediate temperatures with good learning properties; this is true both for synthetic and natural protein sequences.

Abstract

Using a statistical mechanics framework, we investigate the parameter space of transformer models trained on protein sequence data. We sample the loss landscape at varying temperatures using Langevin dynamics to characterize the low-loss manifold, and to understand the mechanisms underlying transformers'superior performance in protein structure prediction. We find that, at variance with networks not based on the attention mechanism, the lack of a first--order--like transition in the loss of the transformer produces a range of intermediate temperatures with good learning properties; this is true both for synthetic and natural protein sequences. We also show that the parameters of most layers are highly conserved at these temperatures if the dimension of the embedding is optimal, and we provide an operative way to find this dimension. Additionally, we show that the attention matrix is more predictive of the contact maps of the protein at higher temperatures and for higher dimensions of the embedding than those optimal for learning. Finally, we showed that the models sampled at intermediate temperatures can predict the free-energy variation upon mutation, better than models obtained through standard optimization techniques.

View source

Similar papers

#protein folding Preprint Aug 2026

Off-Manifold Collapse in Guided Protein Language Models

A cheap density prior is introduced over natural protein activations and keeps only the candidates that remain typical under it, a training-free post-hoc step the authors call Mahalanobis filtering that improves both the property score and the structural plausibility of the sequences it keeps at negligible cost, withou...

Shuibai Zhang, Xin-Chi Liu, Fred Zhangzhi Peng et al. · 0 citations
#machine learning Preprint Aug 2026

Task- and dataset-specific information in protein language models

Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs)....

R. Joeres, Ilya S. Senatorov, A. Kolchina et al. · 0 citations
Preprint Aug 2026

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

It is shown that PLM embeddings encode patterns correlated with biochemical properties and quantify their contribution to predicting protein fitness, and this technique is readily transferable to problem settings beyond protein fitness prediction.

Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer et al. · 0 citations
Aug 2026

SA-MPNN: A Sequence-Aware ThermoMPNN for Accurate Prediction of Mutational Effects on Protein Thermodynamic Stability.

Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding mo...

Xin-Yue Zhang, Xiang-Shan Zheng, Ze-Yuan Dong et al. · 0 citations
Open access Aug 2026

Aligning protein-generative models to experimental fitness with ProteinDPO

This work demonstrates how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences.

Talal Widatalla, Ashir Borah, Samuel H. King et al. · 3 citations
Open access Sep 2026

Modeling Protein Sequence Evolution as an Ornstein-Uhlenbeck Process in a Latent Space

An unsupervised inference model that integrates directed-evolution sequencing time series with natural homologs is presented and the inferred couplings improve structural contact prediction by combining global evolutionary constraints from nature with local, experiment-specific signals.

Matteo De Leonardis, Andrea Pagnani · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.