Skip to content
Open access

Aligning protein-generative models to experimental fitness with ProteinDPO

Aug 2026 · Nature Methods · Vol 23, pp. 1805 - 1813 · 1 citation · 72 references
Medicine

TL;DR

This work demonstrates how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences.

Abstract

Biological generative models can predict biological functions without task-specific training data but often under-perform specialized models. This is due to a fundamental ‘alignment gap’, where the rules learned during unsupervised training are not related to the function of interest. Here we demonstrate how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences. Our aligned model, ProteinDPO, achieves stability prediction competitive to task-specific models and consistently outperforms unsupervised and fine-tuned versions of the model. Notably, ProteinDPO generalizes beyond its training data to enable stabilization and improved binding affinity prediction of large multichain protein complexes. When applied to stabilization of the hemagglutinin trimer, a primary component of influenza vaccines, ~80% of designs achieve increased or similar stability compared with the native hemagglutinin and up to 32 °C improvements from recently emerged mammalian strains. Our results demonstrate how to augment generative models with biophysical information and, more broadly, provide a general framework for the alignment of biological foundation models. This Article demonstrates that direct preference optimization (DPO) can be used to effectively align an unsupervised structure-conditioned language model with biophysical information. The aligned model, ProteinDPO, achieves stability prediction competitive with that of task-specific models and consistently outperforms unsupervised and fine-tuned versions of the model.

Read PDF

Similar papers

Jul 2026

Exploring the Alignment of Generation and Understanding in Protein Structure Modeling

Understanding and generation are often treated as two separate paradigms in training deep neural networks, despite the fact that both are trained with closely related objectives such as denoising and masked prediction. While prior studies have shown that generative models often learn suboptimal representations for understanding tasks in vision, it is less understood whether a similar gap exists in the protein domain. In this work, we systematically investigate this question by benchmarking state-of-the-art protein generative models on widely-used protein understanding tasks, and observe that these models exhibit consistently poor performance compared to existing protein encoders. Furthermore, inspired by the Representation Alignment (REPA) framework, we propose to explicitly align generative protein diffusion models with pretrained protein understanding models during training. Experiments on the MotifBench demonstrate that representation alignment significantly improves functional protein generation, boosting the MotifBench score of Protpardelle-1c from 39.2 to 47.1, corresponding to a 20% relative improvement. Our results suggest that representation alignment provides a general and effective mechanism for bridging understanding and generation in protein structure modeling.

Junde Xu, Yuansheng Huang, Zijun Gao et al. · 0 citations
Jul 2026

Property guidance for protein sequence generative models with ProteinGuide.

ProteinGuide is applied jointly with wet-lab data generation to increase the editing activity of an adenine base editor in vivo, resulting in a base editor with higher editing efficiency than was previously achieved using seven rounds of directed evolution.

Junhao Xiong, Ishan Gaur, Maria Lukarska et al. · 1 citation
Open access Jul 2026

Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer.

Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.

Yo Akiyama, Zhidian Zhang, Olivia Tang et al. · 3 citations
Open access Aug 2026

MAXWELL: Calibrating the probabilistic outputs of protein language models to the mutation-induced stability change landscape

MAXWELL (Matrix-wise Landscape Learning), a novel post-training method that calibrates the probabilistic outputs learned by protein language models during pretraining to generate mutational landscapes that quantify the effects of individual amino acid substitutions on protein stability, is introduced.

Ming-Chen Li, Xiaoran Cheng, Fan Jiang et al. · 0 citations
Open access Jul 2026

ProteinDock: A physics-informed layer to improve protein-protein docking reliability

It is demonstrated that a truncated version of ProteinDock can be used to choose the optimal prediction among outputs from multiple deep learning-based tools, and shown that this strategy is a computationally efficient alternative to increasing the seed quantity for deep-learning predictions.

G. Rajagopal, Søren C. Spina, Joe Bailey et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.