Skip to content
Book Open access

From Structure to Function: Preference Alignment for Function-aware Protein Inverse Folding

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 12066-12077 · 0 citations · 11 references

Abstract

Protein inverse folding models conditioned on structure achieve high sequence recovery but often fail to preserve biological function due to the lack of functional supervision. We propose a function-aware preference alignment framework that improves functional preservation by fine-tuning models to favor function-preserving sequences over function-disrupting alternatives, avoiding the need for explicit function optimization. Our approach constructs reliable preference pairs in silico using hypothesis-driven perturbations of critical residues and model-consistent likelihood constraints, enabling scalable supervision without additional wet-lab measurements. The resulting framework guides protein sequence design models toward generating sequences that better preserve functional integrity, while remaining compatible with existing inverse folding pipelines such as ProteinMPNN and ESM-IF. Extensive experiments on protein design benchmarks and enzyme datasets with established wet-lab validation show that our fine-tuned models consistently outperform pretrained counterparts in preserving functional integrity during protein sequence design. The code is available at https://github.com/EvaFlower/Function-aware-Protein-Inverse-Folding

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding

Protein inverse folding aims to recover amino acid sequences for a given 3D protein structure, underpinning broad applications such as enzyme engineering and drug discovery.Current methods often follow a serial pipeline, in which a structure encoder predicts a coarse sequence, which is then refined by protein language models (PLMs). However, because PLMs only perform post-hoc sequence edits, the refinement is bounded by the quality of upstream predictions.Thanks to recent multimodal protein language models (MPLMs), we could directly encode structure to generate sequences with pretrained structural knowledge, but we observe that they are not effective for inverse folding. Therefore, we introduce a symmetric dual-path architecture that both leverages PLMs for pretrained sequence evolution knowledge and MPLMs for pretrained structural knowledge to iteratively guide protein sequence generation.Through extensive experiments across standard protein inverse folding benchmarks, our method achieves state-of-the-art performance, surpassing prior approaches, and ablation studies validate the rationale of our symmetric design, revealing a promising direction for the community.

Han-Dong Wang, Jiaxin Qi, Baisheng Lai et al. · 0 citations

TTS-Design: Test-Time Compute Scaling for Structure-Guided Protein Design

TTS-Design is proposed, a test-time compute scaling framework that enhances protein sequence design without retraining models or relying on larger training datasets, and can consistently improve sequence recovery and structural reliability across different backbone models, without retraining or increasing model size.

Zizhe Jin, Yi Zheng, Huan Yee Koh et al. · 0 citations
Jul 2026

EnerBridge-DPO: Energy-Aware Markov Bridge Inverse Folding for Protein Sequence Design

Designing protein sequences with favorable predicted energetic properties is an important challenge in protein inverse folding, because many existing deep learning methods are primarily trained by maximizing sequence recovery and do not explicitly incorporate energy-related preferences during generation. In this work, we propose EnerBridge-DPO, an energy-aware inverse folding framework that integrates Markov bridge sequence generation with preference optimization for protein complex design. The framework builds on the Markov bridge inverse-folding process to generate structure-compatible sequences from an informative prior sequence. It then introduces a Bridge-DPO objective that uses energy-related winner-loser preference pairs to bias the generator toward sequences favored by computational or experimental energy-related signals. In addition, we incorporate a quantitative energy-constrained loss based on mutation-induced binding free-energy changes to provide continuous ΔΔG supervision. Evaluations show that EnerBridge-DPO maintains competitive inverse-folding performance while obtaining lower predicted energy scores under selected computational scoring functions for protein complexes. On SKEMPI, EnerBridge-DPO achieves competitive ΔΔG prediction performance, with small numerical gains in several overall metrics that are not statistically conclusive under paired bootstrap analysis. These results suggest that incorporating energy-related preferences into Markov bridge inverse folding can improve computationally predicted energetic profiles, although experimental validation is required to confirm thermodynamic stability.

Dingyi Rong, Haotian Lu, Xupeng Zhang et al. · 0 citations
Open access Aug 2026

Blending physics-based and inverse folding models to disentangle variant effects on stability and function

It is shown that blending IF models with a physics-based coarse-grained potential improves global correlation with experimental ΔΔG and, crucially, reduces IF model bias at functional sites, and is found that disease gain-of-function variants show a distinct functional signature from loss-of-function variants.

Ezequiel A. Galpern, Xavier Soler Sanchis, Charles W. J. Pugh et al. · 0 citations
#protein folding Open access Sep 2026

Inverse FoldDir: Structure-conditioned Protein Sequence Design by Dirichlet Flow Matching

Inverse FoldDir is a structure-conditioned protein redesign method that combines structural recovery, user control, experimental validation, and a natural route toward future property-guided sampling that performs iterative denoising on the amino acid probability simplex.

Alp Tartici, M. Stojkovic, An-Ru Tian et al. · 0 citations
Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.