Skip to content
Open access

Cold-Start Protein–Protein Interaction Prediction Is Bounded by the Frozen Embedding, Not the Classifier: A Leakage-Controlled Audit and a Calibrated Conformal Baseline

2026 · IEEE Access · Vol 14, pp. 112470-112480 · 0 citations · 29 references
Computer Science

TL;DR

A many-objective NSGA-III search is built that attaches a compact head to a frozen ESM-2 t33 backbone and jointly optimizes the head architecture, a physicochemical feature mask, and a contrastive-pretraining curriculum against a fitness vector that explicitly contains the cold-start gap, and adds distribution-free calibrated selective prediction via split conformal, turning a frozen-ESM classifier into a deployable cold-start predictor with finite-sample coverage.

Abstract

Sequence-based protein–protein interaction (PPI) prediction has reached an empirical plateau on leakage-controlled benchmarks, where the dominant signal under the cold-start regime—predicting interactions among proteins absent from training—comes from evolutionary-scale protein language models (PLMs) rather than from the downstream classifier. A natural hypothesis is that elaborate head-only optimization over a frozen PLM can narrow the in-domain-to-cold-start gap cheaply. We test this hypothesis and report a negative-but-constructive result. We build a many-objective NSGA-III search that attaches a compact head to a frozen ESM-2 t33 backbone and jointly optimizes the head architecture, a physicochemical feature mask, and a contrastive-pretraining curriculum against a fitness vector that explicitly contains the cold-start gap; we then subject it to a controlled audit. On a leakage-controlled yeast DIP benchmark with Park–Marcotte stratification (homology-disjoint at 40% identity), the searched contrastive head, trained to a fair budget, reaches cold-start (C3) average precision (AUPR) 0.683, whereas an off-the-shelf random forest on the same embeddings reaches 0.734; gradient boosting and an RBF-SVM also exceed it. Seven design axes—objective, architecture, pair operator, embedding compression, pooling, training budget, and training-set size—each fail to close the gap, indicating the achievable level is set by the frozen representation, not the classifier. On a second, human benchmark (Bernett), simple classifiers again reach the published state-of-the-art range ( $\approx 0.65$ vs. 0.69 AUPR) at a fraction of the cost. Finally, we add distribution-free calibrated selective prediction via split conformal, turning a frozen-ESM classifier into a deployable cold-start predictor with finite-sample coverage. We release all code and reproduction drivers.

Read PDF

Similar papers

Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
#protein folding Preprint Aug 2026

Off-Manifold Collapse in Guided Protein Language Models

A cheap density prior is introduced over natural protein activations and keeps only the candidates that remain typical under it, a training-free post-hoc step the authors call Mahalanobis filtering that improves both the property score and the structural plausibility of the sequences it keeps at negligible cost, without touching the generator, and transfers across different guidance methods.

Shuibai Zhang, Xin-Chi Liu, Fred Zhangzhi Peng et al. · 0 citations
Aug 2026

SA-MPNN: A Sequence-Aware ThermoMPNN for Accurate Prediction of Mutational Effects on Protein Thermodynamic Stability

Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding models (IFMs) upon fine-tuning, such as ThermoMPNN, achieved state-of-the-art performance in predicting thermostability changes in proteins caused by mutations. However, IFMs are limited in their capacity to capture protein deep evolutionary information, whereas protein language models (pLMs) excel. Here, we present SA-MPNN, a lightweight, end-to-end hybrid framework that dynamically integrates the protein sequence representations from a protein language model (ESM2) into the ThermoMPNN architecture to improve protein stability prediction by combining evolutionary representations with geometric structural embeddings. By evaluating various feature fusion strategies, we selected a self-attention-based integration mechanism to effectively combine the two modalities. Trained on the large-scale Megascale data set, SA-MPNN achieved modest but consistent gains over ThermoMPNN on various benchmark data sets, with particularly noticeable improvements in several correlation analysis and screening-oriented evaluations. Finally, wet-lab validation was performed on the top-ranking variants of Acetivibrio thermocellusβ-glucosidase (AtBgl1A) as a case study. The experimental results demonstrated that multiple designed mutants exhibited improved thermostability, and the optimal variant, GC20, achieved a melting temperature (Tm) of 76.98 °C, representing a 5.97 °C increase over the wild-type, thereby supporting the practical applicability of SA-MPNN in protein engineering.

Xin-Yue Zhang, Xiang Zheng, Ze-Yuan Dong et al. · 0 citations
Open access Jul 2026

TEDlm: domain-centric protein language models with optional structural pre-training

TEDlm variants also substantially improve zero-shot Molecular Function prediction over ESM2, while matching it on various biophysical property tasks, indicating that signals are largely domain-intrinsic.

Tiejun Wei, S. Kandathil, Daniel W. A. Buchan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.