Cold-Start Protein–Protein Interaction Prediction Is Bounded by the Frozen Embedding, Not the Classifier: A Leakage-Controlled Audit and a Calibrated Conformal Baseline
A many-objective NSGA-III search is built that attaches a compact head to a frozen ESM-2 t33 backbone and jointly optimizes the head architecture, a physicochemical feature mask, and a contrastive-pretraining curriculum against a fitness vector that explicitly contains the cold-start gap, and adds distribution-free calibrated selective prediction via split conformal, turning a frozen-ESM classifier into a deployable cold-start predictor with finite-sample coverage.
Abstract
Sequence-based protein–protein interaction (PPI) prediction has reached an empirical plateau on leakage-controlled benchmarks, where the dominant signal under the cold-start regime—predicting interactions among proteins absent from training—comes from evolutionary-scale protein language models (PLMs) rather than from the downstream classifier. A natural hypothesis is that elaborate head-only optimization over a frozen PLM can narrow the in-domain-to-cold-start gap cheaply. We test this hypothesis and report a negative-but-constructive result. We build a many-objective NSGA-III search that attaches a compact head to a frozen ESM-2 t33 backbone and jointly optimizes the head architecture, a physicochemical feature mask, and a contrastive-pretraining curriculum against a fitness vector that explicitly contains the cold-start gap; we then subject it to a controlled audit. On a leakage-controlled yeast DIP benchmark with Park–Marcotte stratification (homology-disjoint at 40% identity), the searched contrastive head, trained to a fair budget, reaches cold-start (C3) average precision (AUPR) 0.683, whereas an off-the-shelf random forest on the same embeddings reaches 0.734; gradient boosting and an RBF-SVM also exceed it. Seven design axes—objective, architecture, pair operator, embedding compression, pooling, training budget, and training-set size—each fail to close the gap, indicating the achievable level is set by the frozen representation, not the classifier. On a second, human benchmark (Bernett), simple classifiers again reach the published state-of-the-art range ( $\approx 0.65$ vs. 0.69 AUPR) at a fraction of the cost. Finally, we add distribution-free calibrated selective prediction via split conformal, turning a frozen-ESM classifier into a deployable cold-start predictor with finite-sample coverage. We release all code and reproduction drivers.
Frozen foundation-model embeddings are position as a strong, data-efficient default for affinity ranking in antibody engineering and establish a conservative lower bound that task-adaptive fine-tuning is expected to exceed.
Roger Wang, Kevin Jin, Lurong Pan· bioRxiv· 0 citations
A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.
Carl David Jasper Causin, M. Fyta· APL Machine Learning· 0 citations
S2DA-GO alleviates feature interference and improves prediction for sparsely annotated GO terms, providing a promising framework for large-scale annotation of uncharacterized proteins.
Hai-Long Wang, Fujun Xiang, Jin Zhang et al.· Frontiers in Genetics· 0 citations
A cheap density prior is introduced over natural protein activations and keeps only the candidates that remain typical under it, a training-free post-hoc step the authors call Mahalanobis filtering that improves both the property score and the structural plausibility of the sequences it keeps at negligible cost, without touching the generator, and transfers across different guidance methods.
Shuibai Zhang, Xin-Chi Liu, Fred Zhangzhi Peng et al.· 0 citations
Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding models (IFMs) upon fine-tuning, such as ThermoMPNN, achieved state-of-the-art performance in predicting thermostability changes in proteins caused by mutations. However, IFMs are limited in their capacity to capture protein deep evolutionary information, whereas protein language models (pLMs) excel. Here, we present SA-MPNN, a lightweight, end-to-end hybrid framework that dynamically integrates the protein sequence representations from a protein language model (ESM2) into the ThermoMPNN architecture to improve protein stability prediction by combining evolutionary representations with geometric structural embeddings. By evaluating various feature fusion strategies, we selected a self-attention-based integration mechanism to effectively combine the two modalities. Trained on the large-scale Megascale data set, SA-MPNN achieved modest but consistent gains over ThermoMPNN on various benchmark data sets, with particularly noticeable improvements in several correlation analysis and screening-oriented evaluations. Finally, wet-lab validation was performed on the top-ranking variants of Acetivibrio thermocellusβ-glucosidase (AtBgl1A) as a case study. The experimental results demonstrated that multiple designed mutants exhibited improved thermostability, and the optimal variant, GC20, achieved a melting temperature (Tm) of 76.98 °C, representing a 5.97 °C increase over the wild-type, thereby supporting the practical applicability of SA-MPNN in protein engineering.
Xin-Yue Zhang, Xiang Zheng, Ze-Yuan Dong et al.· Journal of Chemical Informat...· 0 citations
TEDlm variants also substantially improve zero-shot Molecular Function prediction over ESM2, while matching it on various biophysical property tasks, indicating that signals are largely domain-intrinsic.
Tiejun Wei, S. Kandathil, Daniel W. A. Buchan et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.