Skip to content
Open access

Data-Centric Evaluation of Protein Function Prediction Pipelines

Aug 2026 · bioRxiv · 0 citations · 49 references
Biology

TL;DR

Findings show that performance estimates in protein function prediction should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

Abstract

Performance estimates in protein function prediction depend not only on model choice but also on upstream decisions that define the learning problem. Using antioxidant protein classification as a controlled case study, we evaluated how dataset harmonisation, protein representation, redundancy control, and partitioning strategy affect protein machine learning pipelines. We integrated 18,804 records from 12 publicly available dataset entries into a curated consensus dataset of 4,193 protein sequences. One-hot encoding and six pretrained protein language model representations were evaluated as model inputs and as similarity spaces for redundancy reduction and distance-aware splitting. Representation choice substantially altered dataset geometry, retained dataset size, class balance, and downstream evaluation. At representation-specific p90 thresholds, one-hot encoding retained the complete dataset, whereas pretrained embeddings retained between 5% and 25% of sequences. Distance-aware partitioning reduced apparent performance relative to random splitting by up to 0.15 MCC before redundancy control, while this difference narrowed after similarity filtering. Selected configurations nevertheless maintained high performance under stricter evaluation, reaching an MCC of 0.84. These findings show that performance estimates should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

Read PDF

Similar papers

Aug 2026

Beyond random splits: A hierarchical benchmark of transferability and reliability in PROTAC activity prediction.

Computational prediction of PROTAC degradation activity (DC50) has attracted growing interest, yet the reliability of reported model performance remains poorly understood because sufficiently stringent evaluation protocols are rarely applied. Here, we present a hierarchical benchmark designed to expose evaluation pitfalls and quantify the transferability and reliability limits of current PROTAC predictors. Using a curated dataset of 2405 DC50 measurements spanning 22 target proteins and two E3 ligases (CRBN and VHL), we benchmarked classical machine learning (Random Forest, ExtraTrees, Ridge, PLS), gradient-boosted trees (XGBoost), nearest-neighbor retrieval baselines, protein negative controls, and a representative multi-modal deep learning ensemble (HybridMoECrossAttn) across Random, Scaffold, Leave-One-Target-Out (LOTO), and Leave-One-Family-Out (LOFO) splits. Under Random evaluation, a simple Random Forest + ECFP4 baseline achieved pooled R² = 0.693 ± 0.025, indicating that conventional models already approach the apparent ceiling under interpolation-oriented settings. However, all methods collapsed under LOTO (best R² = -0.012), revealing that much of the apparent progress in the literature reflects chemical-neighbor memorization rather than robust target-level generalization. We further show that target-wise error is significantly associated with continuous protein semantic proximity in ProtBERT space (Spearman ρ = -0.461, p = 0.047), whereas coarse family-level descriptors are uninformative. A four-quadrant failure taxonomy reveals that protein shift is more damaging than chemical novelty (MAE 1.09-1.12 vs. 0.85-0.97), and conformal prediction becomes severely overconfident under target extrapolation, with empirical 90% coverage dropping to 63.9-66.7%. These results reposition PROTAC prediction as a problem of transferability and reliability rather than leaderboard optimization and provide practical guidelines for future benchmark design.

Ren-Guang Zhu, Guang-Hao Guo, Lu-Lu Li et al. · 0 citations
Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations
Open access Aug 2026

A Two-Stage ESM-Based Machine Learning Pipeline for Robust Hierarchical Enzyme Function Prediction

Results support the use of pretrained protein language model embeddings as an effective foundation for enzyme annotation by combining large-scale sequence representations with a lightweight supervised classifier and may facilitate functional annotation of protein sequences derived from large genomic and metagenomic datasets.

Xiao Hua, G. Grimaud · 0 citations
Open access Aug 2026

DHST: A Deep Hybrid Structure–Topology Framework for Accurate Protein Function Prediction

DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings.

Bin Lu, Fujun Xiang, Hai-Long Wang et al. · 0 citations
Open access Aug 2026

FuncSeek: Multi-PLM contrastive learning for protein functional similarity search

FuncSeek is described, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities) that each capture a different aspect of protein biology: evolutionary patterns, three-dimensional shape, and functional context.

Leendert J. Cloete, Hugh G. Patterton · 0 citations
Open access Aug 2026

Sequence-centric deep learning druggability prediction using protein language models with multi-scale attention and feature fusion

Overall, DrugPLMFormer provides a reproducible, leakage-aware framework for retrospective sequence-based druggability screening and target prioritization, while prospective validation and experimental confirmation remain necessary before operational deployment.

Z. Kafi, Khosro Rezaee, Hossein Eslami · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.