Skip to content

Author

Ren-Guang Zhu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Beyond random splits: A hierarchical benchmark of transferability and reliability in PROTAC activity prediction.

Computational prediction of PROTAC degradation activity (DC50) has attracted growing interest, yet the reliability of reported model performance remains poorly understood because sufficiently stringent evaluation protocols are rarely applied. Here, we present a hierarchical benchmark designed to expose evaluation pitfalls and quantify the transferability and reliability limits of current PROTAC predictors. Using a curated dataset of 2405 DC50 measurements spanning 22 target proteins and two E3 ligases (CRBN and VHL), we benchmarked classical machine learning (Random Forest, ExtraTrees, Ridge, PLS), gradient-boosted trees (XGBoost), nearest-neighbor retrieval baselines, protein negative controls, and a representative multi-modal deep learning ensemble (HybridMoECrossAttn) across Random, Scaffold, Leave-One-Target-Out (LOTO), and Leave-One-Family-Out (LOFO) splits. Under Random evaluation, a simple Random Forest + ECFP4 baseline achieved pooled R² = 0.693 ± 0.025, indicating that conventional models already approach the apparent ceiling under interpolation-oriented settings. However, all methods collapsed under LOTO (best R² = -0.012), revealing that much of the apparent progress in the literature reflects chemical-neighbor memorization rather than robust target-level generalization. We further show that target-wise error is significantly associated with continuous protein semantic proximity in ProtBERT space (Spearman ρ = -0.461, p = 0.047), whereas coarse family-level descriptors are uninformative. A four-quadrant failure taxonomy reveals that protein shift is more damaging than chemical novelty (MAE 1.09-1.12 vs. 0.85-0.97), and conformal prediction becomes severely overconfident under target extrapolation, with empirical 90% coverage dropping to 63.9-66.7%. These results reposition PROTAC prediction as a problem of transferability and reliability rather than leaderboard optimization and provide practical guidelines for future benchmark design.

Ren-Guang Zhu, Guang-Hao Guo, Lu-Lu Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.