Skip to content

Comprehensive Assessment and Benchmark of Deep Generative Models for Proteolysis TArgeting Chimera (PROTAC) Design

May 2026 · Journal of Chemical Information and Modeling · Vol 66 9, pp. 5301-5314 · 0 citations
Computer Science Medicine

TL;DR

This work aims to discuss the features and the generative performance of different types of molecular generative models for the PROTAC design task and help researchers to better apply these models in practical cases.

View source

Similar papers

Open access Jul 2026

From Ternary Modeling to Predictive PROTAC Design: A Computational Perspective

Proteolysis-targeting chimeras (PROTACs) are heterobifunctional small molecules that induce targeted protein degradation by recruiting an E3 ligase to a protein of interest. Since 2019, publication volume has accelerated, and computational methods have expanded from isolated demonstrations into practical tools for modeling PROTAC-induced ternary complexes, designing linkers, and forecasting degradation-related outcomes. Here, we present a Perspective on computational PROTAC methodologies published from 2019 to the present, organizing the field into two complementary streams: (i) constraint-driven, physics-based workflows that assemble and refine ternary complex models by enforcing geometric feasibility and evaluating pose stability using docking and molecular simulation; and (ii) data-driven workflows, including deep learning predictors and generative models that predict ternary complex structure, degradation end points, or linker chemistry from structural and assay data. We highlight representative approaches spanning restrained/tethered docking, MD-based refinement and dynamic stability scoring, coarse-grained free-energy modeling, SE(3)/E(3)-equivariant structure prediction, supervised degradation efficacy prediction, and generative linker design. We close by emphasizing persistent gaps, fragmented benchmarking, score robustness across targets and E3 ligases, and nonstandard molecular representations that currently limit generalization and reproducible, pipeline-ready deployment.

Joseph M. Schulz, R. Reynolds, Stephan C. Schürer · 0 citations
Open access Jul 2026

Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization

A machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round that may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants.

Li Wan, Mahdi Bagherpoor Helabad, Lena Fraedrich et al. · 0 citations
Aug 2026

Beyond random splits: A hierarchical benchmark of transferability and reliability in PROTAC activity prediction.

Computational prediction of PROTAC degradation activity (DC50) has attracted growing interest, yet the reliability of reported model performance remains poorly understood because sufficiently stringent evaluation protocols are rarely applied. Here, we present a hierarchical benchmark designed to expose evaluation pitfalls and quantify the transferability and reliability limits of current PROTAC predictors. Using a curated dataset of 2405 DC50 measurements spanning 22 target proteins and two E3 ligases (CRBN and VHL), we benchmarked classical machine learning (Random Forest, ExtraTrees, Ridge, PLS), gradient-boosted trees (XGBoost), nearest-neighbor retrieval baselines, protein negative controls, and a representative multi-modal deep learning ensemble (HybridMoECrossAttn) across Random, Scaffold, Leave-One-Target-Out (LOTO), and Leave-One-Family-Out (LOFO) splits. Under Random evaluation, a simple Random Forest + ECFP4 baseline achieved pooled R² = 0.693 ± 0.025, indicating that conventional models already approach the apparent ceiling under interpolation-oriented settings. However, all methods collapsed under LOTO (best R² = -0.012), revealing that much of the apparent progress in the literature reflects chemical-neighbor memorization rather than robust target-level generalization. We further show that target-wise error is significantly associated with continuous protein semantic proximity in ProtBERT space (Spearman ρ = -0.461, p = 0.047), whereas coarse family-level descriptors are uninformative. A four-quadrant failure taxonomy reveals that protein shift is more damaging than chemical novelty (MAE 1.09-1.12 vs. 0.85-0.97), and conformal prediction becomes severely overconfident under target extrapolation, with empirical 90% coverage dropping to 63.9-66.7%. These results reposition PROTAC prediction as a problem of transferability and reliability rather than leaderboard optimization and provide practical guidelines for future benchmark design.

Ren-Guang Zhu, Guang-Hao Guo, Lu-Lu Li et al. · 0 citations
#machine learning Preprint Sep 2026

ProMeta: Few-shot PROTAC-targeted degradation prediction across E3 ligases

Proteolysis-targeting chimeras (PROTACs) have emerged as a transformative therapeutic strategy that selectively degrades historically''undruggable''targets via the ubiquitin-proteasome system. Despite growing efforts to develop computational predictors of PROTAC degradation activity, existing supervised approaches remain severely challenged by data scarcity and imbalance across E3 ligases, limiting their ability to generalize beyond well-studied ligase contexts. In practice, labeled data are heavily concentrated on a few ligases (e.g., CRBN and VHL), while the majority of E3 ligases remain underexplored yet are critical for expanding the design space of targeted degraders. Developing methods that enable robust cross-ligase generalization with minimal labeled data is therefore essential for improving the practical utility of computational PROTAC discovery. We reformulate PROTAC degradation activity prediction across E3 ligases as a few-shot meta-learning problem and present ProMeta, a prototype-based graph neural network trained through episodic meta-learning on source-E3 tasks and evaluated on held-out target-E3 tasks through support-conditioned inference. ProMeta performs inference without updating the encoder by dynamically estimating class prototypes from minimal target-ligase support samples. On the CRBN-to-VHL benchmark, ProMeta achieves AUROC values of 0.796 under K=2, Q=3 and 0.883 under K=2, Q=5, improving by 19.9% and 6.8%, respectively, over the corresponding supervised GNN baseline. Reverse VHL-to-CRBN transfer under the same protocol yielded AUROC values of 0.702 (K=2, Q=3) and 0.821 (K=2, Q=5), confirming bidirectional applicability while revealing direction and data-regime dependence. Together, these results support ProMeta as a practical framework for cross-ligase few-shot prediction under the evaluated support/query protocols.

Yuansheng Liu, Yu-Fei Ye, Tao Tang et al. · 0 citations
Review Aug 2026

The sweet spot in protein design-Where deep learning meets first principles.

It is argued that the highest-confidence candidates emerge where deep learning and first-principles models agree, a "sweet spot" that balances generative flexibility with thermodynamic realism.

Gabriel Cia, Gabriele Orlando, D. Cianferoni et al. · 0 citations
Review Open access Sep 2026

Deep Learning for Protein Modeling: From Single Structure to Bound Complex and Thermodynamic Ensemble, With a Focus on Architectural Design

We survey modern deep‐learning approaches to protein conformational modeling through the lens of architectural design. We organize the literature into three increasingly expressive paradigms: (I) single‐structure prediction, (II) prediction of molecular binding complexes, and (III) conformational ensemble generation. For each paradigm, we outline a representative set of models to sketch a practical taxonomy, and we summarize their key achievements, limitations, and common evaluation practices. Across the paradigms, we highlight recurring design choices that shape performance and generalization, including enforced SE(3) equivariance versus learned symmetry; MSA‐driven coevolution versus protein language model priors; deterministic prediction versus generative sampling; explicit energetic supervision versus implicit learning; and integrative modeling across heterogeneous data modalities. While single‐structure prediction is now relatively well established, comparable maturity has not yet been reached for binding‐complex prediction and, especially, for generating faithful thermodynamic ensembles with reliable population weights, which remains an open challenge. We discuss open challenges in building physically grounded and transferable models, including data availability and fidelity, the choice of inductive biases to pursue generalization, and the need for rigorous model evaluation. Ultimately, we indicate generative kinetics as an aspirational frontier.

Daniele Angioletti, Matteo Carli, Marco S. Nobile et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jul 14, 2026

Helping AI models to meet the real world

Through research and entrepreneurship, Professor Devavrat Shah is helping to design methods that can handle constant decision-making using limited computational resources.

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.