This work presents MG2Act, a structure-independent framework that translates two-step logic into sequential cross-attention, using CRBN-mediated degradation as the most data-rich representative system.
Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its biological context. Although public databases contain thousands of structured molecule-target-E3 records, degradation measurements are available for only a small fraction of them. Existing supervised approaches therefore leave most recorded chemical-biological relationships unused. We introduce DegradeQuery, a context-aware prediction framework that converts these label-missing records into a pretraining signal. Its counterfactual tuple pretraining objective contrasts recorded tuples with alternatives formed by replacing the target, the E3 ligase, or both, enabling the model to learn contextual associations without assigning activity pseudo-labels. The resulting representation is then fine-tuned to predict degradation from the complete molecule-target-E3 context. On the official PROTAC-8K benchmark, DegradeQuery achieves an area under the receiver operating characteristic curve of 0.9065 and an accuracy of 0.8500, outperforming the compared methods. Controlled analyses further show that the improvement is primarily attributable to tuple-level pretraining, can be recovered using only label-missing records, and remains complementary to protein language model representations. These findings demonstrate that incompletely labeled PROTAC databases contain useful relational supervision and provide a practical route for learning context-aware degradation predictors from scarce experimental labels.
Dong Xu, Zhangfan Yang, Jiantao Wu et al.· 0 citations
By combining protein language model embeddings with topology-adaptive geometric reasoning, DiConSite offers a reusable framework for residue-level protein interaction analysis and achieves consistently strong and often best-performing results, while improving robustness to structural uncertainty and cross-modal variation.
Shou-Zhi Chen, Zhenchao Tang, Linlin You et al.· IEEE Transactions on Pattern...· 1 citation
EvoBind-multimer is presented, a deep learning framework for the de novo design of molecular glues directly from protein sequences that generates small macrocyclic peptides that bridge user-defined protein pairs without requiring prior interface knowledge or existing ligands.
Andrä Brunner, Krzysztof Wierbiłowicz, Diandra Daumiller et al.· bioRxiv· 0 citations
Proteolysis-targeting chimeras (PROTACs) are heterobifunctional small molecules that induce targeted protein degradation by recruiting an E3 ligase to a protein of interest. Since 2019, publication volume has accelerated, and computational methods have expanded from isolated demonstrations into practical tools for modeling PROTAC-induced ternary complexes, designing linkers, and forecasting degradation-related outcomes. Here, we present a Perspective on computational PROTAC methodologies published from 2019 to the present, organizing the field into two complementary streams: (i) constraint-driven, physics-based workflows that assemble and refine ternary complex models by enforcing geometric feasibility and evaluating pose stability using docking and molecular simulation; and (ii) data-driven workflows, including deep learning predictors and generative models that predict ternary complex structure, degradation end points, or linker chemistry from structural and assay data. We highlight representative approaches spanning restrained/tethered docking, MD-based refinement and dynamic stability scoring, coarse-grained free-energy modeling, SE(3)/E(3)-equivariant structure prediction, supervised degradation efficacy prediction, and generative linker design. We close by emphasizing persistent gaps, fragmented benchmarking, score robustness across targets and E3 ligases, and nonstandard molecular representations that currently limit generalization and reproducible, pipeline-ready deployment.
Joseph M. Schulz, R. Reynolds, Stephan C. Schürer· Journal of Chemical Informat...· 0 citations
Computational prediction of PROTAC degradation activity (DC50) has attracted growing interest, yet the reliability of reported model performance remains poorly understood because sufficiently stringent evaluation protocols are rarely applied. Here, we present a hierarchical benchmark designed to expose evaluation pitfalls and quantify the transferability and reliability limits of current PROTAC predictors. Using a curated dataset of 2405 DC50 measurements spanning 22 target proteins and two E3 ligases (CRBN and VHL), we benchmarked classical machine learning (Random Forest, ExtraTrees, Ridge, PLS), gradient-boosted trees (XGBoost), nearest-neighbor retrieval baselines, protein negative controls, and a representative multi-modal deep learning ensemble (HybridMoECrossAttn) across Random, Scaffold, Leave-One-Target-Out (LOTO), and Leave-One-Family-Out (LOFO) splits. Under Random evaluation, a simple Random Forest + ECFP4 baseline achieved pooled R² = 0.693 ± 0.025, indicating that conventional models already approach the apparent ceiling under interpolation-oriented settings. However, all methods collapsed under LOTO (best R² = -0.012), revealing that much of the apparent progress in the literature reflects chemical-neighbor memorization rather than robust target-level generalization. We further show that target-wise error is significantly associated with continuous protein semantic proximity in ProtBERT space (Spearman ρ = -0.461, p = 0.047), whereas coarse family-level descriptors are uninformative. A four-quadrant failure taxonomy reveals that protein shift is more damaging than chemical novelty (MAE 1.09-1.12 vs. 0.85-0.97), and conformal prediction becomes severely overconfident under target extrapolation, with empirical 90% coverage dropping to 63.9-66.7%. These results reposition PROTAC prediction as a problem of transferability and reliability rather than leaderboard optimization and provide practical guidelines for future benchmark design.
Ren-Guang Zhu, Guang-Hao Guo, Lu-Lu Li et al.· Computational biology and ch...· 0 citations
Molecular glue degraders (MGDs), such as pomalidomide, induce degradation of non-native substrates by the cullin-RING E3 ligase 4 (CRL4) through its substrate receptor cereblon (CRBN). Here, to explore CRBN programmability, we tested whether reported CRBN-MGD substrates are part of a network of latent CRBN interactors, proteins capable of MGD-induced CRBN binding without detectable degradation. Leveraging a highly parallel protein complementation assay (GluePCA) to measure MGD-induced interaction between CRBN and zinc fingers, we identified ~210 zinc fingers bound to CRBN-pomalidomide, where top binders are already reported as degraded by dedicated MGDs. To map latent CRBN-MGD interactions proteome-wide and define the accessible CRBN interaction space, we combined artificial intelligence-derived protein surface queries (MaSIF-mimicry) with GluePCA. This pipeline identified 6 known and 43 novel CRBN-pomalidomide binders, including orthogonally validated hits. We find that these binders provide privileged starting points for MGD development. We expect this binding-focused workflow to be applicable to other MGD-E3 ligase systems, potentially extending the scope of this emerging drug class.
Pius Galli, Shuhao Xiao, Y. Meng et al.· Nature Biotechnology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.