Drug synergy prediction holds great promise in accelerating combination therapy development and improving treatment efficacy in cancer and other complex diseases. However, progress in this area is hindered by considerable heterogeneity across experimental datasets, including variability in the number of drug combinations, inconsistencies in synergy scoring methodologies, and differences in data quality. Here, we present the first comprehensive benchmarking framework specifically designed to accommodate inter-dataset heterogeneity. This framework integrates 13 independent datasets encompassing 454,794 retained drug combination-cell line entries, 4247 drugs, and 187 cell lines, all of which are cancer cell lines. Our comparative evaluation of seven computational models for drug synergy prediction reveals that model performance strongly depends on both the scale and quality of the datasets. The graph model JointSyn performed favorably on datasets with larger numbers (> 10,000) of retained drug combinations, while the traditional random forest model performed competitively on smaller-scale datasets. We also observe that the ZIP scoring metric yields the highest accuracy in large-scale data, whereas HSA is more effective in sparse-data scenarios. However, different synergy metrics show significant variability in performance across datasets, suggesting that different synergy metrics capture distinct aspects of drug interactions, and the choice of metric can substantially affect model evaluation and cross‑dataset consistency. Furthermore, we find that well-designed small datasets can match or even surpass the performance of larger benchmarks, suggesting that different metrics are applicable to different datasets/testing scenarios. Our benchmark provides a robust foundation for fair model evaluation and paves the way for the development of more generalizable and preclinically relevant drug synergy prediction methods.
Yingjuan Cheng, Qing Ye, Linlong Jiang et al.· Journal of Cheminformatics· 0 citations
Sampling rare conformation transitions between metastable states is a central challenge in atomistic simulations. While the committor function serve as an ideal reaction coordinate for driving enhanced sampling, their high-dimensional inputs and complex functional forms limit the efficacy of standard feedforward neural networks in modeling them. Inspired by recent breakthroughs in biomolecular structure prediction, we propose a novel committor learning framework grounded in the AlphaFold 3 paradigm. By integrating a lightweight, differentiable atom-level embedding with a simplified Pairformer architecture, our method inherently captures intricate dynamical features of diverse biosystems without requiring specialized prior knowledge. We demonstrate the superior expressiveness and accuracy of the proposed framework across multiple atomistic processes. For the folding of the chignolin mini-protein, our model reveals the finer-grained structure of its transition state ensemble (TSE) and a detailed bifurcated reaction mechanism. Furthermore, for calixarene host-guest systems, we develop a unified committor model that elucidates how ligand substituents regulate the ratio between distinct binding pathways, offering new perspectives for structure-based drug design.
Jintu Zhang, Zichang Jin, Huifeng Zhao et al.· 0 citations
Current structure-based drug design generative models often struggle to faithfully recapitulate genuine ligand-protein binding interactions. Instead, under the coupling of implicit learning architectures and biased training data, they tend to learn spurious statistical correlations. To address this, we propose EIP-Diff (Explicit Interaction-Prompted Diffusion), an architecture featuring a novel explicit interaction-prompt embedding mechanism that is better suited for real-world target-specific drug design. This architecture replaces biased implicit learning with explicit, residue-level biological guidance, thereby promoting more fine-grained geometric fidelity and more precise interaction-aware conditioning. To fully realize the capabilities of EIP-Diff and provide a reliable basis for performance evaluation, we further constructed CrystalData set, which provides higher-fidelity and less-biased structural supervision than existing data sets. This explicit architecture markedly improves distribution consistency: even when trained on the crossdocked data set, EIP-Diff achieves the highest alignment with authentic pharmacological distributions among evaluated models. Training on CrystalData set further enhances this alignment and improves 3D geometric accuracy, while retaining strong controllability, high chemical space coverage, and near-perfect uniqueness. In addition, target-based validation on KAT6A and YTHDC1 confirmed that EIP-Diff accurately recapitulates native-like binding modes. Furthermore, in a real-world drug design task against IDO1, we successfully designed a novel lead compound with nanomolar potency (IC50 = 0.31 nM). These results demonstrate that the EIP-Diff architecture can explicitly leverage experimentally derived structural data and biologically meaningful interaction information for target-specific molecular generation, thereby enabling its effective application to real-world structure-based drug design.
Huabin Du, Mingyang Wang, M. Luo et al.· Journal of the American Chem...· 0 citations
This paper introduces Caduceus, a family of MoE-enhanced foundation models built with a hierarchical pre-training paradigm to jointly integrate biological and natural language, and incorporates a multi-task instruction tuning phase, enabling robust protein parsing and natural language question answering.
Mingze Yin, Yiheng Zhu, Jialu Wu et al.· Proceedings of the 32nd ACM...· 0 citations