P Pep2Mol is introduced, a diffusion-based generative model for 3D molecule design that targets orthosteric PPI sites by explicitly incorporating binding peptides or proteins as structural guidance, moving beyond conventional pocket-conditioned generation.
Abstract
Protein-protein interactions (PPIs) are central to biological processes. Designing small molecules that modulate dysregulated PPIs holds strong promise for targeting undruggable proteins. However, existing structure-based drug design approaches focus on well-defined small-molecule binding pockets and struggle to generalize to large, shallow, and chemically complex PPI interfaces. Here, we introduce Pep2Mol, a diffusion-based generative model for 3D molecule design that targets orthosteric PPI sites by explicitly incorporating binding peptides or proteins as structural guidance, moving beyond conventional pocket-conditioned generation. To enable model development and benchmarking, we curate a large-scale, high-quality dataset of 10,956 experimentally resolved protein complex structure pairs, each pairing an orthosteric competitive ligand with a protein binder at overlapping receptor interfaces. Pep2Mol integrates two SE(3)-equivariant graph neural networks that encode protein–ligand and protein–peptide interactions respectively, and fuses these representations via attention-based conditioning to jointly guide the diffusion trajectory. Extensive evaluations demonstrate that Pep2Mol generates chemically valid ligands with state-of-the-art binding affinities, providing a strong foundation for small-molecule inhibitor design against challenging PPI interfaces.
This work used its strategy, termed neural iterative selection–expansion (NISE), to design proteins that, using different folds, specifically bind to two chemically distinct small-molecule drugs, exatecan and apixaban, with success rates of 100%, respectively.
Benjamin Fry, Kaia Slaw, Nicholas F. Polizzi· Nature· 7 citations· ⚡1
How advances in artificial intelligence and computational modeling may reshape the rational design of next-generation peptide therapeutics is explored and an integrated experimental–computational framework is proposed to facilitate the development of clinically actionable candidates is proposed.
Ha Thi Ngoc Nguyen, B. Le, Nhung Thi Hong Van et al.· Pharmaceuticals· 0 citations
Abstract Motivation Identifying protein binding sites in protein–protein complexes is a central challenge in structural biology. Binding sites, consisting of groups of residues, govern how proteins recognize, and interact with protein partners. Thus, identifying them is essential for understanding biological function and guiding the design of effective biomolecules and even drug molecules. Despite major progress in computational approaches, their performance remains limited because most models underrepresent the combined influence of surface properties and residue-level information, leaving room for improvement. Recent advances in state-space models and vision-based deep learning offer an opportunity to address these limitations by efficiently modeling long-range spatial dependencies on protein surfaces. Here, we introduce BiMba (protein Binding site prediction using Vision Mamba), a state-space–driven deep learning framework that leverages the efficient long-range modeling capability of the Vision Mamba architecture to learn from three-dimensional (3D) protein surfaces represented as two-dimensional (2D) geometric or physicochemical grids. Results BiMba integrates complementary sources of information, capturing geometric and physicochemical determinants of molecular recognition as surface patches, encoded as 2D images, along with residue-level descriptors, yielding a unified representation that couples spatial topology with biochemical context. BiMba demonstrates competitive performance across diverse and specialized benchmark datasets, often outperforming existing state-of-the-art methods. In addition, BiMba incorporates perturbation-based and gradient-based interpretability analyses by extracting hidden attentions from Mamba layers, enabling visualization of feature relevance and biologically meaningful residue clusters. Overall, our findings establish state-space models as efficient, interpretable, and scalable architectures for molecular surface learning, advancing the application of deep learning in structural bioinformatics. Availability and implementation The BiMba source code, training, test, and benchmark datasets are available at https://github.com/Azam-Shi/BiMba.
Abstract Summary We introduce DruGUI 2.0, a drug discovery tool for assessing the druggability of proteins, integrated into the ProDy application programming interface (API). DruGUI 2.0 is developed to facilitate the search for druggable sites while allowing for proteins’ conformational flexibility. Simulations in explicit solvent, with an option to include membrane, are carried out in the presence of probe molecules selected from an expanded library of small molecules containing drug-like fragments. Druggable sites beyond orthosteric sites are identifiable, as well as the probes that show high affinity to bind to those sites. Characterization of the composition and position of the probes helps build pharmacophore models and estimate relative binding affinities. As a Python module with enhanced visualization features, DruGUI 2.0 complements, and benefits from, the vast collection of protein sequence, structure, and dynamics analyses modules accessible in ProDy. Case studies in the Supplemental Material showcase the utility of DruGUI 2.0 applied to both soluble targets and membrane proteins. Availability ProDy is open-sourced and freely available under MIT License from https://github.com/prody/ProDy. The code version of DruGUI 2.0 used for simulations is available on Zenodo : 10.5281/zenodo.20511357.
Carlos Ventura, J. Y. Lee, Anthony T. Bogetti et al.· Bioinform.· 0 citations
HyBind-NN is developed, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein–protein and protein–peptide affinity, and it is demonstrated that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets.
E. A. Bogdanova, A. Chernukhin, Alexey K. Shaytan· International Journal of Mol...· 0 citations
Drug discovery and development is time-consuming and resource-intensive, motivating computational approaches such as diffusion models for de novo drug design. Many such models follow the structure-based drug design (SBDD) paradigm, generating molecules to fit a target binding pocket. However, existing diffusion-based SBDD methods typically couple pocket and ligand representation learning, model interactions only at the atom level, and prioritize binding affinity over other developability properties. Here, we introduce conDitar-dev, a conditional diffusion-based SBDD framework for generating ligands with strong binding affinities and favorable ADMET properties. It consists of three modules: msPRL, a pretrained multi-scale pocket representation learning module; conDitar, a pocket-conditioned diffusion model guided by msPRL representations; and paOPT, a generation-time method for optimizing ligand developability. On a newly curated benchmark of human disease targets, conDitar outperforms state-of-the-art SBDD baselines, achieving an average binding score of -8.85 kcal/mol. Across five ADMET properties, conDitar-dev improves performance by up to 73% over conDitar. To further validate the abilities of conDitar-dev to generate developable molecules, we have applied it to two validated druggable targets: programmed death-ligand 1 (PD-L1) and colony-stimulating factor 1 receptor (CSF1R) proteins. Top-ranked generatively designed molecules and their analogs have been experimentally synthesized and biologically tested. Two molecules generated directly by conDitar-dev for PD-L1 exhibited SPR-derived $K_D$ values of 3.49 and 3.75 $\mu$M, respectively. Hit expansion based on conDitar-dev-designed molecules identified selective CSF1R inhibitors with IC$_{50}$ values as low as 200 nM, while also uncovering opportunities for drug repositioning.
Ruoxi Gao, Jiangweizhi Peng, Ziqi Chen et al.· 0 citations