Skip to content
Preprint

Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization

Jul 2026 · 0 citations · 88 references
Computer Science

Abstract

Drug discovery and development is time-consuming and resource-intensive, motivating computational approaches such as diffusion models for de novo drug design. Many such models follow the structure-based drug design (SBDD) paradigm, generating molecules to fit a target binding pocket. However, existing diffusion-based SBDD methods typically couple pocket and ligand representation learning, model interactions only at the atom level, and prioritize binding affinity over other developability properties. Here, we introduce conDitar-dev, a conditional diffusion-based SBDD framework for generating ligands with strong binding affinities and favorable ADMET properties. It consists of three modules: msPRL, a pretrained multi-scale pocket representation learning module; conDitar, a pocket-conditioned diffusion model guided by msPRL representations; and paOPT, a generation-time method for optimizing ligand developability. On a newly curated benchmark of human disease targets, conDitar outperforms state-of-the-art SBDD baselines, achieving an average binding score of -8.85 kcal/mol. Across five ADMET properties, conDitar-dev improves performance by up to 73% over conDitar. To further validate the abilities of conDitar-dev to generate developable molecules, we have applied it to two validated druggable targets: programmed death-ligand 1 (PD-L1) and colony-stimulating factor 1 receptor (CSF1R) proteins. Top-ranked generatively designed molecules and their analogs have been experimentally synthesized and biologically tested. Two molecules generated directly by conDitar-dev for PD-L1 exhibited SPR-derived $K_D$ values of 3.49 and 3.75 $\mu$M, respectively. Hit expansion based on conDitar-dev-designed molecules identified selective CSF1R inhibitors with IC$_{50}$ values as low as 200 nM, while also uncovering opportunities for drug repositioning.

View source

Similar papers

Preprint Jul 2026

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

A clear pattern is revealed in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.

Thomas MacDougall, Maksim Kuznetsov, Roman Schutski et al. · 0 citations
Open access Aug 2026

PockLigGPT: Pocket-Sequence-Conditioned Molecular Generation with GPTs and RL

De novo drug design aims to generate molecules targeting specific protein pockets while retaining chemical plausibility and drug-like properties. Recent 3D structure-based generative methods explicitly model pocket-ligand geometry, but this does not always translate into chemically realistic or practically usable candidate molecules. Molecular language models provide a complementary sequence-based alternative. However, it remains unclear whether sequence-based pocket information can effectively guide ligand generation, whether multistage training improves pocket-specific generation, and whether docking-guided reinforcement learning can be integrated into a practical generation pipeline. We introduce PockLigGPT, a GPT-based framework for pocket-sequence-conditioned molecular generation. Rather than producing fixed 3D coordinates, PockLigGPT formulates ligand design as a sequence-generation problem conditioned on the amino acid composition of the protein pocket. The model is trained in four stages: large-scale chemical pretraining based on ZINC20; bioactivity-oriented adaptation based on ChEMBL; pocket-sequence-conditioned fine-tuning using binding-pocket amino acid sequences paired with ligands; and, finally, pocket-specific docking-guided reinforcement learning using AutoDock Vina-based rewards. PockLigGPT achieves competitive docking-oriented performance under a standardized evaluation protocol while maintaining chemical plausibility, favorable physicochemical profiles, and Lipinski-based drug-likeness. Docking studies on Alzheimer’s disease-associated targets and token-level analyses further support the utility of PockLigGPT for de novo drug design.

Pablo Varas Pardo, Guillermo Marcos-Ayuso, Eugenia Ulzurrun et al. · 0 citations
Review Open access Aug 2026

Geometric Deep Learning‐Based Drug Design Models for Small‐Molecule Drug Discovery

Deep neural network (DNN)‐based in silico models show great promise in predicting the properties and bioactivities of novel compounds, including small molecules. Among traditional approaches, structure‐based drug design (SBDD) remains a fundamental approach for drug discovery using molecular docking, scoring functions, and molecular dynamics simulations. However, these approaches are often constrained by limited flexibility, resolution, and generalizability. Geometric deep learning (GDL) offers a transformative alternative by enabling models to learn directly from non‐Euclidean molecular representations, such as graphs, point clouds, and meshes, capturing critical 3D spatial relationships inherent to protein–ligand interactions. This review highlights the theoretical underpinnings and practical applications of GDL in small‐molecule drug discovery, focusing on tasks including binding affinity prediction, virtual screening, de novo molecule generation, pose prediction, ADMET profiling, and protein flexibility modeling. We explore key GDL architectures, graph neural networks, SE(3)‐equivariant networks, 3D convolutional neural networks, point cloud models, and geometric transformers, and assess their performance across various drug discovery benchmarks. The integration of geometry‐aware AI models with experimental and computational workflows was also highlighted for its potential to streamline hit‐to‐lead optimization and advance rational drug design. Despite remarkable progress, the field faces challenges including limited high‐quality 3D structural datasets, protein flexibility representation, and the interpretability of deep models. Addressing these issues through hybrid modeling approaches, multi‐resolution learning, and self‐supervised training could further elevate GDL's impact. Ultimately, GDL stands at the frontier of AI‐enhanced pharmaceutical innovation, offering unprecedented precision, efficiency, and insight in the pursuit of next‐generation therapeutics.

A. Srivastav, Unnati Modi, Rahul Kumar et al. · 0 citations
Aug 2026

An Explicit Interaction-Prompted Diffusion Framework for High-Fidelity 3D Molecular Generation.

Current structure-based drug design generative models often struggle to faithfully recapitulate genuine ligand-protein binding interactions. Instead, under the coupling of implicit learning architectures and biased training data, they tend to learn spurious statistical correlations. To address this, we propose EIP-Diff (Explicit Interaction-Prompted Diffusion), an architecture featuring a novel explicit interaction-prompt embedding mechanism that is better suited for real-world target-specific drug design. This architecture replaces biased implicit learning with explicit, residue-level biological guidance, thereby promoting more fine-grained geometric fidelity and more precise interaction-aware conditioning. To fully realize the capabilities of EIP-Diff and provide a reliable basis for performance evaluation, we further constructed CrystalData set, which provides higher-fidelity and less-biased structural supervision than existing data sets. This explicit architecture markedly improves distribution consistency: even when trained on the crossdocked data set, EIP-Diff achieves the highest alignment with authentic pharmacological distributions among evaluated models. Training on CrystalData set further enhances this alignment and improves 3D geometric accuracy, while retaining strong controllability, high chemical space coverage, and near-perfect uniqueness. In addition, target-based validation on KAT6A and YTHDC1 confirmed that EIP-Diff accurately recapitulates native-like binding modes. Furthermore, in a real-world drug design task against IDO1, we successfully designed a novel lead compound with nanomolar potency (IC50 = 0.31 nM). These results demonstrate that the EIP-Diff architecture can explicitly leverage experimentally derived structural data and biologically meaningful interaction information for target-specific molecular generation, thereby enabling its effective application to real-world structure-based drug design.

Huabin Du, Mingyang Wang, M. Luo et al. · 0 citations