Aug 2026· Journal of Chemical Information and Modeling· 0 citations· 27 references
TL;DR
DASH is presented, a pocket-aware and objective-aware framework for converting protein-conditioned diffusion outputs into property-tagged molecular libraries for computational prioritization and can adapt inference effort across pockets, rerank generated libraries under different design objectives, and produce standardized molecular libraries for downstream analysis.
Abstract
Structure-based molecular diffusion models have shown considerable potential for de novo drug design. However, their practical use in million-scale candidate-library construction remains limited by fixed target-agnostic inference settings, insufficient support for objective-aware molecular prioritization, and limited integration of scalable property evaluation with standardized library production. Here, we present DASH, a pocket-aware and objective-aware framework for converting protein-conditioned diffusion outputs into property-tagged molecular libraries for computational prioritization. DASH combines pocket-complexity-aware sampling, configurable objective-aware molecular scoring, and a scalable production layer. The sampling strategy adjusts inference effort according to geometric and physicochemical features of the target binding pocket. The Objective-Aware Quality Module (OQM) filters and reranks generated molecules using configurable descriptors, desirability functions, and scoring profiles. The production layer supports million-scale execution through asynchronous GPU–CPU processing, streaming output, molecular scoring, and SDF annotation. We evaluated DASH through pocket-complexity analysis, OQM profile analysis, execution benchmarks, and multi-GPU/multinode scaling experiments. The results show that DASH can adapt inference effort across pockets, rerank generated libraries under different design objectives, and produce standardized molecular libraries for downstream analysis. An EGFR-oriented case study further demonstrates how DASH-generated libraries can support downstream computational prioritization and representative candidate selection. Together, these results demonstrate that DASH extends protein-conditioned diffusion models from raw molecular generation toward practical-scale, objective-aware candidate-library construction and computational hit prioritization.
MolecularCanvas is an interactive system that enables users to iteratively construct an optimization context by integrating high-level goals, structure-level annotations, property constraints, and reference-based preferences that guides the generation of candidate molecules across diverse molecular structures.
Haoyu Dong, Rui Sheng, Shu-Hao Zhang et al.· 0 citations
Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a balance between pocket compatibility, molecular properties, and physical geometry. We propose \textbf{PocketVE}, a protein-pocket-conditioned variance-exploding (VE) diffusion framework that couples stable coordinate denoising with inference-time property guidance. Specifically, PocketVE combines an EDM-style training and sampling setup for 3D denoising, classifier-free guidance for multi-property steering without external property classifiers, and adaptive protein perturbation as a training-time pocket regularizer. Evaluated on CrossDocked2020 under the GenBench3D protocol, PocketVE improves Valid$_{3\text{D}}$ from 58.6 to 80.6 and reduces strain energy from 457.4 to 127.9 relative to its TAGMol architectural baseline, while retaining competitive docking and molecular-property scores under moderate guidance. A guidance-scale study shows that moderate guidance gives a favorable balance between target-related objectives and geometric quality, whereas stronger guidance can degrade geometry and distributional fidelity. Pocket-permutation and PoseCheck diagnostics further support pocket-specific spatial compatibility with reduced steric conflicts. Overall, the results suggest that geometric stability and inference-time property guidance should be considered as coupled design objectives.
Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared binding patterns and intrinsic similarities among protein-ligand complexes. This fragmented perspective constrains their ability to capture the fundamental principles governing molecular recognition and binding specificity. Moreover, the limited availability of high-quality experimental data further hampers model generalization and real-world applicability. To address these challenges, we present READ, a retrieval-alignment molecular generation framework that conditions the generative process on small molecules targeting homologous proteins. Retrieved ligands are aligned with a diffusion model across multiple representational spaces and integrated as conditional guidance throughout successive stages of generation. Under a standardized docking-based evaluation protocol, READ achieves consistently strong performance against state-of-the-art SBDD methods. More importantly, it introduces a retrieval-alignment paradigm for structure-based molecular generation, offering a practical framework for early-stage computational hit generation while leaving prospective experimental validation as future work.
Dong Xu, Zhangfan Yang, Junchuang Cai et al.· IEEE transactions on computa...· 1 citation
A clear pattern is revealed in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.
Thomas MacDougall, Maksim Kuznetsov, Roman Schutski et al.· arXiv.org· 1 citation
Accurate prediction of drug-target binding affinity (DTA) is a key task in virtual screening. However, current computational methods face a key challenge: sequence-based approaches often fail to capture critical spatial information, while structure-based models rely on computationally expensive 3D coordinates, which restrict their scalability. To address this issue, we propose StructuraDTA, a novel multimodal framework that adopts an implicit structure modeling strategy. Instead of using static protein folding data, our method encodes drug molecular graphs via Graph Isomorphism Networks (GINs) to capture fine-grained topological features. Meanwhile, we optimize protein representations by integrating probabilistic structural priors into a pretrained language model, which effectively simulates thermodynamic conformational flexibility without relying on explicit 3D structural data. A bidirectional cross-attention mechanism is then used to dynamically align these heterogeneous feature modalities. Comprehensive evaluations on the Davis and KIBA benchmark datasets show that StructuraDTA stably outperforms state of-the-art comparison methods. Importantly, the model exhibits strong robustness in cold-start scenarios, and can accurately predict binding affinities for previously unseen drugs and targets. By retaining the predictive performance of structure based models while maintaining the high inference efficiency of sequence-based methods, we provide an accurate and scalable solution to accelerate genome-scale drug discovery research.
Junlin Xu, Ye Yuan, Menglong Hu et al.· IEEE journal of biomedical a...· 0 citations
The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.
Huiming Bao, Shouliang Dong· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.