ScrambleBench provides a holistic medicinal chemistry-oriented framework that identifies methodological strengths, limitations, and opportunities for future model development and highlights the importance of evaluating chemical diversity explicitly and using the recently proposed metrics such as Hamiltonian Diversity (HamDiv) which assess both quantity and dissimilarity of a molecular set.
Abstract
Generative artificial intelligence (AI) has rapidly advanced over the past decade in the field of drug discovery, particularly for the de novo design of small molecules based on target protein structures. While generative AI has the potential to complement traditional structure-based drug design (SBDD) to discover novel hits, their performance is often assessed using non-standardized evaluation criteria. As the number of generative AI models continue to grow, it becomes increasingly important to determine whether these tools are sufficiently robust and reliable for integration into medicinal chemistry workflow. Here, we propose ScrambleBench, a benchmarking workflow designed to evaluate structure-based generative AI models that align with medicinal chemists' practical objectives to identify chemically diverse, drug-like hit candidates that adopt plausible binding conformations and exhibit favourable docking affinities. Using six representative models (Pocket2Mol, PocketFlow, Lingo3DMol, DiffSBDD, PMDM, and Chemistry42), we systematically assessed performance across diverse target classes, including two GPCRs, two kinases, and two hydrolases. While some generative models show superior performance for particular evaluation endpoints, none demonstrates overall dominance across all evaluated criteria. Notably, despite benchmarked proteins (e.g., CDK2, GSK3β) being present in the training datasets, the models still show limited generalization to target binding sites, which resulted in high redocking RMSD values and low virtual hit rates. Our results highlight the importance of evaluating chemical diversity explicitly and using the recently proposed metrics such as Hamiltonian Diversity (HamDiv) which assess both quantity and dissimilarity of a molecular set. Furthermore, as many generated ligands fail to meaningfully engage the target active site, we propose that future generative frameworks incorporate improved loss functions that place greater emphasis on drug-like physicochemical properties and correct pharmacophore recognition.Scientific contributionWhile numerous de novo generative models have been proposed for structure-based molecular design, objective comparison between methods remains challenging due to inconsistent benchmarking practices and heterogeneous evaluation criteria. ScrambleBench introduces a unified and reproducible benchmarking workflow that integrates diversity analysis, conformational validity assessment, docking reproducibility, pharmacophore matching, and virtual hit rate evaluation within a single framework. By systematically comparing representative generative models using common datasets and standardized assessment criteria, this work advances the field by enabling transparent evaluation model performance and practical applicability. Overall, ScrambleBench provides a holistic medicinal chemistry-oriented framework that identifies methodological strengths, limitations, and opportunities for future model development.
Virtual screening (VS) is an essential tool in drug discovery to prioritize potential drug candidates from vast chemical space. One key challenge limiting its performance is accounting for protein conformational flexibility. While ensemble docking methods have been developed to address this challenge by incorporating multiple protein conformations, these methods often rely on computationally intensive physics-based simulations to sample the relevant conformational space. Generative machine learning models offer a highly promising, scalable, and high-throughput alternative to overcome the limitations of these traditional approaches. We therefore investigate whether conformational ensembles generated by BioEmu, a recently developed generative model, can improve VS performance for kinase targets. Using the DUD-E benchmark data set and a validated AutoDock-GPU protocol, we generated and analyzed nearly 1300 structures across 26 kinases (approximately 50 structures each). BioEmu produces structurally diverse ensembles with substantial performance variation among individual structures. However, ensemble methods employing consensus or best-score selection fail to improve upon, and often degrade, VS performance compared to crystal structure baselines. To investigate the source of this limitation, we quantified the relationship between KinCoRe-based conformational state classification and screening performance. By calculating the coefficient of determination (R2) across the kinase subset, we found that the structural features governing VS performance differ substantially from those defining standard conformational states, with KinCoRe classifications leaving over 84% of performance variance unexplained. This critical gap demonstrates that structural diversity alone is insufficient to guarantee screening success. We show that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.
Jaeoh Shin, K. Joo, Jejoong Yoo· Journal of Chemical Informat...· 0 citations
Generative molecular models can support early drug discovery by proposing new candidate compounds de novo. In practice, useful candidates must balance target-relevant activity, synthetic accessibility, physicochemical properties, and other multiparameter design constraints. However, metrics commonly used to evaluate molecular generators only weakly reflect whether the generated compounds are medicinally plausible and suitable for downstream computation. This can produce false positives in model evaluation, incorrect assumptions, and inefficient use of computational resources. We introduce HEDGEHOG, a unified six-stage filtration benchmark that is inspired by industrial hit identification workflows: (i) preprocessing; (ii) physicochemical descriptor screening; (iii) structural alerts and graph-sanity checks; (iv) synthesis feasibility; (v) docking and binding affinity estimation; and (vi) three-dimensional pose and interaction checks. We evaluate 23 molecular generators across three model classes under a standardized protocol. Across 230,000 generated molecules, only 0.65% of initial molecules survive all stages. Our results expose a central limitation of current molecular generators: molecules that appear acceptable under isolated criteria rarely satisfy medicinal chemistry, synthesis, docking, and 3D pose filters simultaneously.
Daria A. Ryabchenko, Pavel Gurevich, S. Kadyrov et al.· 0 citations
Structure-based generative models (SBGMs) hold great promises for accelerating drug discovery by enabling target-aware molecular design. However, existing approaches face fundamental challenges: three-dimensional graph-based models can explicitly incorporate protein structural information but often generate chemically implausible molecules due to limited training data, while chemical language models (CLMs) produce chemically plausible molecules but struggle to effectively leverage three-dimensional structural information for structure-conditioned generation and hard to incorporate lead optimization functionality due to the nature of SMILES string. Here, we present StructureSAFE, a structure-aware chemical language model that resolves this trade-off by integrating protein structural and evolutionary encoders with the SAFE molecular representation via pretraining and finetuning training scheme, enabling both de novo hit identification and a comprehensive suite of lead optimization subtasks within a unified framework. Comprehensive benchmarking on the MolGenBench dataset demonstrates that StructureSAFE achieves state-of-the-art (SOTA) performance across multiple metrics, with particularly pronounced improvements in chemical plausibility relative to graph-based models lacking pretraining. Evaluation on a rigorously constructed held-out test set further confirms its ability to generate drug-like, synthetically accessible molecules with competitive predicted binding affinities for previously unseen targets on both hit identification and lead optimization setting. In silico case studies across four therapeutically relevant targets validate its capacity to generate chemically plausible molecules that recapitulate key binding interactions of known high-affinity ligands while proposing novel interactions for potential better affinity and exploring previously unknown regions of chemical space. Taking together, StructureSAFE represents a versatile and practical tool to provide high-quality candidate molecules for augmenting medicinal chemistry workflows in both hit identification and lead optimization campaigns.
Bo Yang, Ke Xu, Chijian Xiang et al.· bioRxiv· 0 citations
Generative chemistry is an emerging discipline that utilizes generative artificial intelligence (AI) models for the automated de novo design of small molecules. By learning patterns from existing chemical data, these models can generate novel structures with desired properties, thereby accelerating drug discovery. However, a significant gap remains between the potential of AI and its successful implementation in practical pharmaceutical applications. This review covers various infrastructural aspects of generative chemistry, including molecular representation, databases, and diverse model architectures such as generative adversarial networks, variational autoencoders, and diffusion models. Furthermore, key challenges associated with data quality, model selection, and synthesis feasibility are critically discussed. The review highlights that generative chemistry has evolved beyond simple structure generation to encompass the entire molecular design pipeline, including automated synthesis planning, retrosynthesis prediction, and multi‐objective optimization. Additionally, the selection of the most suitable model depends on specific objectives and the quality and diversity of the dataset, rather than a single superior architecture. Overall, a critical perspective is provided on how generative models are shaping the future of rational and reliable drug design.
Rania Ehab Koshty, Manar Ahmed Shehata, Ahmed M. Gab Allah et al.· ChemistrySelect· 0 citations
Abstract Summary Recent advances in computational methods for designing biological sequences have sparked the development of metrics to evaluate these methods performance in terms of the fidelity of the designed sequences to a target distribution and their attainment of desired properties. However, a software library implementing these metrics was lacking. In this work we introduce seqme, a modular and highly extendable open-source Python library, containing model-agnostic metrics for evaluating computational methods for biological sequence design. seqme considers three groups of metrics: sequence-based, embedding-based, and property-based, and is applicable to a wide range of biological sequences: small molecules, DNA, ncRNA, mRNA, peptides and proteins. The library offers a number of embedding and property models for biological sequences, as well as diagnostics and visualization functions to inspect the results. seqme can be used to evaluate both one-shot generation and iterative optimization. We show the utility of seqme by performing an antimicrobial peptide benchmark and acquiring mRNA data. Availability and implementation seqme is released at https://github.com/szczurek-lab/seqme under the BSD 3-Clause license.
Rasmus Møller-Larsen, Adam Izdebski, Jan Olszewski et al.· Bioinformatics Advances· 2 citations
Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared binding patterns and intrinsic similarities among protein-ligand complexes. This fragmented perspective constrains their ability to capture the fundamental principles governing molecular recognition and binding specificity. Moreover, the limited availability of high-quality experimental data further hampers model generalization and real-world applicability. To address these challenges, we present READ, a retrieval-alignment molecular generation framework that conditions the generative process on small molecules targeting homologous proteins. Retrieved ligands are aligned with a diffusion model across multiple representational spaces and integrated as conditional guidance throughout successive stages of generation. Under a standardized docking-based evaluation protocol, READ achieves consistently strong performance against state-of-the-art SBDD methods. More importantly, it introduces a retrieval-alignment paradigm for structure-based molecular generation, offering a practical framework for early-stage computational hit generation while leaving prospective experimental validation as future work.
Dong Xu, Zhangfan Yang, Junchuang Cai et al.· IEEE transactions on computa...· 1 citation