Virtual screening (VS) on small molecules aims to identify promising drug candidates against protein targets from expansive chemical libraries by balancing the core requirements of accurate scoring and efficient search against the inherent trade‐off between accuracy and speed. This survey provides a comprehensive review of how Artificial Intelligence and Machine Learning (AI/ML) are redefining this landscape across three critical dimensions. First, we examine the evolution of AI‐driven scoring functions, which utilize AI/ML models to capture complex structure–activity relationships from massive biochemical datasets, significantly enhancing structure‐ and ligand‐based evaluations beyond traditional heuristics. Second, we summarize the emergence of efficient search algorithms that iteratively prioritize informative compounds to reduce search efforts by orders of magnitude. Third, we review the paradigm shift toward generative molecular design, making VS transition from screening fixed libraries to the
de novo
generation of molecules optimized for specific structural contexts and multi‐objective properties. This review outlines the transition toward end‐to‐end, adaptive discovery systems that ensure computational hits are biologically potent, structurally optimized, and synthetically accessible.
Yifei Wang, Nupur Bansal, Shiyun Wa et al.· WIREs Computational Molecula...· 0 citations
Goal-directed optimization is essential for steering molecular generators to propose candidates with desired properties. However, it is often implemented with policy-gradient reinforcement learning, which requires a generation-trajectory log-probability whose form depends on the model architecture and generation procedure. This makes an optimizer difficult to reuse across architectures and conditional generative designs. Supervised fine-tuning needs none of that machinery, but its update is driven by a fixed dataset, so the reward never enters the update. We introduce Elite-Weighted Supervised Fine-tuning (EW-SFT), which uses reward to guide elite selection of high-scoring molecules, and updates the model by its own pretraining loss on that set. Ablations show that reward information is passed primarily through elite selection, rather than through continuous weighting within the selected set. Because the update consumes only scored molecules and the model's native loss, the same rule applies across autoregressive, masked-diffusion, and discrete-flow generators, and across de novo, motif-extension, and linker-design tasks. Under a fixed budget of 3D shape alignment oracle calls on two kinase reference compounds, EW-SFT consistently outperforms the corresponding native optimizers. It further improves goal-directed optimization under a 2D similarity oracle on four held-out references and achieves comparable performance on a sample-efficiency benchmark without a trajectory-level RL formulation. These results demonstrate that EW-SFT is a unified and effective optimizer across molecular generators, design constraints, references, and oracles.
Shiyun Wa, Yifei Wang, A. G. Green et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.