Skip to content

MetalloDock: Decoding Metalloprotein-Ligand Interactions via Physics-Aware Deep Learning for Metalloprotein Drug Discovery.

Jan 2026 · Journal of the American Chemical Society · Vol 148, pp. 3086-3101 · 6 citations · 46 references
Medicine

Abstract

Accurate prediction of metalloprotein-ligand interactions is critical for metalloprotein-targeted drug discovery. Conventional docking tools and existing deep learning (DL) models fail to reliably capture metal-ligand interactions, hampering the discovery of potent metalloprotein inhibitors. Here, we propose MetalloDock, the first DL-based docking framework specially designed for metalloprotein targets. By innovatively integrating an autoregressive spatial decoding engine with a physics-constrained geometric generation paradigm, MetalloDock can precisely reconstruct metal coordination geometries and accurately capture metal-ligand interactions, which enhance both the accuracy of metalloprotein-ligand docking and binding affinity prediction. Extensive evaluations on our custom-built benchmark data set demonstrate that MetalloDock outperforms existing methods, including AlphaFold3, in docking success rate and virtual screening performance for metalloprotein targets. In real-world applications, MetalloDock successfully identified multiple novel hit compounds in a virtual screening campaign targeting the prostate-specific membrane antigen. Additionally, it enabled rational drug design for acidic polymerase endonuclease, leading to the discovery of potent inhibitors. These results highlight the broad applicability of MetalloDock in accelerating metalloprotein-targeted drug discovery and provide a standardized framework for future evaluation of metalloprotein-specific docking algorithms.

View source

Similar papers

Open access Aug 2026

Targeting BCL-2 through Deep Learning-Based Drug Repurposing: A Multimodal Approach Combining Diffusion-Based Generative Modeling, Neural Relational Inference, and In Vitro Validation

Accurate identification of repurposable BCL-2 ligands requires not only plausible bound complex structures but also a dynamic description of how ligand binding reshapes residue-level communication. Here, we present a multimodal BCL-2 repurposing workflow built with diffusion-based generative modeling for ligand-specific complex generation and an extended neural relational inference (NRI) framework for trajectory-level interaction analysis. NeuralPlexer was applied to a library of 3094 FDA-approved drugs to generate BCL-2-ligand complex conformations at scale, yielding 1294 structurally acceptable complexes for downstream prioritization. To complement static scoring, filtered candidates were evaluated by molecular docking, anticancer QSAR classification, all-atom molecular dynamics (MD) simulations, and MM/GBSA binding free-energy calculations. We then extended NRI to protein–ligand trajectories to quantify residue-ligand and residue–residue dynamic couplings, enabling comparison of candidate-specific interaction signatures against the reference BCL-2 inhibitor Venetoclax. Among the prioritized compounds, Relugolix emerged as one of the most compelling hits, combining favorable binding energetics with an NRI-derived interaction pattern closely resembling that of Venetoclax. In vitro experiments supported BCL-2 inhibition by Relugolix in a TR-FRET assay and reduced viability of LN-18 glioma cells (IC50 = 23.55 μM). Together, these results establish a strategy that couples generative complex prediction with graph-based dynamic inference for structure-guided drug repurposing and identify Relugolix as a tractable scaffold for future BCL-2 inhibitor design.

Ehsan Sayyah, H. Tunc, A. Çelebi et al. · 0 citations
Open access Jul 2026

SurroDock: A Deep Learning Surrogate for Accelerated Pre-Docking Ligand Prioritization in Structure-Based Virtual Screening

Results indicate that 2D-based docking-score surrogate modeling can provide a reproducible and retrainable strategy for large-scale structure-based virtual screening by concentrating docking resources on a smaller, enriched subset of compounds.

Jongkeun Choi · 0 citations
Review Open access Aug 2026

Geometric Deep Learning‐Based Drug Design Models for Small‐Molecule Drug Discovery

Deep neural network (DNN)‐based in silico models show great promise in predicting the properties and bioactivities of novel compounds, including small molecules. Among traditional approaches, structure‐based drug design (SBDD) remains a fundamental approach for drug discovery using molecular docking, scoring functions, and molecular dynamics simulations. However, these approaches are often constrained by limited flexibility, resolution, and generalizability. Geometric deep learning (GDL) offers a transformative alternative by enabling models to learn directly from non‐Euclidean molecular representations, such as graphs, point clouds, and meshes, capturing critical 3D spatial relationships inherent to protein–ligand interactions. This review highlights the theoretical underpinnings and practical applications of GDL in small‐molecule drug discovery, focusing on tasks including binding affinity prediction, virtual screening, de novo molecule generation, pose prediction, ADMET profiling, and protein flexibility modeling. We explore key GDL architectures, graph neural networks, SE(3)‐equivariant networks, 3D convolutional neural networks, point cloud models, and geometric transformers, and assess their performance across various drug discovery benchmarks. The integration of geometry‐aware AI models with experimental and computational workflows was also highlighted for its potential to streamline hit‐to‐lead optimization and advance rational drug design. Despite remarkable progress, the field faces challenges including limited high‐quality 3D structural datasets, protein flexibility representation, and the interpretability of deep models. Addressing these issues through hybrid modeling approaches, multi‐resolution learning, and self‐supervised training could further elevate GDL's impact. Ultimately, GDL stands at the frontier of AI‐enhanced pharmaceutical innovation, offering unprecedented precision, efficiency, and insight in the pursuit of next‐generation therapeutics.

A. Srivastav, Unnati Modi, Rahul Kumar et al. · 1 citation
Open access Sep 2026

Cross-docking and redocking reveal distinct determinants of success in physics-based and AI-driven binding pose prediction in protein–ligand complexes

Protein–ligand pose prediction is central to structure-based drug discovery, yet the relative performance of physics-based and AI-driven methods under realistic cross-docking conditions remains insufficiently characterized. Here, we compare physics-based docking methods (AutoDock4, AutoDock Vina, and DOCK 6) with data-driven approaches, including the deep-learning model GNINA 1.3 and the diffusion-based frameworks AlphaFold 3, Boltz-2, and DiffDock. Performance was evaluated using standardised redocking and cross-docking protocols across three Alzheimer's disease targets representing distinct binding-site architectures: acetylcholinesterase (AChE; deep gorge), β-secretase 1 (BACE1; flexible flap-controlled site), and glycogen synthase kinase-3β (GSK-3β; open, solvent-exposed pocket). Physics-based methods were competitive during redocking but showed substantial performance reductions under cross-docking, whereas diffusion-based approaches generally maintained higher cross-docking accuracy. GNINA 1.3 rigid achieved an 87.7% minimum heavy-atom RMSD success rate during redocking, which decreased to 13.5% during cross-docking, whereas AlphaFold 3, Boltz-2, and DiffDock achieved cross-docking success rates of 93.1%, 89.6%, and 85.7%, respectively. AlphaFold 3 consistently outperformed Boltz-2 despite its smaller training set, suggesting that predictive performance is influenced not only by training-data volume but also by factors such as model architecture and confidence calibration. Training-overlap analysis further showed that AI-based methods retained substantial failure rates even for complexes represented in their training data, indicating that training-data overlap alone does not ensure reliable pose prediction. Under the current protocol conditions, rigid docking outperformed flexible protocols, while flexible-docking pocket volumes showed more restricted sampling relative to experimental holo structures. Among the GNINA 1.3 configurations, CNN rescoring with refinement produced the highest pose-recovery success rates, followed by CNN rescoring alone and the default Vina/empirical scoring approach in cross-docking. Receptor conformational preference was target-dependent: holo structures provided higher docking accuracy for AChE and BACE1, whose ligand-bound cavities exhibited greater structural complexity and geometric confinement that favoured pose discrimination, whereas the apo GSK-3β structure contained a larger, more solvent-exposed cavity that improved ligand accessibility and docking performance. Overall, these findings demonstrate the importance of cross-docking and training-overlap-aware evaluation for assessing docking performance under realistic conditions and provide cavity-topology-based considerations for selecting docking strategies in structure-based drug discovery.

Kapali Suri, Anshul Yadav, A. Tripathi et al. · 0 citations
Preprint Aug 2026

PSLL: Persistent Sheaf Laplacian Learning for Protein-Ligand Binding Affinity Prediction

Accurate prediction of protein-ligand binding affinity remains a central challenge in computational drug discovery due to the complex interplay among molecular geometry, physicochemical interactions, and atom-specific charge information. In this work, we introduce a Persistent Sheaf Laplacian learning (PSLL) framework for protein-ligand binding affinity prediction. The proposed approach constructs multiscale topological representations from three-dimensional protein-ligand complexes by incorporating atomic partial charges into sheaf restriction maps over Vietoris-Rips and alpha complex filtrations. To capture chemically diverse protein-ligand interactions, we introduce element-specific and category-specific atom-pair representations within the PSLL framework. Harmonic and non-harmonic spectra extracted from the resulting persistent sheaf Laplacians are used as molecular descriptors. To complement the PSLL-derived molecular representation, we incorporate transformer-based protein embeddings and SMILES-derived ligand descriptors for binding affinity prediction. The scoring power of the proposed multiscale PSLL model is validated against existing state-of-the-art methods on three widely used PDBbind benchmark datasets, including PDBbind-v2007, PDBbind-v2013, and PDBbind-v2016. The computational results indicate that the proposed PSLL model achieves strong predictive performance across benchmark datasets, highlighting its potential as an interpretable and mathematically grounded framework with promising generalizability for molecular machine learning and drug discovery.

Mushal Zia, Benjamin Jones, Guo-Wei Wei · 0 citations
Open access Jul 2026

Machine Learning Integrated Designing and Screening of 8-Hydroxyquinoline Based Metallo-β-Lactamase Inhibitors

The rapid emergence of metallo-b-lactamase-mediated antibiotic resistance has created an urgent need for new inhibitor discovery strategies. In this work, a machine-learning-guided workflow was developed to generate and prioritize potential inhibitors targeting NDM-1. A SMILES-based variational autoencoder was first pretrained on a broad molecular dataset to learn general chemical syntax and latent molecular representations. The model was then fine-tuned on an 8-hydroxyquinoline-enriched dataset to bias molecular generation toward zinc-binding chemical space relevant to metallo-β-lactamase inhibition. Generated compounds were processed through structural filtering and docking-based evaluation to create training data for downstream predictive modeling. Molecular fingerprints and physicochemical descriptors were then used to train XGBoost models for docking score prediction and classification of potential binders. Classification proved especially useful for prescreening because it avoided overinterpreting small differences in noisy docking scores while still enriching for compounds likely to perform well in docking. The resulting workflow demonstrates how generative modeling and supervised machine learning can be combined to reduce chemical search space, prioritize candidate inhibitors, and guide computational drug discovery. Although experimental validation remains necessary, this approach provides a scalable framework for identifying promising zinc-binding compounds for further molecular simulation and inhibitor development that can be expanded in future studies.

Anthony M. Baudino, Kari L. Stone · 0 citations

Related blog posts

GPT-Lab Aug 28, 2026

We built an AI factory for HVAC control

What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.

Microsoft Research Blog Jul 30, 2026

EvoLib: Turning experience into evolving knowledge

LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.