Skip to content

Bridging between Structure-Based and Data-Driven Affinity Prediction.

Jul 2026 · Journal of Chemical Information and Modeling · 1 citation · 33 references
Medicine

TL;DR

This work introduces a method to smoothly transition from physics-based to knowledge-based predictions based on the uncertainty of each model and shows that combining structure-based and ML models significantly improves the prediction accuracy if training data is limited, whereas the weighting smoothly shifts from docking to ML as more data is acquired.

Abstract

Protein-ligand binding affinity prediction is central in computer-aided drug design, and a wide range of physics-based, empirical, and machine learning (ML) tools have been developed for this purpose. Scientists are often tasked with choosing an optimal tool for a given discovery problem, which can be difficult. Physical models typically perform best in the absence of target-specific data but are often outperformed by ML models as the amount of data on a given target grows. Here, we introduce a method to smoothly transition from physics-based to knowledge-based predictions based on the uncertainty of each model. We apply this framework to combine docking scores with predictions from a Gaussian Process model trained on binding affinities from two industrial data sets. We show that combining structure-based and ML models significantly improves the prediction accuracy if training data is limited, whereas the weighting smoothly shifts from docking to ML as more data is acquired. We also show that structure-based methods, being insensitive to the distribution of binders across chemical space, can improve generalizability to new chemistries and increase the hit rate in active learning screens where the binder density varies among scaffolds in a virtual library. Thus, integrating predictions from multiple tools not only optimizes the use of limited experimental data but also ensures more robust performance compared to reliance on a single model.

View source

Similar papers

Open access Aug 2026

Synthetic data for more accurate deep learning models in molecular science: a test case of protein-ligand binding affinity prediction

This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures, and highlights that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.

P. Agrawal, Prathit Chatterjee, U. Priyakumar · 0 citations
Open access Jul 2026

A Preparation-Free Mixture-of-Experts Framework for Protein-Ligand Affinity Prediction

The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.

Huiming Bao, Shouliang Dong · 0 citations
Open access Jul 2026

Reliability of AI Methods in Drug Discovery: Evaluation of Boltz‑2 for Structure and Binding Affinity Prediction

Despite continuing hype about the role of AI in drug discovery, no “AI-discovered drugs” have so far received regulatory approval. Here we assess one of the latest AI-based tools in this domain. Boltz-2, a recently developed biomolecular foundation model, aims to bridge the gap between AI efficiency and physics-based precision through a joint “cofolding” approach. In this study, we provide an extensive evaluation of Boltz-2 using two large-scale data sets: 16780 compounds for 3CLPro and 21702 compounds for TNKS2. We compare Boltz-2 predicted structures with traditional docking and binding affinities with binding free energies derived from the physics-based ESMACS protocol. Structural analysis reveals significant global RMSD variations, indicating that Boltz-2 predicts multiple protein conformations and ligand binding positions rather than a single converged pose. Energetic evaluations exhibit only weak to moderate correlations across the global data sets. Furthermore, a focused analysis of the top 100 compounds yields no significant correlation between the Boltz-2 predictions and the binding free energies from fine-grained ESMACS, alongside frequently observed saturation-state errors in Boltz-2 predicted ligand structures. Our results show that Boltz-2 lacks the energetic resolution required for lead identification. These findings highlight the necessity of employing physics-based methods for the reliability and refinement of AI-derived models.

S. Wan, Xibei Zhang, Xiao Xue et al. · 0 citations
Open access Aug 2026

On the generalization and usability of cofolding models for GPCR drug discovery

Boltz is benchmarked using a curated set of ligand-bound human G protein-coupled receptors from families unseen during training, showing that while Boltz generally predicts receptor backbones accurately, ligand poses can contain significant errors that lead to a limited ability to reproduce experimental affinity data when tested with FEP+.

Lichirui Zhang, R. Friesner, Edward B. Miller et al. · 0 citations
Conference Jul 2026

A Drug-Target Affinity Prediction Model Based on Bayesian Meta-Learning and Uncertainty Fusion

In the early stages of drug discovery, predicting drug-target affinity is a crucial task. Due to the vast scale of genomic and chemical spaces, traditional biological methods are time-consuming, labor-intensive, and resource-demanding. As a result, machine learning-based computational methods have emerged to narrow down the pool of drug candidates. However, machine learning approaches still face several challenges in practical applications, particularly the scarcity of labeled samples and poor model generalization capability. To address these issues, this paper proposes a novel drug-target affinity prediction model, termed MetaBayes-DTA, based on an uncertainty-aware meta-learning framework. The model integrates the few-shot rapid adaptation capability of meta-learning with an uncertainty quantification mechanism to enhance prediction accuracy and reliability. MetaBayes-DTA is evaluated on two benchmark datasets, DAVIS and KIBA. Experimental results demonstrate that the proposed model outperforms existing methods.

Naihan Shi, Yanpeng Zhao, Wanying Li et al. · 0 citations
Jul 2026

Conditional molecular dynamics refinement for protein-ligand affinity prediction.

Protein-ligand affinity prediction is fundamental to structure-based virtual screening and lead discovery. However, most existing methods rely on a single static conformation of the complex and exhibit high sensitivity to pose uncertainty and conformational noise. To address this limitation, CMD-PLA is proposed as a dynamics-aware framework for protein-ligand affinity prediction. Within this framework, pocket-conditioned molecular dynamics refinement is performed, and the conformational evolution of the ligand is explicitly modeled as a pocket-dependent dynamical process rather than an isolated static update. Furthermore, a dual-view atomic representation is adopted to separately capture the intra-molecular covalent structure and the inter-molecular interaction geometry. Global representations of the ligand and the pocket are also incorporated to complement the local geometric modeling. Experimental results demonstrate that CMD-PLA achieves robust performance across various settings. The provided case study further illustrates the interpretability of the model.

Hao Li, Dongjiang Niu, Xiaofeng Wang et al. · 0 citations