Skip to content
Open access

Synthetic data for more accurate deep learning models in molecular science: a test case of protein-ligand binding affinity prediction

Aug 2026 · Journal of Cheminformatics · 0 citations

TL;DR

This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures, and highlights that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.

Abstract

Deep learning models are data-hungry, and synthetic (artificial) data has been shown to be invaluable when data availability is low. While this has been demonstrated in certain technology areas, adopting such an approach is new in machine learning (ML) applications in chemistry, except for some pre-training tasks. In drug discovery, predicting binding energy between proteins and ligands is crucial. Many ML-based studies have been proposed to predict protein-ligand binding affinity using existing experimental data. However, these models suffer from inherent biases. Recent efforts have produced PLAS-20k, a synthetic dataset of multiple protein-ligand complex (PLC) conformations generated using molecular dynamics (MD) simulations as a viable option to complement existing experimental data and improve binding affinity prediction. For the binding affinity prediction task, we employ Pafnucy, a deep convolutional neural network, and propose using multiple structures for each PLC from PLAS-20k for training. We compare four different statistical and ML-based result-aggregation techniques. This work demonstrates the utility of dynamic datasets in enhancing binding affinity predictions, laying the foundation for future improvements in predicting similar protein properties using synthetic datasets and more sophisticated models and methods. We propose that physics-based synthetic datasets can significantly help develop more accurate data-driven methods. Scientific contribution This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures. By systematically evaluating frame selection and prediction aggregation strategies, we demonstrate that training on diverse conformational snapshots significantly enhances generalization and accuracy. Our results highlight that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.

Read PDF

Similar papers

Jul 2026

Bridging between Structure-Based and Data-Driven Affinity Prediction.

This work introduces a method to smoothly transition from physics-based to knowledge-based predictions based on the uncertainty of each model and shows that combining structure-based and ML models significantly improves the prediction accuracy if training data is limited, whereas the weighting smoothly shifts from docking to ML as more data is acquired.

Ažbeta Kubincová, David L. Mobley · 1 citation
Open access Jul 2026

A Preparation-Free Mixture-of-Experts Framework for Protein-Ligand Affinity Prediction

The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.

Huiming Bao, Shouliang Dong · 0 citations
Review Aug 2026

Recent Advances in Deep Learning-Based Drug-Target Binding Affinity Prediction

Computational approaches to drug discovery involve multiple sub-problems, and among them, drug-target binding affinity prediction plays an important role. Despite recent advances, accurately predicting binding affinity remains an open research area. The major objective of our paper is to perform a comprehensive review and comparative analysis of recent machine learning methods for drug-target binding affinity prediction, with a focus on identifying strengths, limitations, and research gaps. We review representative recent deep learning approaches that use common benchmark datasets and evaluation metrics, covering a range of neural network architectures and representation strategies. In addition, we analyze seven widely used benchmark datasets and commonly adopted evaluation metrics for drug-target binding affinity prediction. Our analysis indicates that although many methods report strong performance on standard benchmarks, their effectiveness is often influenced by dataset bias and limited evaluation settings. Furthermore, most methods exhibit reduced performance in cold-start scenarios, highlighting challenges in generalization. We identify several limitations of current approaches, including dataset imbalance, the lack of standardized evaluation, limited real-world applicability, and challenges in cold-start scenarios. We also discuss future research directions, including better dataset design, more robust evaluation methods, improved handling of cold-start problems, and the integration of multimodal representations.

Jafin Khan, Md Hossain Shuvo · 0 citations
Review Open access Aug 2026

Geometric Deep Learning‐Based Drug Design Models for Small‐Molecule Drug Discovery

Deep neural network (DNN)‐based in silico models show great promise in predicting the properties and bioactivities of novel compounds, including small molecules. Among traditional approaches, structure‐based drug design (SBDD) remains a fundamental approach for drug discovery using molecular docking, scoring functions, and molecular dynamics simulations. However, these approaches are often constrained by limited flexibility, resolution, and generalizability. Geometric deep learning (GDL) offers a transformative alternative by enabling models to learn directly from non‐Euclidean molecular representations, such as graphs, point clouds, and meshes, capturing critical 3D spatial relationships inherent to protein–ligand interactions. This review highlights the theoretical underpinnings and practical applications of GDL in small‐molecule drug discovery, focusing on tasks including binding affinity prediction, virtual screening, de novo molecule generation, pose prediction, ADMET profiling, and protein flexibility modeling. We explore key GDL architectures, graph neural networks, SE(3)‐equivariant networks, 3D convolutional neural networks, point cloud models, and geometric transformers, and assess their performance across various drug discovery benchmarks. The integration of geometry‐aware AI models with experimental and computational workflows was also highlighted for its potential to streamline hit‐to‐lead optimization and advance rational drug design. Despite remarkable progress, the field faces challenges including limited high‐quality 3D structural datasets, protein flexibility representation, and the interpretability of deep models. Addressing these issues through hybrid modeling approaches, multi‐resolution learning, and self‐supervised training could further elevate GDL's impact. Ultimately, GDL stands at the frontier of AI‐enhanced pharmaceutical innovation, offering unprecedented precision, efficiency, and insight in the pursuit of next‐generation therapeutics.

A. Srivastav, Unnati Modi, Rahul Kumar et al. · 0 citations
Aug 2026

Deep Learning Foundation Models for Low-Data Regimes from Classical Molecular Descriptors

Fast and accurate data-driven prediction of molecular properties is pivotal to scientific advancements across myriad chemical domains. Deep learning methods have recently garnered much attention, despite their inability to outperform classical machine learning methods when tested on practical, real-world benchmarks with limited training data. This study seeks to bridge this gap by introducing a new avenue for foundation model pretraining. We propose pretraining on low-noise, calculable molecular descriptors via supervised learning to obtain rich, highly transferable molecular representations. We demonstrate this strategy with CheMeleon, a O(10M) parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods in the low-data regime. We evaluate on 58 benchmark data sets spanning a range of properties relevant to small-molecule drug discovery, sourced from the industry-led Polaris benchmarking initiative. Rigorous statistical comparisons show that CheMeleon outperforms classical baselines like Random Forest on molecular fingerprints and descriptors, as well as existing foundation models. We open-source the CheMeleon model and the pretraining framework to encourage adoption and extension of this pretraining strategy across chemical sciences.

Jackson W. Burns, Akshat Shirish Zalte, C. Abreu et al. · 0 citations
Preprint Aug 2026

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Bioassay activity prediction is often data-limited because drug-discovery datasets rely on time-consuming and expensive wet-lab experiments for data generation and evaluation. This challenge has inspired recent research into molecular foundation models (MFMs), which aim to encode general-purpose chemical knowledge into molecular representations that generalize well in data-constrained scenarios. This paper presents Monroe, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 81 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding decorrelation; improved multi-task learning; and the use of a prior-data-fitted model (TabPFN) for downstream in-context prediction. Our evaluations use a principled pairwise comparison framework that measures statistically significant performance differences. Across established Polaris benchmarks, Monroe matches or exceeds existing MFMs, while on activity cliff benchmarks, designed to assess utility for molecular discovery, it achieves significant improvements over prior methods. Finally, ablation and transfer experiments show that PFN-based downstream predictors also substantially improve two leading existing models, MiniMol and CheMeleon, yielding new state-of-the-art variants we call MiniMol_PFN and CheMeleon_PFN, suggesting that our downstream adaptation strategy generalizes beyond Monroe. Source code is at github.com/blazejba/monroe.

Blazej Banaszewski, Andrew W. Fitzgibbon · 0 citations