Synthetic data for more accurate deep learning models in molecular science: a test case of protein-ligand binding affinity prediction
Deep learning models are data-hungry, and synthetic (artificial) data has been shown to be invaluable when data availability is low. While this has been demonstrated in certain technology areas, adopting such an approach is new in machine learning (ML) applications in chemistry, except for some pre-training tasks. In drug discovery, predicting binding energy between proteins and ligands is crucial. Many ML-based studies have been proposed to predict protein-ligand binding affinity using existing experimental data. However, these models suffer from inherent biases. Recent efforts have produced PLAS-20k, a synthetic dataset of multiple protein-ligand complex (PLC) conformations generated using molecular dynamics (MD) simulations as a viable option to complement existing experimental data and improve binding affinity prediction. For the binding affinity prediction task, we employ Pafnucy, a deep convolutional neural network, and propose using multiple structures for each PLC from PLAS-20k for training. We compare four different statistical and ML-based result-aggregation techniques. This work demonstrates the utility of dynamic datasets in enhancing binding affinity predictions, laying the foundation for future improvements in predicting similar protein properties using synthetic datasets and more sophisticated models and methods. We propose that physics-based synthetic datasets can significantly help develop more accurate data-driven methods. Scientific contribution This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures. By systematically evaluating frame selection and prediction aggregation strategies, we demonstrate that training on diverse conformational snapshots significantly enhances generalization and accuracy. Our results highlight that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.