Skip to content
Open access

Machine learning–driven identification of PIM2 kinase inhibitors through QSAR modeling and molecular dynamics simulations

Jul 2026 · Journal of Genetic Engineering and Biotechnology · Vol 24, pp. 100759 · 0 citations · 71 references

TL;DR

An integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics simulations, and pharmacokinetic prediction, highlighting the potential of the identified molecules as promising molecules.

Abstract

The proto-oncogene serine/threonine kinase PIM2 is a critical regulator of cell proliferation, survival, and tumor progression and represents an attractive therapeutic target for several cancers. In this study, an integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics (MD) simulations, and pharmacokinetic prediction. Bioactivity data for PIM2 inhibitors were retrieved from the ChEMBL database, yielding 5953 compounds. After data cleaning, structural standardization, and removal of duplicates and invalid entries, a curated dataset of 1584 compounds was obtained for QSAR modeling. To address dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied before model development. Twelve molecular fingerprint descriptors were generated and used to construct 180 QSAR models using five machine learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Regression (SVR), k-Nearest Neighbors (KNN), and Multilayer Perceptron (MLP). Among these models, the Random Forest–fingerprint model demonstrated the best predictive performance, achieving a mean R2 of 0.971 with low prediction errors (RMSE = 0.271; MAE = 0.125) across training, testing, and cross-validation datasets. The optimized model was subsequently applied to virtual screening of multiple chemical libraries, including FDA-approved drugs, natural product databases, and commercial compound collections. Several promising candidates were identified, including TCMBANKIN000009 (emetine), Amb28533044 (4,6′-Anhydrooxysporidinone), NPC170963 (Lysophosphatidylcholine (15:0)), NPC262768 (Endosulfan), and NPC469603 (8-hydroxyircinialactam A). Molecular docking showed that these compounds bind within the ATP-binding pocket of PIM2 kinase, forming interactions with key residues such as Lys62, Asp125, Asp128, and Glu168. Subsequent molecular dynamics simulations confirmed the stability of selected complexes, demonstrating reduced residue fluctuations, stable protein compactness, and persistent intermolecular interactions during the simulation. Furthermore, ADMET prediction suggested favorable pharmacokinetic and toxicity profiles for several compounds. Collectively, these findings highlight the potential of the identified molecules as promising PIM2 inhibitor candidates, providing valuable leads for future experimental validation and anticancer drug development.

Read PDF

Similar papers

Open access Aug 2026

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Jurica Novak · 0 citations
Open access Jul 2026

Machine Learning-Driven Discovery of Novel HER2 Inhibitors Through Integrated Virtual Screening and Molecular Dynamics Simulations

These findings support further in vitro and in vivo testing for developing new therapeutics against HER2-overexpressing breast cancer, highlighting two scaffolds with promising lead optimization potential.

Alhumaidi B. Alabbas, Safar M. Alqahtani · 0 citations
Jul 2026

Modeling Structure-Activity Relationships with Machine Learning to Identify DPP4 Inhibitors as potential Therapeutics for Type 2 Diabetes

Findings identify CP20 as a promising lead scaffold for the development of novel DPP4 inhibitors and demonstrate the effectiveness of an ensemble machine learning-guided computational framework for accelerating antidiabetic drug discovery.

Iqra Anwar, T. Chohan, Drakhshaan et al. · 0 citations
Sep 2026

Natural Product‐Derived ErbB1 Inhibitors Identified Through Machine Learning‐Based QSAR, Molecular Docking, and Molecular Dynamics Simulations

Cancer represents a major global health burden, with abnormal ErbB1 (EGFR) signaling implicated in multiple tumors. Despite the clinical availability of several ErbB1 inhibitors, their long‐term efficacy is often limited by drug resistance, adverse effects, and restricted chemical diversity, underscoring the need for novel inhibitory scaffolds. Natural products represent a largely underutilized source of structurally diverse bioactive compounds; however, their systematic exploration against ErbB1 has been hindered by the lack of robust, large‐scale predictive approaches. This work developed an integrated computational strategy to identify novel natural product‐derived ErbB1 inhibitors. A machine learning QSAR classification model based on XGBoost was trained on a curated dataset of 6953 ErbB1 inhibitors, achieving strong predictive performance (accuracy = 85.3%, AUC = 0.92). Applying this model to over 80,000 natural products from the ZINC database yielded a focused subset of high‐confidence candidates. Subsequent molecular docking analyses revealed that five compounds engage with key catalytic residues of ErbB1, notably MET793 and ASP855, in a manner comparable to clinically used inhibitors. Molecular dynamics simulations (100 ns) confirmed stable binding for four candidates, while in silico ADMET evaluations supported favorable drug‐like properties for most hits. Overall, this study identifies promising natural scaffolds for ErbB1 inhibition and provides a scalable computational framework to support future experimental validation and lead optimization.

Unknown authors · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.