Skip to content
Open access

Explainable Machine Learning-Guided Virtual Screening and Molecular Docking for Identification of Novel FYN Kinase Inhibitors

Jul 2026 · Intelligent Systems Research and Applications Journal · Vol 2, pp. 340-363 · 4 citations

TL;DR

How explainable ML can be used in combination with structure-based docking to expedite the drug discovery process against FYN kinase is highlighted and can be applied to other kinase targets.

Abstract

FYN kinase is a non-receptor protein tyrosine kinase involved in various cancers and neurodegenerative diseases; however, no selective FYN inhibitor has been approved yet. Here we introduce the explainable Machine Learning (ML) coupled with virtual screening and Molecular Docking (MD) pipeline for fast prediction of new FYN kinase inhibitors. In this study, we constructed the training set of 906 molecules active against FYN kinase from the ChEMBL database. Molecules were encoded with Extended-Connectivity Fingerprints (ECFP4). The classification models Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) were developed, and the latter showed the better performance in test (AUC=0.8118) and 5-fold cross-validation (AUC=0.8297). Based on the SHapley Additive exPlanations (SHAP) values obtained via TreeExplainer, nitrogen-containing heterocycles and hydrogen bond acceptors have been identified as the most important molecular substructures. Using the optimal XGBoost classifier, screening of 2,000 approved drugs has been performed, resulting in 470 hit molecules (23.5% hit rate). Five best molecules were further submitted to the MD procedure using AutoDock Vina to dock to FYN kinase domain (PDB RCSB: 2DQ7), showing binding energies in the interval of -9.57 to -6.32 kcal/mol. Dasatinib Anhydrous (CHEMBL1421) was the second strongest binder (-8.49 kcal/mol), effectively interacting with the ATP binding site. Although CHEMBL1171837 was the strongest binder (-9.57 kcal/mol), it was caught in the ADMET profiling. According to ADMET profiling, the top one inhibitor (CHEMBL1421) satisfies Lipinski’s rule of five and Veber rules. Analysis of hydrogen bond and hydrophobic interactions revealed hydrogen bonding with ASP148, LYS39, and ASN86 and hydrophobic interactions with ALA147, ILE80, and GLY88. Validation by self-docking procedure (self-docking or STS) showed low Root Mean Square Deviation (RMSD)<2.0 Å with a binding affinity of -11.53 kcal/mol. This work highlights how explainable ML can be used in combination with structure-based docking to expedite the drug discovery process against FYN kinase and can be applied to other kinase targets.

Read PDF

Similar papers

Open access Aug 2026

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Jurica Novak · 0 citations
Aug 2026

Harnessing machine learning, docking and molecular dynamics for the virtual screening of compounds as CDK4/6 dual inhibitors.

It is elucidated that HY-18,623 achieves high-affinity binding through persistent hydrogen bonds with hinge residues Val96 (CDK4) and Val101 (CDK6) and Val101 (CDK6) as a promising lead candidate for further therapeutic development in oncology.

Yu-Xin Wang, Lin-Xia Fang, Quan-Fang Liu et al. · 0 citations
Open access Jul 2026

Machine-learning-based pIC50 prediction identifies novel acetylcholinesterase inhibitors among FDA-approved drugs

Acetylcholinesterase (AChE) remains one of the most validated therapeutic targets for the symptomatic treatment of Alzheimer’s disease. Despite decades of research, the repertoire of approved AChE inhibitors is limited, and the identification of novel chemotypes with inhibitory potential is an active area of investigation. We present an automated, reproducible machine-learning pipeline for the prediction of AChE inhibitory activity among approved drugs. Bioactivity records (9731 IC50 measurements) were retrieved from the ChEMBL database for human AChE and preprocessed into a dataset of 6050 compounds. Each molecule was encoded with a dual descriptor set comprising approximately 1600 Mordred 2D physicochemical descriptors and 2048-bit Morgan fingerprints, yielding 3571 features after variance filtering. Two XGBoost models—a regressor for continuous pIC50 prediction and a classifier for binary activity assignment—were independently optimized through Bayesian search over a scaffold-stratified GroupKFold cross-validation scheme to prevent data leakage between related compounds. On a held-out test set of 1076 molecules, the regressor achieved a mean absolute error of 0.661 log units and R2 = 0.642, while the classifier attained an area under the receiver operating characteristic curve of 0.905, an area under the precision-recall curve of 0.895, and an F1 score of 0.796. Model interpretability was assessed via SHapley Additive exPlanations analysis, which highlighted contributions of physicochemical descriptors and topological substructures. An in silico screen of 2062 approved drugs from the DrugBank database identified 256 compounds (12.4%) as predicted active against AChE, including 19 drugs with documented AChE activity in ChEMBL and 235 novel repurposing candidates with no prior AChE record. The pipeline is publicly available as an open-source tool, readily adaptable to other pharmacological targets, facilitating drug-repurposing efforts. Sensitivity analysis confirmed the robustness of the binary activity threshold across alternative cut-offs, and an ablation study demonstrated that the combined Mordred–Morgan feature set yields better cross-validation performance than either descriptor family alone.

Sveva Bonomi, Marco Buglione, Crescenzo Edoardo Mauriello et al. · 0 citations
Open access Aug 2026

Multiclass Machine Learning-Based Discovery of Novel Scaffold Inhibitors Targeting ALK

This integrated ML-to-simulation workflow prioritizes structurally novel candidate hits with predicted ALK inhibitory activity and provides an effective strategy for scaffold discovery and hit prioritization.

José Zarzuelo Romero, M. López-Viota, M. Haque et al. · 0 citations
Open access Jul 2026

Machine Learning-Driven Discovery of Novel HER2 Inhibitors Through Integrated Virtual Screening and Molecular Dynamics Simulations

These findings support further in vitro and in vivo testing for developing new therapeutics against HER2-overexpressing breast cancer, highlighting two scaffolds with promising lead optimization potential.

Alhumaidi B. Alabbas, Safar M. Alqahtani · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.