Skip to content
Open access

Multiclass Machine Learning-Based Discovery of Novel Scaffold Inhibitors Targeting ALK

Aug 2026 · Pharmaceuticals · Vol 19 · 0 citations · 62 references
Medicine

TL;DR

This integrated ML-to-simulation workflow prioritizes structurally novel candidate hits with predicted ALK inhibitory activity and provides an effective strategy for scaffold discovery and hit prioritization.

Abstract

Background: Anaplastic Lymphoma Kinase (ALK) is an oncogenic receptor tyrosine kinase implicated in several cancers. Despite the clinical success of ALK inhibitors, acquired resistance continues to drive the search for novel chemotypes. We developed a multiclass machine learning framework to classify ALK inhibitory activity using a curated ChEMBL dataset. Methods: Models were built using 2D molecular descriptors together with MACCS and ECFP4 fingerprints. Three widely used algorithms, Support Vector Machine (SVM), Random Forest (RF), and XGBoost, were applied for model development. Results: RF and XGBoost models demonstrated the best performance, achieving accuracies of ~0.75–0.79 with consistently high ROC–AUC values, particularly for fingerprint-based features. Bemis–Murcko scaffold analysis identified enriched chemotypes and underexplored scaffolds for further prioritization. The validated models were subsequently used to screen the Maybridge library, and compounds predicted to possess potential ALK inhibitory activity were prioritized for further computational evaluation. Applicability-domain filtering confirmed that the selected compounds occupied the predicted ALK inhibitor chemical space across multiple activity classes. The shortlisted compounds were subsequently evaluated by molecular docking to characterize their binding modes and interactions. Three candidate hits (SCR00078, SCR00073, and AW01085) were selected for further evaluation using 500 ns molecular dynamics simulations alongside the reference inhibitor Brigatinib. Simulation analyses revealed stable protein–ligand complexes and reduced conformational fluctuations relative to apo ALK, while MM/PBSA calculations identified SCR00078 and AW01085 as the most favorable binders. Conclusions: This integrated ML-to-simulation workflow prioritizes structurally novel candidate hits with predicted ALK inhibitory activity and provides an effective strategy for scaffold discovery and hit prioritization.

Read PDF

Similar papers

#protein folding Aug 2026

P1.062. Discovery of Novel Dual-Target Inhibitors for EGFR and PIK3CA From Natural Products via Machine Learning and Molecular Simulation

The natural product compounds CNP0456830 and CNP0467494 exhibited the lowest binding free energies for both EGFR and PIK3CA, identifying them as the most promising dual-target inhibitors.

Si-miao Lu, Yi Zhu, Yong-tao Han et al. · 0 citations
Open access Jul 2026

Computational design and AI driven discovery of anaplastic lymphoma kinase inhibitors for non small cell lung cancer treatment

Anaplastic Lymphoma Kinase (ALK), a tyrosine receptor kinase is of immense importance in non small cell lung cancer (NSCLC). Therefore, there is a need to design novel derivatives with the intention to overcome the limitations of resistance and debilitating side effects associated with current FDA-approved ALK inhibitors. This work therefore employed artificial intelligence and computational techniques to screen a library of ALK tyrosine kinase inhibitors, retrieved from ChEMBL database, with their corresponding IC50 in nM. The compounds’ descriptors were obtained using the Padel descriptors software and screened to reduce dimensionality and remove redundancy and multicollinearity. The compounds’ descriptors and their corresponding IC50 in nM were imported to Google Colab workspace with the necessary Python packages for machine learning (ML) models building. The best model was used to predict the bioactivity of new derivatives of 5-FDA approved drugs and TPX-1301, taking into account their drug-likeness properties for initials screening. The binding affinities and modes of lead compounds at the binding domain of ALK tyrosine kinase receptors were predicted using molecular docking while the binding free energies were obtained using MM-GB/SA calculations. Among all the trained models, the artificial neural network showed the most promising results, with a coefficient of determination (R2) value of 0.84, a root mean square error (RMSE) value of 0.27, a mean squared error (MSE) value of 0.22 and a mean absolute error (MAE) of 0.26. on the training data and an R2 value of 0.62, RMSE of 0.73, MSE of 0.53 and MAE of 0.55 for the test data, an indication of its reliability in making prediction. Cross-docking of cognate lorlatinib against ALK (4CLI) model yielded a binding mode closely aligned with the native conformer of lorlatinib, exhibiting an RMSD of 0.13 Å. Additionally, Induced Fit Docking (IFD) and Prime MM-GB/SA calculations indicated that briga_15 possesses the highest IFD Score of -675.27 kcal/mol, MM-GB/SA binding energy of -61.60 kcal/mol and a predicted pIC50 value of 8.57. The binding of briga_15 is facilitated by critical hydrogen bond networks with His1124 and Met1199 of ALK. Within the crizotinib series, crizo_35 demonstrated an exceptional IFD score of -671.16 kcal/mol, a predicted experimental pIC50 value of 7.75, and a binding energy of -77.71 kcal/mol, with water-mediated hydrogen bond networks involving Asp1203 and Lys1150 of ALK. These findings suggest that briga_15 and crizo_35 are promising leads warranting further optimization, synthesis, and biological evaluation.

O. Oyeneyin, N. Ipinloju, N. Gumede · 0 citations
Open access Jul 2026

Machine learning–driven identification of PIM2 kinase inhibitors through QSAR modeling and molecular dynamics simulations

An integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics simulations, and pharmacokinetic prediction, highlighting the potential of the identified molecules as promising molecules.

A. Fahira, M. Shahab, Zaheer Ud Din et al. · 0 citations
Open access Aug 2026

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Jurica Novak · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.