Skip to content

A synergistic deep learning and machine learning framework for screening heterocyclic compounds against ALDH1A1.

Jul 2026 · Molecular diversity · 0 citations · 46 references
Medicine

TL;DR

An integrated computer-aided drug design (CADD) and artificial intelligence (AI) framework to systematically identify selective ALDH1A1 inhibitors from a heterocyclic compound library is developed and suggests that LDN-27219 exhibits favorable binding characteristics and represents a promising lead candidate for subsequent experimental validation.

View source

Similar papers

Open access Jul 2026

Integrating machine learning and structure-based simulations to prioritise Z4P as a mutation-resilient IRE1α inhibitor for breast cancer.

Breast cancer is one of the most prevalent and lethal malignancies affecting women globally. The increasing resistance to current therapeutic strategies highlights the need for novel molecular targets. Inositol-requiring enzyme 1 alpha (IRE1α), a key sensor in the unfolded protein response (UPR), has emerged as a promising therapeutic target due to its role in tumour progression and survival. This study employed an integrative in silico approach combining machine learning, molecular docking, and molecular dynamics simulations to identify potent, non-toxic IRE1α inhibitors for breast cancer treatment. An initial library of 115 compounds retrieved from ChEMBL and MedChemExpress was used for machine learning-based toxicity modelling. Literature curation identified 44 reported IRE1α inhibitors, which were reduced to 38 unique compounds following duplicate removal. Drug-likeness and ADMET screening using SwissADME and ProTox retained 22 compounds for further evaluation. Molecular docking was performed using AutoDock, followed by Dynamics simulations in GROMACS to assess stability. Machine Learning (ML) models were developed for both toxicity regression and binary toxicity classification analyses. Toxicity prediction models were developed using twenty physicochemical and pharmacokinetic descriptors. In the regression analysis, Random Forest demonstrated the strongest cross-validation performance (R² = 0.5998 ± 0.3439), while the stacking ensemble achieved the highest test-set performance (R² = 0.9765), although differences among ensemble methods were not statistically significant. In the complementary classification analysis, the Support Vector Machine (SVM) achieved the highest discriminative performance with an ROC-AUC value of 0.98. Docking studies revealed that Z4P exhibited the strongest binding affinity (- 7.93 kcal/mol) to the wild-type IRE1, compared with the control drug MKC8866 (- 6.7 kcal/mol). Additionally, Z4P exhibited a higher binding energy of - 9.5 kcal/mol, whereas MKC8866 had a binding energy of - 6.94 kcal/mol. MD simulations over 200 ns confirmed the stability of the IRE1-Z4P complex, with favourable RMSD, RMSF, Rg, and SASA profiles relative to the control. These findings highlight Z4P as a promising mutation-resilient IRE1 inhibitor and validate the effectiveness of the integrated computational pipeline for identifying potential anti-cancer therapeutics.

Nithisha L Bastin, P. K. Praveen Kumar, B. Ethiraj et al. · 0 citations
Open access Aug 2026

An Integrated Consensus Machine Learning and Structure-Based Workflow for the Discovery of Novel Tankyrase 1 Inhibitors

Background: Tankyrase 1 (TNKS1) is a poly(ADP-ribose) polymerase involved in Wnt/β-catenin signaling, telomere maintenance, and genomic stability, making it an attractive therapeutic target in oncology. This study aimed to develop and apply an integrated computational workflow to identify novel TNKS1 inhibitor candidates. Methods: A curated dataset of experimentally validated TNKS1 inhibitors and property-matched DUD-E decoys was used to develop a consensus supervised machine learning (ML) model prioritization framework for ligand-based virtual screening, integrating Morgan fingerprints with three complementary classifiers. The model screened more than 700,000 compounds, and prioritized hits were evaluated by structure-based virtual screening (SBVS), Prime MM-GBSA binding free-energy refinement, and 500 ns molecular dynamics simulations (MDs). The top candidates were subsequently tested in an in vitro TNKS1 enzymatic inhibition assay. Results: The consensus ML framework prioritized 670 compounds, yielding five candidates for experimental testing. Compound 3 displayed the most favorable computational profile and was experimentally confirmed as a TNKS1 inhibitor candidate, exhibiting approximately 80% TNKS1 inhibition at 0.1 μM, whereas the remaining candidates showed only limited activity. Conclusions: The proposed workflow efficiently reduced a large chemical space to a focused set of TNKS1 inhibitor candidates while substantially reducing the experimental screening burden. Compound 3 represents a promising starting point for future structure–activity relationship studies and lead optimization in the context of TNKS1 inhibition. Moreover, this work highlights the value of integrating consensus ML, SBVS, and experimental validation to accelerate early-stage hit discovery for TNKS1 and other therapeutic targets.

M. Bilotta, Adriana Gargano, R. Rocca et al. · 0 citations
Open access Aug 2026

Discovery of a potent TDP1 inhibitor through machine learning-driven predictive modeling combined with structure-based virtual screening and experimental validation

Tyrosyl-DNA phosphodiesterase I (TDP1) repairs topoisomerase I (TOP1)–mediated DNA damage and is a promising anticancer target, particularly in combination with TOP1 inhibitors. However, the discovery of potent and drug-like TDP1 inhibitors remains challenging due to the limited structural diversity of known active compounds. Here, we developed an integrated computational framework combining machine learning (ML), deep learning (DL), and structure-based docking with experimental validation. A curated dataset of 2040 compounds (857 active, 1183 inactive) was assembled and analyzed by scaffold composition. A total of 40 binary classification models were constructed using six ML algorithms and a deep neural network (DNN), each paired with five molecular fingerprint representations, along with five graph neural network architectures (GCN, GAT, MPNN, AttentiveFP, and FPGNN). The SVM::RDKitDes model performed best (AUC = 0.89, F1 = 0.78, BA = 0.80), with robustness confirmed by Y-scrambling and randomized-split analyses, and SHAP analysis identified 20 key descriptors of TDP1 inhibition. The model was deployed as a web application (http://drugpred.top:5050) and standalone desktop applications (.exe) are available at https://github.com/zenghuang8006/TDP1-inhibitor-prediction. The validated model was applied to screen 201 231 compounds, followed by drug-likeness filtering and hierarchical docking, yielding 16 candidates. Biological evaluation identified compound AO65 as a potent TDP1 inhibitor (IC50 = 0.80 ± 0.02 µM), and quantum chemical calculations and docking elucidated its electronic properties and binding within the catalytic domain. This work demonstrates the value of integrating ML-driven prediction with structure-based approaches and identifies AO65 as a promising lead for further TDP1-focused investigation.

Huang Zeng, Manyi Zhang, Bo Qiu et al. · 0 citations
Open access Jul 2026

Machine Learning Integrated Designing and Screening of 8-Hydroxyquinoline Based Metallo-β-Lactamase Inhibitors

The rapid emergence of metallo-b-lactamase-mediated antibiotic resistance has created an urgent need for new inhibitor discovery strategies. In this work, a machine-learning-guided workflow was developed to generate and prioritize potential inhibitors targeting NDM-1. A SMILES-based variational autoencoder was first pretrained on a broad molecular dataset to learn general chemical syntax and latent molecular representations. The model was then fine-tuned on an 8-hydroxyquinoline-enriched dataset to bias molecular generation toward zinc-binding chemical space relevant to metallo-β-lactamase inhibition. Generated compounds were processed through structural filtering and docking-based evaluation to create training data for downstream predictive modeling. Molecular fingerprints and physicochemical descriptors were then used to train XGBoost models for docking score prediction and classification of potential binders. Classification proved especially useful for prescreening because it avoided overinterpreting small differences in noisy docking scores while still enriching for compounds likely to perform well in docking. The resulting workflow demonstrates how generative modeling and supervised machine learning can be combined to reduce chemical search space, prioritize candidate inhibitors, and guide computational drug discovery. Although experimental validation remains necessary, this approach provides a scalable framework for identifying promising zinc-binding compounds for further molecular simulation and inhibitor development that can be expanded in future studies.

Anthony M. Baudino, Kari L. Stone · 0 citations