Skip to content

Applicability-domain-aware cheminformatics and structure-based prioritization of HER2 kinase inhibitor analogs.

Aug 2026 · Computational biology and chemistry · Vol 125, pp. 109288 · 0 citations · 49 references
Medicine

Abstract

HER2/ERBB2 kinase inhibition provides a useful case study for connecting curated bioactivity data, ligand-based modelling, chemical-space design, and structure-based triage. This study developed an applicability-domain-aware workflow for prioritizing scaffold-constrained HER2 kinase inhibitor analogs. Exact nanomolar IC50 records from ChEMBL target CHEMBL1824 were cleaned, converted to pIC50 values, aggregated at molecule level, and encoded as 2,048-bit Morgan fingerprints. Tree-ensemble classifiers gave the strongest ligand-based performance, and repeated scaffold-split validation was used to estimate scaffold-level generalization. Four HER2-associated Bemis-Murcko frameworks were selected for constrained R-group enumeration, yielding 500,000 valid analogs. Composite screening combined predicted HER2 activity, maximum-reference Tanimoto similarity as an applicability-domain measure, preferred physicochemical filters, and parent-scaffold priority, followed by Butina diversity selection of 100 docking-ready ligands. Synthetic-feasibility triage indicated that 94 of the 100 final ligands passed conservative RDKit-based medicinal-chemistry criteria. Wild-type docking, Prime MMGBSA rescoring, interaction-fingerprint comparison, and profiling against L755S, T798I, V777L, and V842I HER2 mutant receptors prioritized GEN_0015905 as the leading wild-type MMGBSA analog and highlighted GEN_0268162, GEN_0394004, and GEN_0015892 as follow-up candidates with smaller predicted wild-type-to-mutant score shifts. The workflow provides a reproducible computational route for selecting experimentally testable HER2 analogs while keeping model extrapolation, chemical diversity, and structure-based interpretation explicit. Repeated validation also showed the expected drop from random-split to scaffold-split performance, with random forest ROC-AUC decreasing from 0.973 ± 0.006 under repeated random splits to 0.956 ± 0.015 under repeated scaffold splits.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.