Skip to content
Open access

ERβ-Score: An Interpretable Machine Learning-Based Scoring Function and Web Server for Estrogen Receptor β-Guided Drug Discovery in Triple-Negative Breast Cancer

Aug 2026 · International Journal of Molecular Sciences · Vol 27, pp. 7089 · 0 citations · 46 references
Medicine

TL;DR

ERβ-Score is introduced, an interpretable machine learning scoring system developed using a curated dataset of 1699 ERβ bioactive chemicals obtained from ChEMBL, characterized by 39 physicochemical and three-dimensional molecular descriptors, indicating strong and balanced discrimination between active and inactive ERβ modulators.

Abstract

Triple-negative breast cancer (TNBC) is the most clinically aggressive subtype of breast cancer, characterized by the absence of targetable hormone receptors and HER2 amplification, significantly constraining treatment choices. Estrogen Receptor Beta (ERβ) has emerged as a biologically relevant yet underutilized target in TNBC, with its re-expression linked to tumor suppression and improved prognosis, prompting the development of selective ERβ modulators as a precision therapeutic approach. We introduce ERβ-Score, an interpretable machine learning scoring system developed using a curated dataset of 1699 ERβ bioactive chemicals obtained from ChEMBL, characterized by 39 physicochemical and three-dimensional molecular descriptors. After implementing scaffold-disjoint train/test partitioning to avert structural data leakage, a Gradient Boosting Classifier, fine-tuned through Bayesian hyperparameter optimization, attained in five-fold cross-validation a Precision–Recall AUC (Area Under the Curve) of 0.891, a ROC-AUC (Receiver Operating Characteristic) of 0.888, a Matthews Correlation Coefficient of 0.664, an F1-score of 0.838, and a balanced accuracy of 0.831; on the scaffold-disjoint hold-out test set it attained a Precision–Recall AUC of 0.905, a ROC-AUC of 0.864, and a Matthews Correlation Coefficient of 0.578, indicating strong and balanced discrimination between active and inactive ERβ modulators. We note explicitly that this scaffold-disjoint hold-out constitutes internal validation, since it derives from the same curated ChEMBL workflow used for model development, and it is therefore reported throughout as scaffold-disjoint internal validation rather than as independent external validation. The applicability domain boundaries were established using a k-nearest-neighbor Tanimoto-similarity method with ECFP4 (Extended-Connectivity Fingerprint with a Diameter of 4) fingerprints, offering a quantitative confidence metric that identifies structurally new molecules beyond the model’s reliable prediction range. External validation against independent Tox21 ERβ bioassay data confirmed genuine, statistically significant predictive signal (ROC-AUC = 0.71) while revealing reduced sensitivity for structurally novel active compounds. The model was subsequently used for extensive virtual screening of natural product and drug-like compound libraries, with prioritized candidates undergoing structure-based molecular docking against the ERβ co-crystal structure (PDB: 7XWQ) using Smina, facilitating a comprehensive evaluation of hits based on both ligand and structural properties. To enhance accessibility, the complete pipeline was implemented as an open-access interactive web application utilizing Streamlit, enabling researchers to input any SMILES string and obtain, in real time, an activity prediction with a probability score, applicability domain classification, Lipinski drug-likeness assessment, interactive three-dimensional visualization of protein–ligand interactions, and on-demand docking within the ERβ active site.

Read PDF

Similar papers

#protein folding Aug 2026

P1.062. Discovery of Novel Dual-Target Inhibitors for EGFR and PIK3CA From Natural Products via Machine Learning and Molecular Simulation

The natural product compounds CNP0456830 and CNP0467494 exhibited the lowest binding free energies for both EGFR and PIK3CA, identifying them as the most promising dual-target inhibitors.

Si-miao Lu, Yi Zhu, Yong-tao Han et al. · 0 citations
Open access Sep 2026

Interpretable machine learning for individualized survival prediction in node-positive, non-metastatic prostate cancer: a population-based study.

BACKGROUND Prostate cancer with regional lymph node involvement but no distant metastasis (N1M0) has heterogeneous prognosis. This study aimed to develop and validate an interpretable machine learning model for predicting cancer-specific survival (CSS). METHODS Data from 18,287 N1M0 patients (2000-2022) in the SEER database were divided into training (n = 4780), internal testing (n = 1193), and two temporal validation cohorts (n = 3149; n = 9165). Cox proportional hazards and four machine learning models were compared using C-index, time-dependent AUC, and Integrated Brier Score. SHAP was used for interpretability, and IPTW for sensitivity analysis. RESULTS Random Survival Forest (RSF) outperformed all models, achieving the highest C-index (0.697 internal; 0.740 temporal validation). RSF-stratified risk groups showed significant survival differences (P < 0.001). SHAP revealed radical prostatectomy as the strongest protective factor, followed by lower T stage and PSA. IPTW confirmed survival benefits of aggressive local control. CSS outperformed overall survival in discriminative accuracy. CONCLUSION The RSF model provides accurate, robust, and interpretable prognostication for N1M0 prostate cancer. Deployed as a web-based calculator, it enables precise risk stratification and individualized treatment planning.

Unknown authors · 0 citations
Open access Jul 2026

Discovery of Novel AURKA Inhibitors for Triple-Negative Breast Cancer (TNBC) Therapy via a Hybrid Virtual Screening Pipeline, Biological Evaluation and Molecular Dynamics Simulation

A cascaded AI-driven virtual screening pipeline is developed, integrating sequence-based affinity prediction, equivariant deep learning docking (KarmaDock), and geometric rescoring (DeepDock) to identify novel AURKA inhibitor candidates.

Pei Liu, Yanqing Liu, Xin Zhang et al. · 0 citations
Open access Jul 2026

PredictRx: AI based decision support tool for molecular screening for breast cancer drug recommendation

PredictRx shows how AI-driven predictive modeling which can speed up molecular screening and early-stage breast cancer medication discovery shows how AI-driven predictive modeling can speed up molecular screening and early-stage breast cancer medication discovery.

Ritu Chauhan, Neha Pandey, M. Zuhairi · 0 citations
Review Aug 2026

The 21-gene recurrence score assay as a tool for predicting recurrence risk and guiding adjuvant treatment selection in early breast cancer

ABSTRACT Introduction Estrogen receptor-positive (ER+), HER2-negative breast cancer is the most common breast cancer subtype. While adjuvant endocrine therapy reduces recurrence risk, identifying which patients benefit from the addition of chemotherapy remains a key clinical challenge. The Oncotype DX® 21-gene Recurrence Score assay (Exact Sciences, via Genomic Health, Inc.) was developed to address this by quantifying distant recurrence risk and informing chemotherapy decisions in early-stage ER+/HER2− disease. Areas covered This diagnostic profile reviews the development, validation, and clinical evidence for Oncotype DX, including findings from the TAILORx and RxPONDER prospective trials and the subsequent development of hybrid tools integrating genomic and clinicopathological data. Alternative multiparameter molecular tests (MammaPrint, Prosigna, EndoPredict, Breast Cancer Index) are summarized and compared. We review international guideline recommendations, decision impact studies, cost-effectiveness evidence, and ongoing trials. Expert opinion Oncotype DX has strong prognostic evidence and has meaningfully reduced chemotherapy use, though its case as a biomarker predictive of therapeutic effect from chemotherapy rests on trial designs with important limitations. Its independent prognostic contribution beyond comprehensive clinicopathological assessment requires further clarification, and cost-effectiveness varies substantially by indication and healthcare setting.

C. Martínez-Pérez, Giovanni Tramonti, C. Kay et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.