Back to #machine learning

Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

Aug 2026 · 0 citations · 31 references
Biology Computer Science

TL;DR

This work substantially expand the IMPROVE benchmark through large-scale integration of pharmacogenomic data, primarily from PharmacoDB, together with additional smaller data sources, which includes millions of drug response measurements, broader multi-omics coverage, and a major increase in chemical diversity, adding more than 50,000 compounds.

Abstract

Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs. However, their predictive performance is often constrained by limited dataset scale and insufficient coverages of cancer and chemical spaces. In addition, inconsistent benchmarking practices hinder reliable comparison across models. Standardized frameworks, such as the Innovative Methodologies and New Data for Predictive Oncology Model Evaluation (IMPROVE) project, provide unified data schemas and evaluation protocols for consistent benchmarking, but improving model generalizability requires larger and more diverse training data. In this work, we substantially expand the IMPROVE benchmark through large-scale integration of pharmacogenomic data, primarily from PharmacoDB, together with additional smaller data sources. The expanded resource includes millions of drug response measurements, broader multi-omics coverage, and a major increase in chemical diversity, adding more than 50,000 compounds. To evaluate the impact of the new dataset compared to the original IMPROVE benchmark dataset, we trained DRP models using the two datasets and assess their prediction performance using a common test set and several evaluation strategies, including drug-blind, cancer-blind, and disjoint data splits. While cancer-blind performance remained comparable to the original benchmark, models trained on the expanded dataset showed consistent improvements in drug-blind and disjoint settings, indicating enhanced generalization to previously unseen compounds. These results position the expanded dataset as a community resource that provides a richer foundation for developing DRP models intended to aid in the discovery of novel anticancer drugs.

View source

Similar papers

Open access Jun 2026

Unified heterogeneity-aware benchmark of drug synergy prediction: a cross-study analysis of traditional machine learning and graph deep learning models.

Drug synergy prediction holds great promise in accelerating combination therapy development and improving treatment efficacy in cancer and other complex diseases. However, progress in this area is hindered by considerable heterogeneity across experimental datasets, including variability in the number of drug combinations, inconsistencies in synergy scoring methodologies, and differences in data quality. Here, we present the first comprehensive benchmarking framework specifically designed to accommodate inter-dataset heterogeneity. This framework integrates 13 independent datasets encompassing 454,794 retained drug combination-cell line entries, 4247 drugs, and 187 cell lines, all of which are cancer cell lines. Our comparative evaluation of seven computational models for drug synergy prediction reveals that model performance strongly depends on both the scale and quality of the datasets. The graph model JointSyn performed favorably on datasets with larger numbers (> 10,000) of retained drug combinations, while the traditional random forest model performed competitively on smaller-scale datasets. We also observe that the ZIP scoring metric yields the highest accuracy in large-scale data, whereas HSA is more effective in sparse-data scenarios. However, different synergy metrics show significant variability in performance across datasets, suggesting that different synergy metrics capture distinct aspects of drug interactions, and the choice of metric can substantially affect model evaluation and cross‑dataset consistency. Furthermore, we find that well-designed small datasets can match or even surpass the performance of larger benchmarks, suggesting that different metrics are applicable to different datasets/testing scenarios. Our benchmark provides a robust foundation for fair model evaluation and paves the way for the development of more generalizable and preclinically relevant drug synergy prediction methods.

Yingjuan Cheng, Qing Ye, Linlong Jiang et al. · 0 citations
Open access Jul 2026

Essentiality-driven prediction of anticancer drug responses in preclinical and clinical contexts

Summary Precision oncology relies on tumor molecular profiles to predict drug responses. Instead of using conventional molecular features directly, we construct predictive signatures based on gene essentiality. Here, we present DrGee, an essentiality-centered platform that infers drug sensitivity solely from gene expression profiles. The built-in DeepEEAA model integrates gene expression, gene essentiality, drug-protein affinity, and drug-gene associations to quantitatively predict IC50 values. DeepEEAA achieved competitive predictive performance on independent cell line datasets (R2 = 0.764; MSE = 0.9345), outperforming recent benchmark deep learning methods. DrGee prioritized four candidate drugs for the 95-D lung cancer cell line, among which BI-97C1 and trimetrexate were validated by in vitro assays and mouse xenograft experiments. Robust predictive performance was further confirmed in OVCAR8 ovarian cancer cells. In TCGA cohorts, essentiality-driven predictions stratified patients with significantly different overall survival outcomes (AUC-PR = 0.825), highlighting the translational potential of DrGee.

Hongtu Cui, Xiaohui Du, Hai-Xia Guo et al. · 0 citations
Conference Jul 2026

Explainable Multi-Omic Machine Learning Framework for Predicting Drug Response in Breast Cancer

Accurate prediction of drug sensitivity in cancer cell lines is vital for precision oncology and patient-specific therapies. However, many computational approaches fail to integrate multi-modal biological and chemical features and often struggle with high-dimensional, imbalanced pharmacogenomic data, limiting predictive accuracy and interpretability. To address these challenges, we developed a machine learning framework that integrates pharmacogenomic profiles-including mutation status, copy number alterations, and microsatellite instabil-ity-with molecular fingerprints and descriptors of 85 anticancer drugs, generated using PaDEL from SMILES strings. Data from 40 breast cancer cell lines in the Genomics of Drug Sensitivity in Cancer (GDSC) dataset were employed. A threestage feature selection strategy combining Boruta, mRMR, and XGBoost was applied to reduce drug feature dimensionality while retaining 130 cell line features. Multiple models were trained, and LightGBM, optimized with grid search, class weighting, and 3-fold cross-validation, demonstrated superior performance in handling severe class imbalance (233 sensitive vs. 3167 resistant samples). LightGBM achieved training AUROC $=0.9455$, AUPRC $\boldsymbol{=} \mathbf{0. 5 1 4 8}$, Accuracy $\boldsymbol{=} \mathbf{0. 8 4 1 5}$, F1-score = 0.4481, Recall = 0.9409, and MCC = 0.4732, underscoring its suitability for sparse biomedical datasets. Model interpretation with SHapley Additive exPlanations (SHAP) highlighted BRCA-related features, identifying cnaBRCA25 (not mutated) as a resistance marker and cnaBRCA47 (mutated) as a context-dependent biomarker, consistent with their roles in DNA repair pathways. Overall, this framework demonstrates the value of multi-modal integration and interpretable machine learning in pharmacogenomics. While results are promising, validation on larger and independent cohorts is essential to establish clinical relevance.

D. Kumari, Aiman, Sakshi Singh et al. · 0 citations
Review Open access Jul 2026

Advanced Artificial Intelligence and data science in bioinformatics-driven drug discovery for cancer: Pathways toward shorter and less toxic treatment

Cancer remains one of the leading causes of death worldwide, with the GLOBOCAN estimates placing the 2022 global burden at close to 20 million new cases and 9.7 million deaths (Bray et al., 2024), a burden projected by the American Cancer Society (2024) to rise to roughly 35 million annual cases by 2050. Conventional cytotoxic chemotherapy, though still central to treatment for many tumor types, is frequently associated with prolonged treatment courses, non-specific systemic toxicity, and reduced quality of life. This review synthesizes recent literature on the application of artificial intelligence (AI) and data science within bioinformatics-driven cancer drug discovery, examining how these tools are reshaping target identification, molecular design, biomarker discovery, and treatment personalization. The analysis shows that deep learning-based protein structure prediction (Jumper et al., 2021), generative molecular design (Gangwal & Lavecchia, 2024), multi-omics target identification (Bhinder et al., 2021; Wei et al., 2023), digital pathology and radiomics (Bera et al., 2022; Lu et al., 2024), and machine learning models for predicting chemotherapy toxicity (Huang et al., 2024; Moslemi et al., 2025) are collectively shortening discovery timelines, improving the precision of treatment selection, and reducing treatment-related adverse effects in reported studies. Case evidence is presented, including a generative-AI-designed molecule that reached Phase I clinical trials in under 30 months (Insilico Medicine, 2022) and the 2024 Nobel Prize in Chemistry awarded for the AlphaFold protein-structure-prediction system. While these advances present a credible pathway toward shorter, more targeted, and less toxic cancer treatment, and in specific molecular contexts may reduce reliance on conventional chemotherapy, the evidence does not yet support claims that AI will universally eliminate chemotherapy; rather, it points toward an increasingly personalized standard of oncologic care. The review concludes by discussing the ethical, regulatory, and data-governance barriers that must be addressed for these gains to be realized safely and equitably.

Yejide Eniola Dabiri · 0 citations
Review Aug 2026

Advancing cancer drug discovery through the integration of machine learning and high-throughput screening.

Cancer drug discovery is a complex process that requires identifying compounds that selectively target malignant cells. While high-throughput screening (HTS) is essential for testing large libraries, it generates vast datasets that are difficult to interpret. Recently, the integration of artificial intelligence (AI), particularly deep learning (DL), has significantly accelerated drug candidate selection. This review highlights the synergy between AI and HTS, emphasizing DL techniques such as convolutional neural networks for bioactivity prediction, recurrent neural networks for de novo design, and reinforcement learning for property optimization. These methods streamline preclinical research by enabling rapid multi-omics analysis and prediction of drug-target interactions. However, challenges regarding data quality, model interpretability, and ethics persist. Emerging paradigms like Explainable AI and federated learning aim to enhance transparency and collaboration while safeguarding privacy. Ultimately, overcoming these barriers through AI-HTS integration holds transformative potential to reduce development costs and improve clinical outcomes for cancer patients.

K. Herbetko, Katarzyna Herbetko, Magdalena Mikołajek et al. · 0 citations
Open access Jul 2026

ICBcDrug: An online resource and tool for screening and predicting immunotherapy combination drugs.

BACKGROUND Combining small molecules with immune checkpoint blockade (ICB) therapy is an effective strategy for improving therapeutic efficacy. However, the systematic identification of small molecules that effectively potentiate ICB remains a significant challenge. METHODS We curated 276 literature-supported compounds known to enhance the efficacy of ICB. For each compound, we calculated its Core and Minor gene set score (CM-score), a quantitative metric defined by the CM-Drug method. The median CM-score of these compounds was then established as the reference standard. We subsequently used the CM-Drug framework to screen 2036 candidate drugs from the Library of Integrated Network-based Cellular Signatures (LINCS) database against this reference to identify potential ICB enhancers. RESULTS We developed ICBcDrug, a resource that integrates 2311 reported or predicted compounds across 18 cancer types. With respect to the reported drugs, ICBcDrug displays compounds with similar chemical structures or transcriptomic profiles, and for the predicted drugs, they are prioritized for further validation across various cancer types. A "Search" module allows users to retrieve compounds using customizable filters. Furthermore, ICBcDrug offers an "Analysis" module to predict the ICB combination efficacy of novel compounds on the basis of user-provided expression data. Using these modules, we identified promising candidate drugs and accurately predicted both the efficacy and potential mechanisms of the known ICB enhancer entinostat. CONCLUSIONS ICBcDrug (https://guolab.wchscu.cn/ICBcDrug) is a freely accessible and valuable resource for advancing ICB combination therapy. These findings will facilitate the discovery and development of novel ICB-enhancing drugs, contributing to the improvement of cancer immunotherapy.

Yu Lin, Wen Sun, Yun Xia et al. · 0 citations

Related blog posts