Skip to content
Open access

Smiles-based bioactivity prediction through molecular encoder selection and data augmentation.

Jul 2026 · Journal of Cheminformatics · 0 citations
Medicine

TL;DR

The results show that the use of suitable encoder-regressor pairs together with embedding-level mix-up augmentation improves model generalizability without requiring SMILES-level augmentation, and could be applied more broadly to IC50 prediction for other kinase inhibitors.

Abstract

Quantitative prediction of inhibitor potency can accelerate early-stage drug discovery. Recently, data-driven approaches have gained widespread interest in drug discovery, as evidenced by a growing number of benchmarking challenges and open competitions. In this context, we developed a machine learning-based methodology that can find the most effective way of predicting IC50 values against ASK1 from SMILES, for "Jump AI(.py) 2025: 3rd AI Drug Discovery Competition", hosted by the Korea Pharmaceutical and Bio-Pharma Manufacturers Association (KPBMA) on the Dacon platform. Applying our methodology achieved the highest overall predictive performance among all participating teams. Beyond this competition setting, we present a compact SMILES-based modeling workflow comprising (i) a pre-trained encoder, (ii) regression models, (iii) data augmentation, and (iv) hyperparameter tuning. We systematically compared molecular representations from sequence- and graph-based models, including ChemBERTa-2 and MolCLR. Across encoder-regressor combinations, ChemBERTa-77 M-MLM embeddings paired with support vector regression (SVR) yielded the strongest predictive performance. Embedding-level mix-up augmentation and SVR hyperparameter tuning further improved predictive performance. Our findings highlight that careful SMILES preprocessing and encoder selection have a critical influence on IC50 values and provide a reproducible benchmark for single-target bioactivity prediction, thus contributing to a more efficient drug discovery process. Scientific Contribution In this study, we propose a machine learning methodology for predicting the IC50 values of ASK1 inhibitors from SMILES representations, with a systematic comparison of molecular encoders and regression models. Our results show that the use of suitable encoder-regressor pairs together with embedding-level mix-up augmentation improves model generalizability without requiring SMILES-level augmentation. This strategy would be particularly useful for settings with imbalanced labels or limited data, and could be applied more broadly to IC50 prediction for other kinase inhibitors.

Read PDF

Similar papers

Review Aug 2026

Harnessing the Power of AI: A Modern Review on the Prediction of ADMET Properties in Drug Discovery.

Drug discovery is frequently limited by high attrition rates, and poor absorption, distribution, metabolism, excretion, and toxicity (ADMET) profiles are a major cause of late-stage failure. Therefore, precise ADMET property prediction is necessary to develop safe and effective drug candidates. Traditional experimental assays and rule-based computational procedures are limited by their poor predictive power, cost, and time, despite providing valuable insights. Innovative strategies to deal with these issues have been introduced by developments in artificial intelligence (AI), such as machine learning (ML), deep learning (DL), graph neural networks (GNNs), generative models, and multi-task learning (MTL). AI techniques can better generalize scaffolds, capture interdependencies between pharmacokinetic and toxicological endpoints, and model complex nonlinear relationships by leveraging large, diverse datasets. Explainable AI (XAI) enhances transparency by detecting biological and structural characteristics that are relevant to predictions, even if integrated pipelines combine predictive modeling with molecular creation and optimization. AI-driven ADMET prediction is becoming a vital tool in lowering attrition, speeding up candidate prioritization, and influencing the direction of rational drug development, despite persistent issues with data quality, regulatory acceptance, and synthetic viability.

Satyam Kumar Vishwash, Ram Babu Soni, Ratima Sood et al. · 0 citations
Conference Jul 2026

A Drug-Target Affinity Prediction Model Based on Bayesian Meta-Learning and Uncertainty Fusion

In the early stages of drug discovery, predicting drug-target affinity is a crucial task. Due to the vast scale of genomic and chemical spaces, traditional biological methods are time-consuming, labor-intensive, and resource-demanding. As a result, machine learning-based computational methods have emerged to narrow down the pool of drug candidates. However, machine learning approaches still face several challenges in practical applications, particularly the scarcity of labeled samples and poor model generalization capability. To address these issues, this paper proposes a novel drug-target affinity prediction model, termed MetaBayes-DTA, based on an uncertainty-aware meta-learning framework. The model integrates the few-shot rapid adaptation capability of meta-learning with an uncertainty quantification mechanism to enhance prediction accuracy and reliability. MetaBayes-DTA is evaluated on two benchmark datasets, DAVIS and KIBA. Experimental results demonstrate that the proposed model outperforms existing methods.

Naihan Shi, Yanpeng Zhao, Wanying Li et al. · 0 citations
Review Aug 2026

Gaps in AI-Driven Pharmacokinetic Property Prediction for Early Drug Development: A Scoping Review.

This scoping review investigates the current state of PK property prediction of small molecules in drug discovery using machine learning methods and a combination of machine learning and mechanistic models and proposes leveraging the pattern recognition capabilities of deep learning models in conjunction with the biological interpretability provided by mechanistic approaches.

Lucille Tomin, Vida Bodaghi-Namileh, D. Schwartz et al. · 0 citations
Open access Aug 2026

Machine learning-based prediction of SARS-CoV-2 bioactivity: integrating IC50 regression and activity classification using multi-task neural networks

Accurate prediction of compound bioactivity is essential for accelerating antiviral drug discovery and reducing experimental costs. Machine learning (ML) methods have shown considerable promise in modeling structure–activity relationships and compound potency. In this study, we present an integrated ML framework for predicting IC50 and pIC50 values of compounds active against SARS-CoV-2, key indicators of antiviral potency. The proposed framework comprises three complementary approaches: (i) a regression model for quantitative IC50 prediction validated against experimental data; (ii) a classification model that categorizes compounds into active and inactive classes to support compound prioritization; and (iii) a multi-task neural network that jointly performs IC50 regression and activity classification, enhancing predictive performance and interpretability. A distinctive feature of this work is the incorporation of ligand efficiency (LE) as a criterion for activity classification, offering an alternative perspective on compound prioritization that has not been previously explored in SARS-CoV-2 bioactivity modeling. The proposed models demonstrate strong predictive capability, achieving a coefficient of determination (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$R^2$$\end{document}) of 0.77 using a neural network with feature selection, while the Random Forest classifier attains an accuracy, precision, and recall of approximately 0.92. These results highlight the potential of integrated regression, classification, and multi-task learning approaches as scalable and cost-effective tools for SARS-CoV-2 bioactivity prediction and antiviral drug discovery.

Aya I. Maiyza, Sohila Osama, Hanan A Hassan · 0 citations
Open access Aug 2026

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Jurica Novak · 0 citations
Aug 2026

Dual-Attention Multimodal Framework for Molecular Property Prediction.

A novel Dual-Attention Multimodal framework for Graphs and Sequence-based representations, so-called DAM-GS, which provides a promising solution for molecular property prediction with broad applications in drug discovery and computational molecular science.

Bay Van Nguyen, Vinh Truong, Ha Duong Thi Hong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.