Skip to content
Open access

3Br-MGD: few-shot toxicity prediction with a three-branch deep encoder and meta-learning framework.

Jul 2026 · Scientific Reports · 0 citations
Medicine

TL;DR

3Br-MGD, a novel three-branch framework that integrates deep learning and meta-learning for molecular toxicity prediction, demonstrates that 3Br-MGD consistently outperforms conventional baselines in predictive accuracy, robustness, and generalization.

Abstract

Predicting the toxicity of pharmaceutical compounds remains a major challenge in drug discovery. Early and accurate toxicity assessment is essential for eliminating harmful candidates before costly preclinical and clinical testing, thereby improving patient safety, reducing development costs, and accelerating the drug development process. Despite advances in computational toxicology, existing methods often struggle to capture complex molecular characteristics and maintain robust performance under limited-data conditions. To address these challenges, we propose 3Br-MGD, a novel three-branch framework that integrates deep learning and meta-learning for molecular toxicity prediction. The architecture combines complementary molecular representations: FingerprintMLP encodes Morgan fingerprint descriptors, Graph Convolutional Networks (GCNs) capture structural information from molecular graphs, and one-dimensional Deep Convolutional Neural Networks (1D-CNNs) extract sequential features from SMILES strings. These embeddings are integrated within a Prototypical Network-based few-shot learning framework, enabling rapid adaptation to new prediction tasks with limited labeled samples and improving generalization in low-resource settings. Experimental results on benchmark toxicity datasets demonstrate that 3Br-MGD consistently outperforms conventional baselines in predictive accuracy, robustness, and generalization. Furthermore, the integration of heterogeneous molecular encoders reduces dependence on large training datasets while enhancing interpretability through the exploitation of complementary chemical information from multiple molecular views.

Read PDF

Similar papers

Jul 2026

MSMPP: Molecular Property Prediction by Integrating Multi-scale Multi-view information with pretrained 3D molecular large model representation.

Evaluations on eight MoleculeNet datasets show that MSMPP significantly outperforms state-of-the-art models, demonstrating its effectiveness in integrating multi-view intra-molecular features, inter-molecular features and cross-task information.

Jiongfeng Chen, Yulian Ding, Yan Yan et al. · 0 citations
Open access Aug 2026

LMF-CP: An Interpretable Multimodal Late-Fusion Framework for Compound Carcinogenicity Prediction

Accurately predicting the carcinogenicity of compounds is of great significance for drug discovery, clinical drug safety, and chemical risk assessment. Traditional methods for assessing carcinogenicity rely on animal testing, which suffers from limitations such as time-consuming processes, high costs, significant interspecies differences, and low predictive throughput. In recent years, computational modeling-based prediction methods (such as Quantitative Structure–Activity Relationships, QSAR) have made some progress, but they still face challenges such as insufficient molecular feature information and poor model interpretability. To overcome these barriers, the multimodal deep learning framework LMF-CP (Late Multimodal Fusion of Carcinogenicity Prediction) is proposed to enhance the performance and interpretability of compound carcinogenicity prediction. First, to comprehensively characterize the structural and physicochemical properties of compounds, a multimodal representation system based on four molecular modalities is constructed, namely SMILES sequences, molecular fingerprints, molecular images, and molecular graph structures. Specifically, Text Convolutional Neural Network (TextCNN), Multi-Layer Perceptron (MLP), Visual Geometry Group Network (VGGNet), as well as Molecular Graph Attention Network (MGAT) are employed to process this information, respectively. Second, to integrate information from different molecular representations, a late-stage fusion strategy based on Lasso stacking is employed. On the test set, LMF-CP achieves an area under curve (AUC) of 0.828, an accuracy (ACC) of 0.782, an F1 score of 0.786, a sensitivity (SEN) of 0.786, and a specificity (SPE) of 0.779. In addition, this paper combines Shapley Additive Explanations (SHAP) analysis with Bemis–Murcko scaffold analysis to interpret the model results from two perspectives. Finally, a visual online platform for predicting the carcinogenicity of compounds is designed, providing a convenient tool for the rapid assessment of compound carcinogenicity and structural interpretation.

Yingjie Zhu, Liu-Jie He, Xin-Jie Liang · 0 citations
Open access Jul 2026

Development and validation of an attention-based cGAS-specific deep learning scoring function for structure-based virtual screening

DeepCGASPred is a cyclic GMP-AMP synthase (cGAS)-specific deep learning scoring function that integrates three-dimensional convolutional neural networks with multi-head attention mechanisms and composite structural descriptors, including Structural Protein–Ligand Interaction Fingerprints (SPLIF), hydrogen bond features, and extended connectivity fingerprints (ECFP).

Muhammad Junaid, Muhammad Zeeshan, Abbas Khan et al. · 0 citations
Open access Jul 2026

Mol2Image: an enhanced DDI prediction framework leveraging drug molecular descriptors

Drug–drug interactions (DDIs) are a critical safety issue in clinical practice, as they can lead to severe and often unpredictable adverse effects. This risk becomes significantly higher in multi-drug therapies, which are increasingly used in the treatment of complex and chronic diseases such as cancer, cardiovascular disorders, and diabetes. However, identifying DDIs through in vivo studies is costly and time-consuming. In this study, a novel DDI prediction model, Mol2Image, has been proposed that utilizes chemical structure features derived from Simplified Molecular Input Line Entry System (SMILES) representations, including molecular property descriptors and structural fingerprints. The proposed model combines chemical structure information with automated feature learning. Molecular descriptors and structural fingerprints extracted from SMILES representations are converted into visual patterns that capture key chemical characteristics of each drug. These images are then processed by a Convolutional Neural Network (CNN) to learn high-level structural features associated with drug–drug interactions. The model is trained and evaluated using two benchmark DDI datasets: the Drugbank dataset, which consists of 443,046 interactions, and ChCh-Miner, which consists of 48,514 DDIs. Experimental results demonstrate that the proposed model (Mol2Image) achieves competitive performance compared with several state-of-the-art methods. Experimental results demonstrate that the proposed model consistently outperforms existing approaches, achieving accuracies of 0.9608 and 0.9683 using the Drugbank dataset and ChCh-Miner dataset, respectively. Ultimately, Mol2Image provides a highly scalable, strictly structure-centric framework that ensures superior predictive accuracy with minimal computational overhead, operating entirely independently of clinical data.

Nourhan Helmy, H. A. Maghawry, N. Badr · 0 citations
Open access Sep 2026

HFEDTI: A DTI Prediction Model Integrating Local–Global Feature Fusion and Weighted Ensemble Learning

Drug–target interaction (DTI) prediction is a critical step in drug discovery, and accurate prediction of potential interactions can significantly accelerate the drug-development process. Although deep-learning approaches have achieved promising performance in DTI prediction, two challenges remain: single models often fail to comprehensively capture heterogeneous sequence information, resulting in limited stability and generalization, while insufficient integration of local and global features restricts interaction representation. To address these limitations, we propose HFEDTI, a DTI prediction model that integrates hierarchical feature fusion and weighted ensemble learning. Specifically, a residual convolutional neural network (ResCNN) is employed to extract local structural features of drugs and targets, while a self-attention-based hierarchical bidirectional long short-term memory network (SAHBiLSTM) captures global contextual dependencies. Furthermore, a hierarchical heterogeneous attention mechanism is introduced to align and fuse multi-level cross-modal representations, and a weighted ensemble strategy based on validation performance ranking is developed to enhance model robustness and generalization. Experimental results on three benchmark datasets demonstrate the effectiveness of HFEDTI. On the DrugBank dataset, HFEDTI achieves an AUC of 0.9238 and an AUPR of 0.9327, improving the best-performing baseline by 0.90 and 1.40 percentage points, respectively. Moreover, HFEDTI consistently achieves strong performance on the C. elegans and Human datasets, further validating its effectiveness and generalization capability for DTI prediction.

Unknown authors · 0 citations
#machine learning Preprint Aug 2026

ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction

Molecular toxicity prediction is increasingly used to prioritise compounds before experimental testing, but conventional benchmark performance can overstate practical utility when structurally related molecules occur across training and test folds. We introduce ToxLens, a reproducible multi-task graph-learning framework for 11 toxicity endpoints spanning Ames mutagenicity, acute oral toxicity, hERG inhibition, and Tox21 nuclear-receptor and stress-response assays. The workflow combines conservative chemical curation, sphere-exclusion filtering, a leakage-aware UMAP-HDBSCAN split, parallel graph and global-feature encoders joined by late concatenation, temperature-scaled Monte Carlo dropout with conformal-style prediction sets, applicability-domain analysis, and SHAP-guided toxicophore discovery with occlusion controls. On the leakage-controlled test fold, a five-seed soft-voting ensemble achieved a Matthews correlation coefficient score of 0.44, an area under the receiver operating characteristic curve score of 0.83, and an area under the precision-recall curve score of 0.58. It exceeded four ECFP4-based shallow baselines on all 11 endpoints under the same split and validation-based threshold-selection protocol. Controlled ablations showed that the global pathway was important, whereas late concatenation outperformed the tested gated and feature-wise linear modulation fusion variants. Conformal-style prediction sets revealed substantial endpoint-specific variation in set efficiency, and discrimination and calibration improved with similarity to the training domain. Retraining on fixed published Tox21 Challenge and TDA folds produced competitive, but not uniformly state-of-the-art, performance. SHAP-guided occlusion and consensus subgraph mining yielded model-derived structural hypotheses, 44 of which contained at least one occurrence that passed the predefined counterfactual criteria.

Magnus H. Strømme, A. D. de Sá, David B. Ascher · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.