Jul 2026· International Conference on Climate Informatics· Vol 42· 0 citations· 48 references
TL;DR
The proposed MM2Vec model consistently outperforms classical machine learning (ML)‐based models and recent state‐of‐the‐art deep learning (DL)/GNN‐based methods in terms of accuracy, robustness, and generalization.
Abstract
For many years, graph representation learning plays a pivotal role in bioinformatics and cheminformatics; as a result, supporting a wide range of tasks such as drug discovery, toxicity prediction, and compound–protein interaction analysis. However, existing approaches often focus solely on either sequential molecular fingerprints or graph‐based structural features, which limit their ability to capture both local chemical substructures and global molecular topology. To address this issue, we propose MM2Vec, a novel multi‐viewed molecular representation learning framework that integrates local rich‐feature embedding with graph neural network (GNN)‐based structural learning. Specifically, each molecular graph is first processed through an MLP‐based embedding layer that encodes sub‐structural fingerprint information extracted from radius‐based subgraphs, capturing fine‐grained chemical and physiochemical features. Simultaneously, a multi‐layered GNN encoder learns topological relationships from the molecular graph structure; therefore, focusing more on geometric and relational information among atoms. The outputs from both embedding branches are then fused using a learnable linear mechanism to produce unified, high‐quality molecular embeddings in a shared latent space. These fused representations are used to drive task‐specific prediction layers for addressing various learning objectives. We validate the proposed MM2Vec model on multiple graph learning tasks, including drug‐induced liver injury (DILI) classification and lethal dose (LD) molecular regression problems. Experimental results show that MM2Vec consistently outperforms classical machine learning (ML)‐based models and recent state‐of‐the‐art deep learning (DL)/GNN‐based methods in terms of accuracy, robustness, and generalization. Our findings in this highlight the importance of combining both sub‐structural and graph‐structural perspectives and demonstrate the versatility and effectiveness of our MM2Vec model for a wide range of molecular analysis tasks.
Chemoresistance is a major contributor to cancer treatment failure, and microRNAs (miRNAs) play a critical role in mediating this resistance by regulating gene expression. Therefore, identifying miRNA-drug associations is of great significance for advancing cancer therapy. However, existing computational models face significant challenges, including heterogeneous feature integration and data sparsity. To overcome these limitations, we propose a novel Multi-view Graph Learning Framework with Spectral Encoding and Sparse Cross-Attention (MVGSCA) for predicting miRNA-drug associations. The model constructs node features based on miRNA sequence similarity and drug SMILES similarity. Then it builds two distinct graphs: a gene-mediated functional graph from miRNA-drug target interactions and an association-guided structural graph from known miRNA-drug associations. These two graphs are linearly combined to produce a collaborative feature representation. To capture both local and global topological features, the model applies local power filtering and global heat kernel diffusion, followed by spectral encoding via Poisson-Charlier polynomial approximation to enhance the feature representation. Furthermore, a sparse cross-attention mechanism is introduced to dynamically weight and integrate heterogeneous features from multiple sources. On a benchmark dataset with 8,720 associations, MVGSCA achieves an AUC of 96.32% and an AUPR of 95.69% under five-fold cross-validation, significantly outperforming six state-of-the-art methods. Experimental results show that MVGSCA effectively integrates heterogeneous biological information and achieves superior prediction performance, offering valuable insights into cancer resistance mechanisms and supporting drug discovery efforts.
Ru Nie, Ying Fu, Zhengwei Li et al.· IEEE transactions on computa...· 0 citations
Molecular property prediction is a fundamental task in drug discovery and chemical biology, where effective molecular representations are essential for accurate prediction. Learning transferable motif-level representations remains challenging because explicit motif annotations are scarce and existing representations may be altered during downstream supervised optimization. In this study, we propose MSGRL, a motif-driven self-supervised graph representation learning framework for interpretable molecular property prediction. MSGRL represents each molecule through a hierarchical graph structure, consisting of a motif-based graph for inter-motif organization and motif-specific atom-based graphs for intra-motif atomic structure. Its central design is to decouple label-agnostic intra-motif structural learning from label-dependent inter-motif property learning. An MPNN-GRU encoder is pretrained on motif-specific atom-based graphs using a variational motif graph autoencoder (VMGAE), which reconstructs the internal bond topology of motifs in a self-supervised manner. After pretraining, the intra-motif encoder is kept frozen, while the downstream module pools atom-level feature matrices into motif vectors, propagates information over the motif-based graph, and applies attention-based pooling for molecular property prediction. This design keeps the pretrained intra-motif representations fixed while allowing the inter-motif network and prediction head to adapt to individual downstream tasks. Experiments on eight MoleculeNet benchmark datasets show that MSGRL achieves the highest ROC-AUC scores on all five classification datasets and the lowest RMSE on Lipophilicity, while its performance on ESOL and FreeSolv is more mixed. Ablation studies further support the contributions of encoder freezing, motif-based graph construction, and attention-based pooling. Motif-level attribution analyses provide qualitative and dataset-level evidence regarding the substructures emphasized by the model. These results demonstrate the effectiveness of the proposed hierarchical representation strategy, particularly for the evaluated molecular classification tasks.
You Wu, Yu-Xin Jiang, Xiao-Yun Qi et al.· Molecules· 0 citations
A diffusion-enhanced inductive link prediction framework that combines Graph Diffusion Convolution (GDC), structural node descriptors, and neighborhood aggregation from GraphSAGE is proposed that achieves higher accuracy than the other models on the benchmark datasets.
An Algebraic Graph Neural Network model designed to encode molecular structures into a low-dimensional graph representation while preserving critical biochemical interactions is introduced, demonstrating superior performance in binding affinity prediction compared to state-of-the-art scoring functions.
Augustine Ouru, Xi Chen, Cameron Yeagle et al.· Computational and Mathematic...· 0 citations
Results indicate that integrating heterogeneous structural cues through coarse- and fine-grained feature interaction provides an effective and scalable solution for DDI prediction.
The DRL-DSP is proposed, a novel dual representation learning framework designed to enhance drug synergy prediction by integrating molecular-level features from SMILES sequences with graph-level relational information from reconstructed molecular networks.
Juanzi Zhou, Xiaoliang Yang, Yin Zhang et al.· Intelligent Data Analysis· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.