Aug 2026· IEEE journal of biomedical and health informatics· Vol PP· 0 citations
Medicine
TL;DR
By holistically integrating atomic, motif, and global fingerprint information via hypergraph modeling, HyperMolFusion offers a more reliable computational tool to enhance the efficiency and accuracy of drug development pipelines.
Abstract
Molecular property prediction is a critical task in accelerating drug discovery. While deep learning has shown promise, prevailing single-modal methods struggle to integrate multi-source (e.g., atomic graph and molecular fingerprints), heterogeneous chemical knowledge, thereby failing to holistically represent molecular structures and capture the high-order synergistic interactions governing their functions. To address these challenges, we present HyperMolFusion, a hypergraph-enhanced multi-modal fusion model for molecular property prediction. Compared with traditional graphs limited to pairwise atomic bonds, HyperMolFusion models chemical motifs as hyperedges to explicitly capture high-order structural correlations and encode complex molecular interactions. The framework comprises three core representation learning modules: AtomConv for local atomic interaction learning via attention-enhanced message passing, HyperConv for motif-level high-order correlation extraction via hypergraph convolution with GRU gating, and a mixed molecular fingerprint module that adaptively integrates MACCS, PubChem, and Pharmacophore fingerprints. A chemically guided attention (CGA) mechanism then dynamically fuses these multi-level features into hierarchical molecular representations, alleviating over-smoothing and preserving structural information effectively. Evaluated on eight MoleculeNet benchmarks (covering regression and classification tasks), HyperMolFusion achieves promising performance. For regression, it achieves an RMSE of 0.611 in lipophilicity, 0.653 in ESOL, and 0.951 in FreeSolv. For classification, it achieves a ROC-AUC of 0.935 in ClinTox, 0.907 in BBBP, and 0.689 in SIDER. This work provides a systematic and effective solution for molecular property prediction: by holistically integrating atomic, motif, and global fingerprint information via hypergraph modeling, HyperMolFusion offers a more reliable computational tool to enhance the efficiency and accuracy of drug development pipelines.
Molecular property prediction provides an important computational basis for compound screening and drug development by estimating physicochemical characteristics and biological activities from molecular structures. Although deep learning has improved molecular modeling, existing methods often describe molecules through a limited structural view or combine multiple views without sufficiently exploiting their complementary relationships. In addition, graph-based approaches commonly concentrate on local atomic connectivity, making it difficult to represent chemically meaningful structural units that may strongly influence molecular properties. This paper develops MG-CMIF, a multi-granularity cross-modal framework for molecular property prediction. The proposed model describes each molecule from symbolic, topological, and spatial perspectives and learns an integrated representation through three key designs. First, hierarchical graph modeling combines detailed atomic interactions with substructure-level chemical patterns to enrich topology-oriented features. Second, interaction across molecular views enables information relevant to property prediction to be exchanged selectively rather than merged through shallow operations. Third, alignment-oriented training objectives encourage representations derived from the same molecule to preserve compatible chemical semantics during fusion. Experiments on multiple public benchmark datasets show that MG-CMIF achieves better prediction results than competitive methods in both classification and regression settings. Further ablation analyses confirm that hierarchical structural modeling and cross-view integration both contribute to the effectiveness of the proposed framework.
Unknown authors· Journal of Machine Learning...· 0 citations
Evaluations on eight MoleculeNet datasets show that MSMPP significantly outperforms state-of-the-art models, demonstrating its effectiveness in integrating multi-view intra-molecular features, inter-molecular features and cross-task information.
Jiongfeng Chen, Yulian Ding, Yan Yan et al.· IEEE journal of biomedical a...· 0 citations
A novel Dual-Attention Multimodal framework for Graphs and Sequence-based representations, so-called DAM-GS, which provides a promising solution for molecular property prediction with broad applications in drug discovery and computational molecular science.
Bay Van Nguyen, Vinh Truong, Ha Duong Thi Hong et al.· Journal of Chemical Informat...· 0 citations
A novel multimodal alignment framework for joint modeling of molecular graphs and sequences, called Mol-ME, which employs ensemble learning to predict on extracted representations, which captures complex nonlinear relationships and compensates for the modeling limitations of single shallow networks.
Bao-Ren Huang, Mu Chen, Jun-Jie Luo et al.· Journal of Chemical Informat...· 0 citations
Active learning provides an efficient strategy for molecular property prediction by iteratively prioritizing compounds for experimental evaluation. However, the effectiveness of active learning pipelines depends strongly on the choice of molecular representation, and systematic understanding of how representation families affect the active learning process in terms of uncertainty and predictive performance remains limited. In this work, we introduce ActiveFusion, a framework for integrating heterogeneous molecular representations within active learning workflows for molecular property prediction. The framework enables systematic evaluation of physicochemical descriptors, molecular fingerprints, learned graph neural network (GNN) representations, and pretrained Transformer-based representations, as well as feature-level fusion strategies that combine complementary chemical information sources. ActiveFusion evaluates models across four molecular property prediction regression tasks. Across datasets, we demonstrate that feature fusion between learned graph representations and physicochemical descriptors consistently improves prediction performance and discovery (average final iteration R2 of 0.71, 0.67, and 0.54 for our overall best representations Chemprop+RDKit, Chemprop, and RDKit, respectively). We show that exploration-driven acquisition strategies enhance scaffold coverage and promote sampling of structurally novel regions of chemical space, and that model-agnostic acquisition of new compounds based on diversity has strong performance. Notably, both pretrained and finetuned Transformer-based embeddings do not consistently outperform physicochemical features or GNN-learned representations in our setting, highlighting the continued relevance of chemically interpretable features and learned features from supervised, task-specific models for active learning applications in molecular property prediction. Overall, ActiveFusion provides a systematic framework for studying representation-acquisition interactions in molecular discovery with representation fusion capabilities. Our study offers practical guidance for designing active learning pipelines that balance prediction accuracy with chemical space exploration.
Nelson Evbarunegbe, Shiyun Wa, Luke Taylor et al.· Journal of Chemical Informat...· 0 citations
This review provides a systematic overview of recent advances in SSL-based molecular property prediction and analyzes how multimodal molecular representation learning by integrating sequence, graph, three-dimensional structure, and textual information can improve the quality and expressiveness of molecular representations.
Shuning Yang, Lei Deng· Journal of Chemical Informat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.