Unimodal vs Multimodal
Learning: A Systematic Evaluation
of Fusion Strategies and Model Design for Molecular Property Prediction
and Uncertainty Quantification
Aug 2026· Journal of Chemical Information and Modeling· 0 citations· 61 references
TL;DR
The results demonstrate that successful multimodal learning depends on the coordinated selection of complementary molecular representations, fusion strategy, and learning algorithm rather than simply increasing the number of integrated modalities.
Abstract
Accurate prediction of molecular properties is fundamental to environmental chemistry, yet remains challenging when experimental data are limited. Multimodal fusion provides a promising strategy for integrating complementary molecular representations; however, the relative contributions of molecular representation, fusion strategy, and learning algorithm to the predictive accuracy and uncertainty remain poorly understood. Five molecular modalities (RDKit descriptors, Mol2Vec embeddings, graph neural network embeddings, SMILES representations, and MS2 fragmentation spectra) were evaluated by using early and late fusion strategies with four learning algorithms (LightGBM, RF, AttentiveFP, and DMPNN). Across 14 physicochemical properties, multimodal models exhibited modest numerical improvements over the best unimodal models, although these differences were generally not statistically significant. In contrast, uncertainty quantification revealed clearer distinctions among the modeling strategies. Multimodal integration significantly improved the alignment between prediction error and estimated uncertainty. The fusion strategy had a modest influence on epistemic uncertainty, with significant early versus late differences observed only for selected modality combinations, whereas meta-learner selection had the greatest effect on uncertainty calibration. Ablation, grouped SHAP, and RDKit descriptor reduction analyses showed that RDKit descriptors remained consistently informative despite substantial descriptor reduction, while Mol2Vec, SMILES, GNN embeddings, and MS2 contributed in a property-dependent and partially redundant manner. Computational cost increased substantially with multimodal complexity, whereas predictive accuracy exhibited diminishing returns, indicating that intermediate multimodal configurations often provided the most favorable balance among computational efficiency, predictive performance, and uncertainty reliability. Overall, the results demonstrate that successful multimodal learning depends on the coordinated selection of complementary molecular representations, fusion strategy, and learning algorithm rather than simply increasing the number of integrated modalities. Multimodal integration may provide particular value by improving the reliability of uncertainty estimation.
This review provides a systematic overview of recent advances in SSL-based molecular property prediction and analyzes how multimodal molecular representation learning by integrating sequence, graph, three-dimensional structure, and textual information can improve the quality and expressiveness of molecular representations.
Shuning Yang, Lei Deng· Journal of Chemical Informat...· 0 citations
Trimole-Hybrid is presented, a task-wise multimodal framework that addresses ADMET heterogeneity by selecting or combining predictors built from complementary molecular representations, and shows sensitivity to changes in essential functional motifs, suggesting its ability to capture ADMET-relevant molecular substructures.
This work develops the Bayesian Class-Attentive Transformer Network, a unified Bayesian framework that integrates classification-to-regression knowledge fusion, uncertainty quantification, and active learning for data-efficient molecular property prediction and establishes a generalizable paradigm for uncertainty-aware molecular modeling.
Shiyang Bian, Yukun Luo, Hongqiao Wang et al.· Journal of Chemical Informat...· 0 citations
A two-stage contrastive learning framework integrating drug structures, protein sequences, and Cell Painting morphological profiles into a unified embedding space, which reveals pathway-specific morphological signatures associated with drug targets, providing biologically interpretable insights into drug mechanisms.
Accurately predicting the carcinogenicity of compounds is of great significance for drug discovery, clinical drug safety, and chemical risk assessment. Traditional methods for assessing carcinogenicity rely on animal testing, which suffers from limitations such as time-consuming processes, high costs, significant interspecies differences, and low predictive throughput. In recent years, computational modeling-based prediction methods (such as Quantitative Structure–Activity Relationships, QSAR) have made some progress, but they still face challenges such as insufficient molecular feature information and poor model interpretability. To overcome these barriers, the multimodal deep learning framework LMF-CP (Late Multimodal Fusion of Carcinogenicity Prediction) is proposed to enhance the performance and interpretability of compound carcinogenicity prediction. First, to comprehensively characterize the structural and physicochemical properties of compounds, a multimodal representation system based on four molecular modalities is constructed, namely SMILES sequences, molecular fingerprints, molecular images, and molecular graph structures. Specifically, Text Convolutional Neural Network (TextCNN), Multi-Layer Perceptron (MLP), Visual Geometry Group Network (VGGNet), as well as Molecular Graph Attention Network (MGAT) are employed to process this information, respectively. Second, to integrate information from different molecular representations, a late-stage fusion strategy based on Lasso stacking is employed. On the test set, LMF-CP achieves an area under curve (AUC) of 0.828, an accuracy (ACC) of 0.782, an F1 score of 0.786, a sensitivity (SEN) of 0.786, and a specificity (SPE) of 0.779. In addition, this paper combines Shapley Additive Explanations (SHAP) analysis with Bemis–Murcko scaffold analysis to interpret the model results from two perspectives. Finally, a visual online platform for predicting the carcinogenicity of compounds is designed, providing a convenient tool for the rapid assessment of compound carcinogenicity and structural interpretation.
Yingjie Zhu, Liu-Jie He, Xin-Jie Liang· International Journal of Mol...· 0 citations
EQTri-DTI is designed to integrate three modality-specific networks to encode 1D protein sequences, 2D molecular images, and 3D drug structures and develops a joint uncertainty quantification scheme by calculating the weight summation of evidential uncertainty and prediction entropy from the aforementioned output, enabling a more comprehensive and nuanced assessment of uncertainties.