ClipMol is proposed, a scalable and structurally faithful solution for molecular representation learning and IM-MS-related CCS prediction in analytical chemistry that jointly models local chemical microenvironments and global structural constraints of molecules.
Abstract
Although existing molecular pretraining models have achieved favorable performance on various downstream tasks, their reliance on explicit three-dimensional conformer sampling or external natural-language corpora often increases computational cost and affects structural fidelity. Here, we propose ClipMol, a molecular representation learning framework based on SMILES-InChI dual-view chemical-language alignment. Without requiring explicit three-dimensional conformers or external corpora, ClipMol jointly models local chemical microenvironments and global structural constraints of molecules. Benchmark results show that ClipMol and its scaled variant, ClipMol-XL, achieve strong overall performance on both classification and regression tasks. More importantly, for collision cross-section (CCS) prediction in ion mobility-mass spectrometry, ClipMol shows stable and competitive performance on two independent benchmark data sets, METLIN-CCS and ALLCCS, while maintaining robustness across different adduct compositions and diverse chemical categories. Compared with state-of-the-art and competitive CCS prediction models, the ClipMol models achieved the best or highly competitive averaged performance on both data sets, with ClipMol-XL showing the strongest overall R2 and root-mean-square error performance. Overall, ClipMol provides a scalable and structurally faithful solution for molecular representation learning and IM-MS-related CCS prediction in analytical chemistry.
MolPACL is proposed, a prompt-augmentation-based supervised contrastive learning framework that incorporates high-level chemical semantics while preserving molecular identity, and achieves strong performance on both classification and regression tasks while reducing training cost.
Ali Forooghi, Luis Rueda, A. Ngom· IEEE transactions on computa...· 0 citations
FragBERTa is introduced, a molecular fragment-aware transformer-based representation language model pretrained using masked language modeling on Sequential Attachment-based Fragment Embedding (SAFE) representations, suggesting that fragment-based string representations offer advantages over atom-level representations for scaffold-sensitive and interaction-driven tasks.
The representation of molecular structure in textual form plays a central role in data-driven cheminformatics, not just for deep learning models that often rely on sequence-based inputs, but also for classic machine learning pipelines which are still relevant in this field. The Simplified Molecular Input Line Entry System (SMILES) representation and its explicit variants differ substantially, yet the implications of this difference for tokenization and downstream property prediction remain insufficiently characterized. This paper systematically investigates the role of explicitness in SMILES representation and its impact on aqueous solubility prediction. Four SMILES variations ranging from basic and canonical to explicit and explicit canonical are reviewed to illustrate how increasing explicit chemical detail alters the available information content. Building on this analysis, a novel tokenization approach termed ExACT (Explicit Atom level Context Tokenization) is introduced, which directly leverages atom level explicitness and local chemical context. In addition, a general strategy for enhancing existing tokenization methods through increased explicitness is proposed and demonstrated by upgrading the current state of the art atom-in-SMILES method to an explicit variant. All approaches are evaluated on the AqSolDB dataset using consistent machine learning pipelines and cross validation protocols. The results show that increased explicitness systematically improves predictive performance across tokenization strategies, yielding higher coefficients of determination, lower mean absolute and root mean squared errors, and reduced variance across folds. The proposed ExACT method achieves the best overall performance, largely outperforming traditional molecular fingerprint representations while remaining statistically comparable to the best competing tokenization methods. Furthermore, the efficiency analysis demonstrated that ExACT offers additional computational advantages, including lower dimensionality and the elimination of the decoding step, resulting in faster feature computation. These findings demonstrate that explicit SMILES representations encode chemically meaningful information that can be effectively exploited through context-aware tokenization, providing an interpretable and computationally efficient alternative representation for molecular property prediction.
Scientific contribution
This work investigates the role of explicitness in SMILES and its impact on aqueous solubility prediction, proposes a novel ExACT tokenization method and a general strategy for enhancing existing tokenization methods through increased explicitness.
D. Begušić, D. Pintar, Zlatko Smole et al.· Journal of Cheminformatics· 0 citations
Evaluations on eight MoleculeNet datasets show that MSMPP significantly outperforms state-of-the-art models, demonstrating its effectiveness in integrating multi-view intra-molecular features, inter-molecular features and cross-task information.
Jiongfeng Chen, Yulian Ding, Yan Yan et al.· IEEE journal of biomedical a...· 0 citations
Experiments show that PWAV generally improves over classical fingerprint descriptors within learned models and achieves competitive performance relative to established external baselines on several endpoints, positioning PWAV as a competitive and chemically transparent component for hybrid molecular property prediction, rather than as a replacement for domain-specific benchmark systems.
M. Afzal, S. Siddiqi· Physica Scripta· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.