Skip to content

ClipMol: A Molecular Representation Learning Framework for CCS Prediction via SMILES-InChI Dual-View Chemical Language Alignment.

Jul 2026 · Analytical Chemistry · Vol 98, pp. 22451-22461 · 0 citations · 34 references
Medicine

TL;DR

ClipMol is proposed, a scalable and structurally faithful solution for molecular representation learning and IM-MS-related CCS prediction in analytical chemistry that jointly models local chemical microenvironments and global structural constraints of molecules.

Abstract

Although existing molecular pretraining models have achieved favorable performance on various downstream tasks, their reliance on explicit three-dimensional conformer sampling or external natural-language corpora often increases computational cost and affects structural fidelity. Here, we propose ClipMol, a molecular representation learning framework based on SMILES-InChI dual-view chemical-language alignment. Without requiring explicit three-dimensional conformers or external corpora, ClipMol jointly models local chemical microenvironments and global structural constraints of molecules. Benchmark results show that ClipMol and its scaled variant, ClipMol-XL, achieve strong overall performance on both classification and regression tasks. More importantly, for collision cross-section (CCS) prediction in ion mobility-mass spectrometry, ClipMol shows stable and competitive performance on two independent benchmark data sets, METLIN-CCS and ALLCCS, while maintaining robustness across different adduct compositions and diverse chemical categories. Compared with state-of-the-art and competitive CCS prediction models, the ClipMol models achieved the best or highly competitive averaged performance on both data sets, with ClipMol-XL showing the strongest overall R2 and root-mean-square error performance. Overall, ClipMol provides a scalable and structurally faithful solution for molecular representation learning and IM-MS-related CCS prediction in analytical chemistry.

View source

Similar papers

Open access Aug 2026

MolPACL: Molecular Property Prediction Based on Prompt Augmentation and Contrastive Learning.

MolPACL is proposed, a prompt-augmentation-based supervised contrastive learning framework that incorporates high-level chemical semantics while preserving molecular identity, and achieves strong performance on both classification and regression tasks while reducing training cost.

Ali Forooghi, Luis Rueda, A. Ngom · 0 citations
Open access Jul 2026

FragBERTa: a fragment-aware molecular representation model with sequential attachment-based fragment embeddings

FragBERTa is introduced, a molecular fragment-aware transformer-based representation language model pretrained using masked language modeling on Sequential Attachment-based Fragment Embedding (SAFE) representations, suggesting that fragment-based string representations offer advantages over atom-level representations for scaffold-sensitive and interaction-driven tasks.

Neerav Kaushal, Ajay Mnv Penmatsa · 0 citations
Review Open access Aug 2026

Explicitness in SMILES representation via ExACT: improved tokenization for aqueous solubility prediction

The representation of molecular structure in textual form plays a central role in data-driven cheminformatics, not just for deep learning models that often rely on sequence-based inputs, but also for classic machine learning pipelines which are still relevant in this field. The Simplified Molecular Input Line Entry System (SMILES) representation and its explicit variants differ substantially, yet the implications of this difference for tokenization and downstream property prediction remain insufficiently characterized. This paper systematically investigates the role of explicitness in SMILES representation and its impact on aqueous solubility prediction. Four SMILES variations ranging from basic and canonical to explicit and explicit canonical are reviewed to illustrate how increasing explicit chemical detail alters the available information content. Building on this analysis, a novel tokenization approach termed ExACT (Explicit Atom level Context Tokenization) is introduced, which directly leverages atom level explicitness and local chemical context. In addition, a general strategy for enhancing existing tokenization methods through increased explicitness is proposed and demonstrated by upgrading the current state of the art atom-in-SMILES method to an explicit variant. All approaches are evaluated on the AqSolDB dataset using consistent machine learning pipelines and cross validation protocols. The results show that increased explicitness systematically improves predictive performance across tokenization strategies, yielding higher coefficients of determination, lower mean absolute and root mean squared errors, and reduced variance across folds. The proposed ExACT method achieves the best overall performance, largely outperforming traditional molecular fingerprint representations while remaining statistically comparable to the best competing tokenization methods. Furthermore, the efficiency analysis demonstrated that ExACT offers additional computational advantages, including lower dimensionality and the elimination of the decoding step, resulting in faster feature computation. These findings demonstrate that explicit SMILES representations encode chemically meaningful information that can be effectively exploited through context-aware tokenization, providing an interpretable and computationally efficient alternative representation for molecular property prediction. Scientific contribution This work investigates the role of explicitness in SMILES and its impact on aqueous solubility prediction, proposes a novel ExACT tokenization method and a general strategy for enhancing existing tokenization methods through increased explicitness.

D. Begušić, D. Pintar, Zlatko Smole et al. · 0 citations
Jul 2026

MSMPP: Molecular Property Prediction by Integrating Multi-scale Multi-view information with pretrained 3D molecular large model representation.

Evaluations on eight MoleculeNet datasets show that MSMPP significantly outperforms state-of-the-art models, demonstrating its effectiveness in integrating multi-view intra-molecular features, inter-molecular features and cross-task information.

Jiongfeng Chen, Yulian Ding, Yan Yan et al. · 0 citations
Open access Jul 2026

Path-weighted atom vectors and ChemBERTa fusion for predicting physicochemical properties

Experiments show that PWAV generally improves over classical fingerprint descriptors within learned models and achieves competitive performance relative to established external baselines on several endpoints, positioning PWAV as a competitive and chemically transparent component for hybrid molecular property prediction, rather than as a replacement for domain-specific benchmark systems.

M. Afzal, S. Siddiqi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.