Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· 0 citations· 80 references
TL;DR
A geometry-adaptive Riemannian framework for molecular representation learning, which explicitly models motifs as the basic units and learns their embeddings across multiple constant-curvature spaces, offering a general and scalable framework for scientific molecular modeling.
Abstract
Molecular properties are often governed by a small number of local substructures, or motifs, whose topologies can vary drastically across molecules. Existing molecular representation learning approaches typically embed all motifs into a single Euclidean or fixed-curvature space, which fails to capture the motif-level topological heterogeneity and leads to geometric mismatch, impairing property prediction. To address this challenge, we propose a geometry-adaptive Riemannian framework for molecular representation learning, which explicitly models motifs as the basic units and learns their embeddings across multiple constant-curvature spaces. Each motif is adaptively aligned with the geometric space that best fits its intrinsic topology, enabling simultaneous modeling of cyclic, hierarchical, and tree-like structures. Motif embeddings are then aggregated into molecule-level representations, emphasizing functional substructures while suppressing irrelevant background. Extensive experiments on benchmark molecular property prediction datasets demonstrate that our approach outperforms state-of-the-art baselines, shows strong generalization under distribution shifts, and provides interpretable motif-level insights, offering a general and scalable framework for scientific molecular modeling. Our code is available at https://github.com/qimuya/mo-mi-r.
Results show that explicitly teaching the relation between a molecule and its structural core can reliably shape the organization of molecular embedding space, while the extent of usefulness of this organization remains task dependent.
David Sulu, Lorenzo Di Fruscia, Jana M. Weber· 0 citations
OmniScore is introduced, a universal structure-based framework that learns a shared geometry-aware representation of complexes once and then adapts it to downstream scoring through lightweight task-specific heads, suggesting that geometry-aware pretraining can provide a reusable scoring backbone for tasks that depend o...
Medical image analysis remains fundamentally challenging because of the intricate geometric and topological structures present in medical data. Conventional convolutional neural networks model images as regular Euclidean grids, limiting their ability to preserve geometric relationships and higher-order structural infor...
A. Wachira, Xiang Liu, Zhe Su et al.· arXiv.org· 0 citations
This work shows that informative embeddings can be derived without complicated model design and gradient-based training, and suggests that informative graph embeddings can arise from carefully chosen topological transformations before any learning operation is applied.
Meng Qin, Jin-Qiang Cui, Hong-Wei Zheng et al.· 0 citations
This work proposes a novel approach that learns hierarchical discrete representations of protein structures using vector quantization, and outperforms state-of-the-art models such as ESMDiff across challenging benchmark datasets, including BPTI MD trajectories and conformational-changing pairs.
Seokjun On, Yujin Jeong, Kanghyeon Kim et al.· Bioinformatics· 0 citations
Analysis of HiFi-Mol reveals that fragment-aware masking improves graph representation quality, and classification results demonstrate dataset-dependent strengths of the individual graph and fingerprint variants, confirming that the two views provide complementary predictive signals.
Gwang-Hyeon Yun, Jong-Hoon Park, Bing Hu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.