Skip to content

Procedural Pretraining for Molecular Property Prediction

Sep 2026 · 0 citations · 17 references
Computer Science

TL;DR

It is found that procedural pretraining can improve molecular property prediction even after subsequent molecular pretraining, and procedural data can provide transferable structure for molecular learning and offer a complementary route to improving performance when labeled molecular data are limited.

Abstract

Molecular property prediction is often limited by the small size of labeled downstream datasets, motivating pretraining on large corpora of unlabeled molecules. In this work, we ask whether useful inductive biases can instead be learned from abstract, procedurally generated data before a model sees any molecular data. We introduce a three-stage training pipeline consisting of procedural pretraining, molecular pretraining on SMILES, and downstream fine-tuning, and evaluate several procedural tasks spanning sequence structure, cellular automata, and graph reasoning. We find that procedural pretraining can improve molecular property prediction even after subsequent molecular pretraining: on Lipophilicity, \textsc{Reverse} reduces test error by 4.8\%. For context, the magnitude of this improvement is roughly 90\% of the performance difference between our 250K-molecule baseline and the publicly released MoLFormer checkpoint pretrained on approximately 100M molecules. Our analysis shows that the benefit is strongest under downstream data scarcity, depends on the structure of the procedural data rather than only surface-level statistics, and does not increase monotonically with additional procedural training. Instead, transfer typically peaks at an intermediate procedural budget and deteriorates as the model approaches convergence on the procedural task. We further find that, for several tasks, much of the transferable information is localized in the attention layers, while feed-forward layers can contribute to over-specialization. These results show that procedural data can provide transferable structure for molecular learning and offer a complementary route to improving performance when labeled molecular data are limited.

View source

Similar papers

Preprint Aug 2026

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is presented, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 81 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding de...

Blazej Banaszewski, Andrew W. Fitzgibbon · 1 citation
#machine learning Preprint Sep 2026

Self-Supervised Pretraining of Molecular Graph Encoders with LeJEPA

LeJEPA is adapted to molecular graphs, evaluating GPS and Chemprop-style D-MPNN encoders on the antibiotic-activity dataset and ogbg-molhiv using a multi-seed, bootstrap-based protocol, and pretraining supplies complementary information best realised through feature-level combination, while finetuning gains are weak an...

Michał Kulczykowski, Rafał Łabędzki · 0 citations
Preprint Aug 2026

Multi-Granular Rationale-Guided Molecular LLM for Property Prediction

This is the first method to expose GNN-derived attributions to an LLM as evidence for property prediction, and achieves the best overall results among generalist models and narrows the gap to specialist models tuned for each task.

Junwoo Park, Minyoung Shin, C. Lee et al. · 0 citations
#machine learning Preprint Sep 2026

WEECFP-SuRGE: A Position-Aware Substructure Encoding Method for Molecular Property Prediction

Computational molecular property prediction requires representations that capture local chemistry, long-range interactions, and molecular topology. Conventional fingerprints provide efficient local substructure features, whereas learned graph and sequence models can represent broader context but often rely on pretraini...

Robert W. Epps · 0 citations
Open access Aug 2026

MolPACL: Molecular Property Prediction Based on Prompt Augmentation and Contrastive Learning.

MolPACL is proposed, a prompt-augmentation-based supervised contrastive learning framework that incorporates high-level chemical semantics while preserving molecular identity, and achieves strong performance on both classification and regression tasks while reducing training cost.

Ali Forooghi, Luis Rueda, A. Ngom · 0 citations
#artificial intelligence Preprint Aug 2026

CoMPASS: Collaborative Molecular Property Prediction via Adaptive Small-Large Model Synergy

CoMPASS is presented, a retrieval-calibrated framework for small-large model collaboration that retains a graph attention network as the predictive anchor, retrieves locally relevant training molecules, provides attention-grounded evidence to an LLM, and converts its proposal into a bounded correction through an agreem...

Wen-Tao Li, Jiang-Jie Qiu, Yi-Jun Li et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.