Skip to content

Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design

Jun 2026 · arXiv.org · Vol abs/2606.28578 · 0 citations · 51 references
Physics Computer Science

Abstract

Closed-loop materials discovery iterates between proposing candidate structures and evaluating their properties, and property evaluation dominates the cost. In the generative variant, a learned prior proposes candidate crystals and a property oracle scores them; we ask whether a cheap probabilistic surrogate can triage the generator's output, and what such a surrogate must do well. Across three architecturally distinct pretrained diffusion priors (MatterGen, CrystalFlow, ADiT) and two targets (room-temperature heat capacity and bulk modulus), we insert a Gaussian process acquisition gate between structure generation and the oracle in an RL-steered generative workflow. The gate matches or exceeds ungated fine-tuning of the generative model while capping oracle calls at a fixed per-cycle budget. Budget-matched ablations isolate the mechanism. At an identical four-call budget, ranking-based selection outperforms arbitrary selection, confirming that the gain comes from the surrogate's choice; the gate comes within $\sim$9\% of exhaustive oracle spending at roughly one-fifth of the calls. A density-functional-theory check of the bulk-modulus discoveries confirms the learned oracle to within 2.5\% on average and the surrogate's ranking of the generated structures at Spearman $\rho = 0.94$. A cross-factorial benchmark of surrogate performance spanning mechanical, electronic, and vibrational properties identifies pretrained ORB embeddings with a Gaussian process as the most reliable combination, which we adopt as the building blocks of the proposed workflow. The complete pipeline is released as open-source software.

View source

Similar papers

Preprint Jul 2026

Chemical filters for ultra-high-throughput materials screening and generation

Generative artificial intelligence is rapidly transforming materials design by enabling de novo exploration of immense chemical spaces. Yet a large proportion of AI-generated compositions remain implausible, violating established chemical principles, which limits the reliability and interpretability of generative materials design. Here, we introduce a chemical validity operator that recasts heuristic chemical rules as a configurable algorithmic prior for evaluating and guiding generative materials discovery. Built on the open-source SMACT package, a data-informed oxidation-state model exposes tunable thresholds, allowing users to interpolate continuously between permissive and conservative chemical constraints, while supporting both exploratory and conservative materials-design workflows. Benchmarking six state-of-the-art generative models for inorganic crystals shows that most reproduce stoichiometry but under-represent realistic oxidation-state combinations, and that filtering removes compositions reliant on rarely observed oxidation states while preserving low-energy compounds near the convex hull. Beyond screening, the same operator can also serve as a reinforcement-learning reward, steering a latent diffusion model towards chemically grounded compositions. By encoding chemical heuristics and observations, this work establishes a foundation for oxidation-state-aware generative models.

Kinga O. Mastej, Panyalak Detrattanawichai, Hyunsoo Park et al. · 0 citations
Preprint Jul 2026

Transfer Learning Architectures for Scalable Multi-Fidelity Bayesian Optimization

Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, expensive ones, with a predictive surrogate at its core. Gaussian processes (GPs) are the default choice, but they scale poorly as data accumulate and assume a smooth landscape that molecular and materials search spaces routinely violate. Transfer learning offers an alternative suited to this regime: it learns a representation from abundant cheap data and adapts it to sparse expensive data. Despite its use in property prediction, transfer learning has not been tested as the engine of a closed-loop optimization. Here we benchmark eleven transfer-learning surrogates against four GP methods under an identical selection rule, fidelity budget, and model size, across nine tasks spanning synthetic functions to real chemistry and materials problems. GPs win on smooth, low-dimensional functions but perform worst on molecular and materials problems, where transfer-learning surrogates reach substantially better solutions using far less computation. Because acquisition policy is held fixed across surrogates, this advantage is attributable to the surrogate itself. Uncertainty-driven exploration is not reliably beneficial, and calibration does not predict optimization performance, so greedy exploitation of the transfer-learned mean is the more robust default. Transfer learning is therefore the surrogate of choice for molecular and materials MFBO.

Jaewook Lee, Ethan Errington, Christian D. Lorenz et al. · 0 citations
Preprint Jul 2026

ATLAS: A Foundation Neural Sampler for Amorphous Materials

Amorphous materials exhibit exceptional mechanical and functional properties, yet their rugged energy landscapes are notoriously difficult to sample. Below the glass-transition temperature, conventional molecular dynamics and Monte Carlo become inefficient because equilibration relies on rare barrier-crossing events, while data-driven generative models are constrained by scarce and biased reference ensembles. Here, we introduce ATLAS, an efficient sampler that learns a diffusion process to generate Boltzmann-distributed amorphous structures directly from a target energy function. Parameterized by an equivariant graph neural network, ATLAS generalizes across system size, temperature, and composition. By exploiting the time reversal of the diffusion process, it enables efficient estimation of thermodynamic quantities and steering toward target observables. In two-dimensional Kob-Andersen systems, ATLAS reproduces parallel tempering Markov chain Monte Carlo structural distributions, free energies and entropies, achieving below 0.2% free energy error in the low-temperature glass regime with over 500-fold fewer energy evaluations. In Cu-Zr and Cr-Co-Ni metallic glasses, ATLAS recovers experimentally observed short-range-order trends and steers structures toward prescribed order parameters and optimized bulk moduli. Moreover, composition-amortized pretraining outperforms composition-specific training from scratch, reduces inverse-design costs by several hundred-fold, and enables sampling with expensive universal machine learning interatomic potentials. Coupled to a large language model agent, ATLAS searches an eight-element space for high-entropy metallic glasses balancing stiffness and ductility, identifying a converged Pareto frontier within 480 oracle evaluations. Together, these results establish ATLAS as a foundation model for sampling, steering and designing amorphous materials.

Mouyang Cheng, Denis Blessing, Botao Yu et al. · 1 citation
Review Open access Aug 2026

Generative Models in Inorganic Crystals Discovery and Inverse Design

A review of physically constrained, multimodal and closed‐loop workflows as the clearest route from generative crystal models to experimentally actionable candidates and defines the inverse‐design problem and main generation tasks.

Tao Li, Xiaolin Liu, Fei Wang et al. · 0 citations
Open access Nov 2025

Guiding generative models to uncover diverse and novel crystals via reinforcement learning

A reinforcement learning framework that guides latent denoising diffusion models in finding diverse and novel, yet thermodynamically viable, crystalline compounds and demonstrates enhanced property-guided design that preserves chemical validity while targeting desired functional properties.

Hyunsoo Park, Aron Walsh · 27 citations · ⚡3
Preprint Jul 2026

Projected Energy Matching for Generative 3D Priors

Energy Matching has emerged as a powerful generative framework that combines flow model efficiency with the explicit likelihood of Energy-Based Models (EBMs) via a single, time-independent scalar potential. However, directly training this potential on high-dimensional 3D data remains computationally challenging. While distilling a pre-trained flow model circumvents some of the initial training costs, we demonstrate that velocity fields inevitably contain non-conservative rotational artifacts (curl). Forcing a strictly conservative scalar potential to match this unconstrained field creates a"structural conflict", which degrades generation quality and mode coverage. To solve this, we propose Projected Energy Matching, a scalable framework that resolves these structural and computational bottlenecks. We introduce Helmholtz Distillation, a structural relaxation that leverages a Hutchinson trace estimator to explicitly absorb rotational noise into an auxiliary residual network. We subsequently refine this landscape using Negative Caching, a memory-efficient strategy that reuses negative samples across micro-batches, rendering sampling tractable during contrastive training with gradient accumulation. We deploy our method as an unconditional prior for real-world medical CT inverse problems, specifically sparse-view reconstruction. Ultimately, our amortized pipeline reduces total compute to a small fraction of that required by standard energy matching, while achieving high-fidelity reconstructions and successfully resolving severe measurement artifacts.

Daniel Barco, M. Balcerak, Suprosanna Shit et al. · 0 citations