Skip to content
Preprint

Synthesizing like a chemist: an iterative, feedback-driven loop for materials discovery

Aug 2026 · 0 citations · 2 references
Physics

Abstract

Most computationally predicted materials are never synthesized because conventional synthesis optimization is slow, expertise-dependent, and iterative. Here we present a closed-loop framework that automates this expert workflow by placing human tacit knowledge in the loop through a large language model (LLM) that distills synthesis knowledge from the literature, high-throughput hyperspectral imaging for rapid film evaluation, and multi-objective Bayesian optimization guided by experimental feedback. In a paired optimization campaign, LLM-assisted initialization produced more Pareto-optimal samples and higher hypervolume than a Latin hypercube sampling baseline at matched trial counts, and this advantage persisted throughout iterative optimization. We demonstrate the framework by synthesizing the previously unreported perovskite-inspired compound Rb3BiI6 as thin films and validating the optimized films by optical bandgap analysis and X-ray diffraction. The framework transforms synthesis prediction from single-shot recommendation to iterative learning, providing a generalizable strategy to accelerate automated and fully autonomous experimental materials discovery.

View source

Similar papers

Open access Aug 2026

Closed-Loop Solid-State Synthesis Planning for Materials Discovery With Large Language Models.

Developing reliable synthesis routes for complex materials remains a major bottleneck in accelerating materials discovery. This study establishes a large language model-based framework for predicting and optimizing synthesis conditions directly from the literature data. Key synthesis information, including target compounds, precursors, and processing parameters, was systematically extracted from 4407 open-access solid-state synthesis papers and organized into a structured recipe dataset. Using a retrieval-augmented generation (RAG) approach, the system first retrieves similar recipes from the corpus and then generates a new candidate recipe conditioned on those exemplars. The generated recipes were benchmarked against literature data using quantitative scoring metrics, achieving strong agreement with experimentally reported conditions. To validate the predictive capability, the framework was applied to unreported solid-state electrolyte candidates identified through first-principles screening, and multiple oxy-selenide compounds were successfully synthesized through iterative feedback between the model and experiment. The recipe generator accurately refined synthesis parameters over successive trials, demonstrating its ability to reproduce phase-pure products while minimizing trial-and-error. This approach establishes a data-driven, feedback-optimized route to accelerate synthesis design, offering a generalizable paradigm for integrating language models into experimental materials research.

Dong Won Jeon, Dong Hwi Kim, Taeyang Jeon et al. · 0 citations
Jul 2026

(Invited) Autonomous Experiments for Solid Materials: From Thin Films to Bulk Synthesis

Autonomous experiments that integrate machine learning and robotics are reshaping materials research. By automating experimental workflows and efficiently searching high-dimensional parameter spaces, these approaches markedly accelerate materials discovery and process optimization. Here, we report a modular self-driving laboratory (SDL) for solids and thin films [1–4]. The SDL orchestrates all stages of the experimental cycle—including sample transfer, synthesis, characterization, and iterative optimization. Data acquisition spans X-ray diffraction, scanning electron microscopy, Raman spectroscopy, electrical conductivity and optical transmittance measurements. A Bayesian optimization enables autonomous exploration of the parameter space and rapid identification of optimal conditions. We demonstrate the platform by synthesizing thin films of TiO₂ and LiCoO 2 . We further show that the same workflow supports the discovery of new ionic conductors. These results highlight the potential of autonomous experimentation to accelerate research in solid-state materials. Ongoing efforts extend the SDL to bulk-materials synthesis, aiming to unify thin-film and bulk workflows within a single autonomous framework. [1] "Autonomous experimental systems in materials science" N. Ishizuki, R. Shimizu, and T. Hitosugi, STAM Methods 3, 2197519 (2023). [2] "Autonomous materials synthesis by machine learning and robotics" R. Shimizu, T. Hitosugi et al. , APL Mater. 8111110 (2020). [3] “Autonomous exploration of an unexpected electrode material for lithium batteries“ S. Kobayashi, T. Hitosugi et al. , ACS Materials Lett. 5, 2711–2717 (2023). [4] “Digital laboratory with modular measurement system and standardized data format” K. Nishio, T. Hitosugi et al. , Digital Discovery 4, 1734-1742 (2025). Figure 1

T. Hitosugi · 0 citations
Open access Aug 2026

Machine-Learning-Assisted Discovery and Accelerated Synthesis of Metal Phosphosulfides.

Metal phosphosulfides have emerged as unique multifunctional materials, but they present unique synthesis challenges compared to more established material classes such as oxides and nitrides. As a consequence, experimental development and theoretical understanding of phosphosulfides have focused on individual compounds rather than on accelerated broad-range exploration. In this work, we first evaluate the synthesizability and band gaps of 909 hypothetical ternary phosphosulfides by density functional theory. We find 19 previously unknown thermodynamically stable compounds, including the first Si- and Ge-based phosphosulfides. For rapid band gap prediction, we then develop a multi-fidelity machine learning model to translate semilocal density functional theory band gaps into experimentally calibrated band gaps. Importantly, we extend the accelerated material development workflow to the experimental domain by demonstrating a route to high-throughput synthesis and characterization of virtually any phosphosulfide material system. The method is based on thin-film combinatorial libraries and yields over 100 unique compositions in each experiment, enabling us to synthesize four distinct phosphosulfide compounds in only four combinatorial experiments without prior synthesis recipes and without compromising on material quality. Thus, we argue that accelerated materials development workflows combining theory, artificial intelligence, synthesis, and characterization can be viable even for experimentally challenging inorganic materials.

J. Sanz Rodrigo, Nicholas A. Kryger-Nelson, Lena A Mittmann et al. · 0 citations
Open access Jul 2026

Accelerating sustainable glass discovery: integrating molecular dynamics, machine learning, and robotic synthesis

The Simulation-Calibrated Active Learning Estimator (SCALE), a closed-loop framework uniting high-throughput molecular dynamics, machine learning, and robotic synthesis to bridge the gap between simulation and experiment, is introduced.

Felix Arendt, T. Waurischk, Stefan Reinsch et al. · 0 citations
Preprint Jul 2026

Chemical filters for ultra-high-throughput materials screening and generation

Generative artificial intelligence is rapidly transforming materials design by enabling de novo exploration of immense chemical spaces. Yet a large proportion of AI-generated compositions remain implausible, violating established chemical principles, which limits the reliability and interpretability of generative materials design. Here, we introduce a chemical validity operator that recasts heuristic chemical rules as a configurable algorithmic prior for evaluating and guiding generative materials discovery. Built on the open-source SMACT package, a data-informed oxidation-state model exposes tunable thresholds, allowing users to interpolate continuously between permissive and conservative chemical constraints, while supporting both exploratory and conservative materials-design workflows. Benchmarking six state-of-the-art generative models for inorganic crystals shows that most reproduce stoichiometry but under-represent realistic oxidation-state combinations, and that filtering removes compositions reliant on rarely observed oxidation states while preserving low-energy compounds near the convex hull. Beyond screening, the same operator can also serve as a reinforcement-learning reward, steering a latent diffusion model towards chemically grounded compositions. By encoding chemical heuristics and observations, this work establishes a foundation for oxidation-state-aware generative models.

Kinga O. Mastej, Panyalak Detrattanawichai, Hyunsoo Park et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.

B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al. · 0 citations