The integration of machine learning (ML) into materials science offers a transformative pathway toward fully autonomous synthesis workflows. For precise thin-film deposition techniques like molecular beam epitaxy (MBE), this automation is critical to overcome the time-consuming, manual navigation of high-dimensional thermodynamic phase spaces. Existing approaches for ML-assisted thin film growth predominantly rely on continuous Bayesian optimization (BO) models that assume smooth parameter landscapes. Consequently, they struggle to capture the abrupt crystallographic phase boundaries and narrow growth windows inherent to binary quantum materials. Here, we demonstrate an active learning protocol based on Sequential Model-Based Optimization (SMBO) designed specifically for the closed-loop MBE of such compounds. To overcome the limitations of continuous models while retaining the efficient exploration-exploitation logic of traditional BO, we combine a random forest surrogate model capable of capturing highly non-linear phase transitions and thermodynamics constraints of the growth process with an expected improvement function to predict optimum growth parameters. We apply this combined SMBO framework to the MBE of the topological Weyl ferromagnet Fe$_3$Sn, which exists as a metastable line compound. Using a small initial training set of fewer than twenty growth iterations, our active learning loop rapidly navigates a complex optimization landscape to identify an optimum growth window bounded by sharp transitions. Within only four active learning iterations, the absolute predictive error is halved to $\approx10\%$. This data-efficient framework paves the way for the autonomous discovery and thin-film synthesis of functional quantum materials.
The Simulation-Calibrated Active Learning Estimator (SCALE), a closed-loop framework uniting high-throughput molecular dynamics, machine learning, and robotic synthesis to bridge the gap between simulation and experiment, is introduced.
Felix Arendt, T. Waurischk, Stefan Reinsch et al.· npj Computational Materials· 0 citations
Most computationally predicted materials are never synthesized because conventional synthesis optimization is slow, expertise-dependent, and iterative. Here we present a closed-loop framework that automates this expert workflow by placing human tacit knowledge in the loop through a large language model (LLM) that distills synthesis knowledge from the literature, high-throughput hyperspectral imaging for rapid film evaluation, and multi-objective Bayesian optimization guided by experimental feedback. In a paired optimization campaign, LLM-assisted initialization produced more Pareto-optimal samples and higher hypervolume than a Latin hypercube sampling baseline at matched trial counts, and this advantage persisted throughout iterative optimization. We demonstrate the framework by synthesizing the previously unreported perovskite-inspired compound Rb3BiI6 as thin films and validating the optimized films by optical bandgap analysis and X-ray diffraction. The framework transforms synthesis prediction from single-shot recommendation to iterative learning, providing a generalizable strategy to accelerate automated and fully autonomous experimental materials discovery.
Fang Sheng, Steven B. Torrisi, Amanda A. Volk et al.· 0 citations
Machine-learning-accelerated materials discovery has yielded large numbers of computationally stable compounds, yet many remain experimentally unrealized, underscoring a persistent gap between prediction and synthesis. Here, we introduce a hierarchical screening framework that combines PBE-based thermodynamic stability, efficient dynamical-stability screening enabled by universal machine-learning interatomic potentials, and SCAN-based thermodynamic refinement. Applying this protocol to the 894 stable materials previously reported in Sci. Data 9, 302 (2022), we first curate 603 unique structures, of which only 298 remain thermodynamically stable on the complete PBE phase diagrams, demonstrating the critical role of competing phases in stability assessment. Dynamical screening then identifies 166 materials stable under both harmonic-phonon and finite-temperature molecular dynamics criteria, and SCAN phase diagrams further narrow this set to 109. Finally, by combining decomposition enthalpy with chemical-space completeness, we prioritize 25 candidates as high-confidence targets for experimental synthesis. This work provides a practical protocol for translating stability predictions into experimentally actionable synthesis targets, closing a key gap in machine-learning-driven materials discovery.
Yuqi An, Sihong Zhu, Joseph H. Montoya et al.· 0 citations
The development of quantum chemistry has long been shaped by a central tension: while the laws governing electronic structure are known, their exact application quickly becomes computationally prohibitive for realistic molecular systems. Over decades, this challenge has driven the design of increasingly sophisticated approximations that balance predictive accuracy with computational affordability. More recently, machine learning (ML) has emerged as a new addition to this methodological landscape, offering the possibility of reproducing high-level quantum chemical results at a fraction of the cost.
This thesis explores how ML can contribute to this long-standing objective in a particularly resource-conscious way. Rather than treating ML purely as a black-box substitute for quantum chemistry, the work asks a broader methodological question: under finite budgets for data generation, training and inference, what is the most efficient way to use data-driven models to accelerate quantum chemical simulations? Across the different applications studied here, the guiding principle has been to identify the simplest effective strategy for the problem at hand while retaining as much physical structure and reusing as much existing data as possible.
A first key result of this thesis is that substantial acceleration can, in some settings, be achieved with remarkably simple models. In the context of basis set extrapolation for GW quasiparticle energies, a linear regression model based on molecular orbital descriptors was shown to recover near-complete basis set accuracy from finite-basis calculations. This demonstrates that when the underlying quantum chemical representations already contain the essential physical information, even lightweight statistical models can provide acceleration while maintaining high-level quantum chemical accuracy.
As many ML applications, especially neural networks, critically depend on sufficiently broad and reliable training data, part of this work focused on constructing a large-scale dataset of quasiparticle self-consistent GW (qsGW) quasiparticle energies, GW Bethe--Salpter equation (GW-BSE) neutral excitation energies, transition dipole moments and oscillator strengths across a chemically diverse space of organic molecules. Building on this foundation, a graph neural network was trained for the prediction of charged and neutral excitation energies. A central finding is that transfer learning from lower-fidelity but already widely available data sources, such as molecular orbital energies from density functional theory (DFT) and excitation energies from time-dependent DFT (TDDFT), can substantially improve the prediction of the high-fidelity qsGW and GW-BSE targets. In this way, previously generated computational data become a powerful resource for reducing the cost of expensive reference calculations.
The question of how simple ML models can remain while still being effective was further investigated for solvation energies and geometry optimization in solution. Here, the results show that relatively simple graph neural network architectures can already yield accurate predictions of Gibbs solvation energies for highly charged molecules. At the same time, ML was also used to parametrize and correct established physically grounded solvation models. These results suggest that, in such settings, hybrid strategies that combine explicit physical models with learned components can be as effective as fully data-driven approaches while retaining the robustness and interpretability of the underlying physical description.
Taken together, the work presented in this thesis shows that efficient ML for quantum chemistry does not rely on a single universally optimal model class. Instead, the most effective strategies arise from matching the complexity of the statistical model to the physical structure of the problem, reusing data across different levels of quantum chemical theory and retaining established physical models wherever they already provide reliable inductive bias. In this sense, ML serves not simply as a faster predictor but as a flexible methodological tool for extending the practical reach of quantum chemical simulations through resource-efficient acceleration.
Metal phosphosulfides have emerged as unique multifunctional materials, but they present unique synthesis challenges compared to more established material classes such as oxides and nitrides. As a consequence, experimental development and theoretical understanding of phosphosulfides have focused on individual compounds rather than on accelerated broad-range exploration. In this work, we first evaluate the synthesizability and band gaps of 909 hypothetical ternary phosphosulfides by density functional theory. We find 19 previously unknown thermodynamically stable compounds, including the first Si- and Ge-based phosphosulfides. For rapid band gap prediction, we then develop a multi-fidelity machine learning model to translate semilocal density functional theory band gaps into experimentally calibrated band gaps. Importantly, we extend the accelerated material development workflow to the experimental domain by demonstrating a route to high-throughput synthesis and characterization of virtually any phosphosulfide material system. The method is based on thin-film combinatorial libraries and yields over 100 unique compositions in each experiment, enabling us to synthesize four distinct phosphosulfide compounds in only four combinatorial experiments without prior synthesis recipes and without compromising on material quality. Thus, we argue that accelerated materials development workflows combining theory, artificial intelligence, synthesis, and characterization can be viable even for experimentally challenging inorganic materials.
J. Sanz Rodrigo, Nicholas A. Kryger-Nelson, Lena A Mittmann et al.· Small· 0 citations
The ElemeNet software package enables the training of advanced ML models for diverse properties and datasets with an enlarged range of elemental compositions, and introduces moiety predictions, a unified, general-purpose software package for molecular machine learning.
Jacob W. Toney, S. Darouich, Yiran Wang et al.· arXiv.org· 0 citations