Skip to content
Preprint

Computed materials proposals depart from the structural memory of experimental discovery

Jun 2026 · 0 citations
Physics

TL;DR

This work embeds 167,500 Inorganic Crystal Structure Database entries in a continuous structural-similarity space, partition it into graph communities, and replay them in time to define a historical synthesizability prior for triaging computed materials.

Abstract

Generative AI and high-throughput DFT pipelines propose millions of inorganic crystal structures, but lack a calibrated reference frame against experimentally realized chemistry. Here we embed 167,500 Inorganic Crystal Structure Database entries in a continuous structural-similarity space, partition it into graph communities, and replay them in time. Experimental discovery shows strong structural memory: 82.9% of new formulas enter pre-existing communities; new-community formation falls from 40.2% (1930s) to 2.6% (2010s). The communities are chemically meaningful, positively identifying nine textbook field-defining renaissances, including cuprates, colossal-magnetoresistance manganites, MAX phases, and Li-ion battery cathodes. Projecting GNoME, MatterGen-public, Materials Project, JARVIS-DFT, and Alexandria-PBE into frozen historical maps yields a cutoff-robust ordering: held-out ICSD>MatterGen>{GNoME ~ MP-theoretical}>JARVIS>Alexandria. Structural departure from experimental basins is not specific to generative AI but general across the tested computed sets. Combining structural proximity with reduced-formula precedent defines a historical synthesizability prior for triaging computed materials.

View source

Similar papers

Open access Oct 2026

A century of structures: historical fidelity and computational fitness in the Inorganic Crystal Structure Database (ICSD).

The earliest entries in the Inorganic Crystal Structure Database (ICSD) are both primary historical documents and active computational inputs. They record how the Braggs, Goldschmidt, Vegard, and others established the structural foundations of inorganic chemistry; they are also served, without distinction, to automated pipelines, machine-learning training sets, and high-throughput workflows. We assess 464 ICSD structures from 1913 to 1929 using two independent tools: Eir v1.4.2 (bond valence sums, global instability index) and MaplePy (Madelung part of lattice energy). The pre-1925 median global instability index is 0.077 valence unit (v.u.), below that of modern rock salt reference data (mean 0.192 v.u.); the Braggs' 1913 NaCl gives 0.005 v.u., Bragg's 1922 corundum 0.001 v.u. Early crystallographers were doing simple structures and getting them right. The problem is not quality but metadata. A small number of entries carry genuine topological errors (BeO in NaCl type, SnTe in zinc blende) whose corrective information already exists in database Comments fields but is not machine-readable, not reflected in quality classifications, and relatively invisible to downstream users. We propose six automated fitness-for-purpose flags that surface these cases without removing or modifying any historical record.

Peter Gross, Nik Reeves-McLaren · 0 citations
Preprint Jun 2026

Adaptive fine-tuning of foundation models for crystal structure prediction: Discovery of high-pressure phases in the CaFeNi system

The prediction of crystal structures is a key challenge in chemistry and materials science, but evolutionary crystal structure prediction (CSP) remains computationally expensive because it relies on repeated \textit{ab initio} relaxations and energy ranking. Machine learning interatomic potentials (MLIPs) can accelerate CSP, yet their use is limited by the need for large training sets and by the difficulty of choosing which candidate structures should be labeled by density functional theory (DFT). Here we introduce a self-consistent, foundation-model-assisted CSP workflow that combines evolutionary search with adaptive data selection and fine-tuning. Starting from a pretrained MLIP, the algorithm rapidly explores configuration space while iteratively selecting compact, representative, and physically relevant subsets of structures for DFT labeling, thereby reducing redundant calculations and improving a system-specific potential. We apply the method to the chemically complex Ca--Fe--Ni ternary system. The workflow reproduces the known low-pressure convex hull and enables efficient high-pressure exploration. It predicts a previously unreported compound, Ca$_6$FeNi, which becomes thermodynamically stable above 100~GPa. These results show that foundation-model-based, data-efficient CSP can greatly reduce computational cost while preserving accuracy and enabling the discovery of new materials in complex multicomponent systems.

N. Chtchelkatchev, M. Magnitskaya, R. Ryltsev · 0 citations
Preprint Jul 2026

Correcting DFT formation energies towards experimental accuracy using foundational MLIPs and latent-feature delta-learning

Crystal structure databases curated by high-throughput density functional theory calculations typically serve as the starting point for computational materials discovery efforts. Thermodynamic stability data, such as formation energies and the energy above the convex hull, are important quantities to guide the search for novel materials, enabling filtering for (meta)stable structures. Here, we present the thermodynamic stability of the fully open-source, reproducible, and experimentally focused Materials Cloud three-dimensional crystals database (MC3D). We compare against two other DFT databases, the Open Quantum Materials Database (OQMD) and the Materials Project (MP), as well as against experimental formation enthalpies. We then demonstrate how recent foundational machine learning interatomic potentials (MLIPs) trained at the r$^2$SCAN level (specifically, we test PET-OMATPES here) can be leveraged to improve the agreement of formation energies with experiment, reducing the mean absolute error by more than 40% relative to GGA without requiring any additional DFT calculation. Our results validate and extend the established practice of combining PBEsol geometries with meta-GGA energies to the era of foundational MLIPs. Finally, we train classical machine learning models to further correct the formation energies in a delta-learning framework, where we use the information-rich latent features of the foundational MLIP. These models further reduce the mean absolute error below 50 meV/atom, bringing it down to values comparable with the experimental uncertainty itself. Notably, compared to purely compositional features, the latent features (combined with carefully tuned regularization) simultaneously reduce the prediction error and limit the impact of the learned corrections on the relative phase stability.

Timo Reents, Marnik Bercx, Giovanni Pizzi · 0 citations
Preprint Jul 2026

Predicting Novel Stable Materials for Experimental Synthesis

Machine-learning-accelerated materials discovery has yielded large numbers of computationally stable compounds, yet many remain experimentally unrealized, underscoring a persistent gap between prediction and synthesis. Here, we introduce a hierarchical screening framework that combines PBE-based thermodynamic stability, efficient dynamical-stability screening enabled by universal machine-learning interatomic potentials, and SCAN-based thermodynamic refinement. Applying this protocol to the 894 stable materials previously reported in Sci. Data 9, 302 (2022), we first curate 603 unique structures, of which only 298 remain thermodynamically stable on the complete PBE phase diagrams, demonstrating the critical role of competing phases in stability assessment. Dynamical screening then identifies 166 materials stable under both harmonic-phonon and finite-temperature molecular dynamics criteria, and SCAN phase diagrams further narrow this set to 109. Finally, by combining decomposition enthalpy with chemical-space completeness, we prioritize 25 candidates as high-confidence targets for experimental synthesis. This work provides a practical protocol for translating stability predictions into experimentally actionable synthesis targets, closing a key gap in machine-learning-driven materials discovery.

Yuqi An, Sihong Zhu, Joseph H. Montoya et al. · 0 citations
Jul 2026

Active-Learning Discovery of Superionic Compositions Using High-Throughput EIS and Structure-Aware Descriptors

We introduce an active-learning framework that closes the loop between high-throughput EIS measurements and structure-aware composition descriptors to discover superionic candidates under realistic processing constraints. Starting from a small seed set, Gaussian-process and tree-based models propose batched experiments that maximize information gain on conductivity and activation energy while enforcing uncertainty-aware Kramers–Kronig quality gates. Descriptor families integrate interpretable features: ionic radius mismatch, framework softness, site connectivity from simple graph-derived motifs, and processing proxies (grain size from Scherrer, porosity, interphase penalty terms). We demonstrate rapid convergence to high-conductivity regions in multi-component chalcogenide and halide spaces using the automated multi-site EIS workflow described separately. Across three material spaces, the approach reduces experiments ~3× versus grid sampling while yielding candidates with improved conductivity at moderate temperatures and stable impedance upon cycling. We release a lightweight, reproducible stack (metadata schema, analysis notebooks, and synthetic datasets) to encourage community benchmarking without proprietary infrastructure. The result is a pragmatic path to self-driving electrolyte discovery that prioritizes experimental tractability and interpretability—features that matter for industrial translation and cross-lab reproducibility. Keywords: active learning; Bayesian optimization; EIS QC; interpretable descriptors; high-throughput screening; solid electrolytes

Progna Banerjee · 0 citations
Preprint Jul 2026

Stoichiometric cluster learning for few-shot property prediction of multi-ionic integrated energetic materials

Multi-ionic materials pose a distinct representational challenge in machine learning-driven materials design. Different from single-molecule or composition-based materials, their properties arise from how charged building blocks aggregate into specific assemblies. Here, we show how pretrained machine-learned interatomic potentials (MLIPs) can bypass full crystal-structure prediction and support pre-synthesis screening from stoichiometric ionic clusters using multi-ionic integrated explosives (MIXs) as a synthesis-facing example. This strategy combines a stoichiometric ionic-cluster representation, which represents each candidate material by a non-periodic, stoichiometry-preserved formula-unit cluster, with multi-task fine-tuning (MT-FT), which adapts a pretrained atomistic backbone while retaining the energy--force objective as physical regularization for the sparse detonation-velocity labels. With the pretrained backbone regularized by MT-FT, this surrogate provides a cross-validated screen across only 25 structurally curated perovskite-type energetic materials (PEMs) with experimentally derived Kamlet--Jacobs (K--J) detonation velocities. Representation probes show that the learned descriptors implicitly retain site-aware ionic organization, density information, and coarse packing compatibility, implying why non-periodic clusters can remain predictive before full crystal structures are known. The surrogate extends known PEMs chemistry to three newly synthesized ABX$_4$ materials with both unseen ABX$_4$ stoichiometry and an unseen ethylenediammonium B-site cation, yielding three-point concordance with K--J reference velocities and a mean absolute error (MAE) of 92~m$\cdot$s$^{-1}$ without retraining. Together, these results establish stoichiometry-preserved cluster learning as a synthesis-facing screening strategy for data-scarce multi-ionic materials.

Ming-Yu Guo, W. Zou, Yu Shang et al. · 0 citations