Interpretable machine learning for Curie temperature prediction of magnetic materials: compositional descriptors, shap analysis, and a perovskite case study
The rational design of magnetic materials with targeted Curie temperatures ( $$T_C$$ ) remains a central challenge in materials science. In this study, we leverage the Northeast Materials Database (NEMAD), comprising 33,668 experimentally reported magnetic compounds, to develop machine learning models that predict $$T_C$$ from composition-derived descriptors rooted in molecular-level elemental properties, supplemented by a small set of coarse crystal-system and structure-family indicators. After rigorous data curation—including deduplication of 13,186 unique compounds and extraction of 49 physics-informed features encompassing electronegativity, atomic radii, ionization energy, and valence electron statistics—we benchmark four gradient-boosted and ensemble algorithms. XGBoost achieves the best predictive performance with a five-fold cross-validated $$R^2$$ of 0.79 and a mean absolute error of 70.2 K. SHapley Additive exPlanations (SHAP) analysis reveals that the fractional content of 3 d magnetic elements (Mn, Fe, Co, Ni, Cr) dominates the prediction landscape, contributing an average of 85 K to the SHAP value, followed by the average first ionization energy and minimum atomic number of the constituents. A dedicated deep-dive into 2,724 perovskite-structured compounds demonstrates that the global model generalizes well to this technologically important family ( $$R^2$$ = 0.73). Finally, we screen 20 hypothetical perovskite compositions and identify candidates—notably Sr $$_2$$ FeMoO $$_6$$ and La $$_{0.7}$$ Ba $$_{0.3}$$ MnO $$_3$$ —predicted to exhibit room-temperature ferromagnetism. These findings provide an interpretable, data-driven bridge between molecular-scale electronic descriptors and macroscopic magnetic ordering, offering a practical tool for accelerated discovery of high- $$T_C$$ magnetic materials.