Predicting thermal stability during handling and storage is essential for the design of safe and reliable energetic materials. However, experimental measurements vary significantly across laboratories due to differences in protocols and analysis methods, making it difficult to train reliable predictive models. We address this challenge through differential learning. Rather than predicting absolute decomposition temperatures, we instead train message passing neural networks to predict relative differences between pairs of molecules. This approach reduces sensitivity to systematic experimental errors and achieves>85% accuracy in ranking compounds by thermal stability, outperforming conventional regression methods on the same heterogeneous dataset. To understand what drives these predictions, we compare neural network models with interpretable alternatives built from descriptors derived from ab initio calculations and cheminformatics software. This analysis identifies bond dissociation enthalpy as a key determinant of thermal stability rankings, providing further insight into the complex chemistry of thermal decomposition. The differential learning framework generalizes across model architectures, from graph neural networks to classical descriptor-based approaches. Our results demonstrate that learning relative properties rather than absolute values offers a practical solution for modeling noisy experimental data, with direct applications in materials design where thermal stability predictions inform safety protocols.
It is argued that developing thermodynamics-informed ML constitutes one of the most important and least explored frontiers in materials discovery and that the next generation of ML models must move beyond static energy predictions towards a thermodynamic description of materials behaviour under realistic operating conditions.
Pol Benítez, Cibr'an L'opez, Claudio Cazorla· 0 citations
While feature importance analysis in chemical and materials machine learning can often be sensitive to both the predictive model and the attribution rule, the robustness of these rankings is rarely quantified before they are used to gain physical insights. Here, we compare 26 feature importance pipelines spanning data-driven, model-based, and formula-based analyses on a metal-support interaction data set anchored by an explicit SISSO equation, and we examine whether the same qualitative behavior recurs in high-entropy-alloy and halide perovskite data sets. Across the three benchmarks, we observe high intrafamily agreement but substantial interfamily variance. While a small subset of features remains stable across multiple families, several midranked features are highly family dependent, with their apparent importance shifting according to the underlying modeling assumptions. To ensure robust interpretability, we recommend that feature importance be reported by method family or correlation-based clusters, supplemented by resampling intervals.
Ruilin Lai, Xiaotong Liu, Yuhang Wang et al.· Journal of Chemical Informat...· 0 citations
Reactivity ratios are a key metric for understanding copolymer microstructure, yet they are challenging to predict a priori. Data-driven methods have recently been employed for the prediction of free radical copolymerization reactivity ratios, but model extrapolation remains modest with respect to accuracy. This has elicited discussions on variability within published reactivity ratio datasets and a call for standardization. Here, we seek to identify the source of accuracy limitation in data-driven reactivity ratio prediction through the addition of new model features, evaluation of the model strategy, and estimate of the quality of literature reactivity ratio data. Systematic studies using a literature-mined dataset of >450 reactivity ratios demonstrated that the inclusion of more relevant transition state features in multivariate linear regression models and the use of more complex machine learning models for reactivity ratios led to marginal improvements in model accuracy, evaluated through mean absolute error. This motivated the evaluation of the accuracy of literature-extracted reactivity ratios through the analysis of a representative library of 100 reactivity ratio values from 47 distinct studies. Independent measurements of the same comonomer pair disagree by an average standard deviation of 0.26 kcal/mol in ΔΔG‡, an estimate of the irreducible noise in the training labels. The error of every model evaluated here, and most literature models, is comparable to this dispersion. Therefore, the study concludes that the accuracy floor in data-driven methods for reactivity ratio prediction is set by the data rather than by the descriptors or architecture and improving predictive models for reactivity ratios will require a large, standardized dataset of reactivity ratios measured using modern synthesis, characterization, and statistical analysis. More generally, the concept presented here of estimating the measurement–noise floor of a literature-mined dataset provides an approach to determine whether a prediction task is limited by the model or by the data, and we propose that it can be generalized beyond reactivity ratios to polymer properties that are compiled from heterogeneous literature.
Caroline M. Coxwell, Dylan M Anstine, O. Isayev et al.· Macromolecules· 0 citations
The Northeast Materials Database is leveraged to develop machine learning models that predict magnetic materials with targeted Curie temperatures from composition-derived descriptors rooted in molecular-level elemental properties, supplemented by a small set of coarse crystal-system and structure-family indicators.
F. Uçar, Nida Katı· Scientific Reports· 0 citations
Thermal decomposition temperature is a crucial factor in evaluating the thermal stability and practicality of polymer materials. In this study, we explore data-driven approaches for predicting the thermal decomposition temperature of polymers using both classical machine learning (CML) models and a small language model (SLM). We use experimental polymer datasets from the PolyInfo database to train Random Forest and XGBoost models, which utilize molecular fingerprints as structured input features. In contrast, the SLM-based approach directly uses polymer expressed as Simplified Molecular Input Line Entry System (SMILES) strings in textual form, eliminating the need for feature engineering or explicit preprocessing of molecular descriptors. This presents an alternative modeling framework for predicting polymer properties, where structure-property relationships are learned directly from raw chemical representations. In addition to thermal decomposition temperature, we also apply this framework to predict glass transition temperature using a dataset previously reported in our work, demonstrating its potential applicability to multiple polymer thermal properties. Overall, our results suggest that small language models can serve as a valuable alternative modeling strategy for predicting polymer thermal properties, providing a complementary perspective to traditional methods.
N. T. T. Duyên, Ngo T. Que, Hanh Bich Vu· Journal of Physics, Conferen...· 0 citations
A global active learning framework is demonstrated to map these landscapes efficiently by coupling genetic algorithms with deep neural networks trained on density functional theory data, which provides a scalable and resource-efficient strategy for high-throughput materials discovery in applications such as hydrogen storage and ammonia synthesis.
Johnathan von der Heyde, Walter Malone, A. Kara· Journal of Physical Chemistr...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.