Aug 2026· Journal of Chemical Physics· Vol 165 7· 0 citations· 52 references
Medicine
Abstract
Melting point (MP) is an important thermophysical property for the chemical process industry, yet accurate prediction of MP for organic compounds in the absence of experimental data remains challenging due to the complex interplay between molecular packing, intermolecular interactions, and electronic structure. Traditional group contribution and quantitative structure-property relationship models, which rely primarily on static molecular descriptors, often fail to capture these critical condensed-phase effects. In this study, we present a hybrid machine learning framework that integrates cheminformatics descriptors with quantum chemical features and dynamic condensed-phase descriptors derived from molecular dynamics (MD) simulations. Using a curated subset of the DIPPR 801 database, multiple machine learning architectures, including light gradient boosting machine (LightGBM) and graph convolutional networks, were evaluated with feature sets of increasing physical fidelity. The best-performing model, based on LightGBM trained on Dragon descriptors augmented with MD and quantum chemical features, achieves a mean absolute error of 22.5 K, outperforming descriptor-only models and structure-based deep learning baselines. Shapley additive explanations interpretability analysis reveals that melting behavior is governed primarily by molecular topology, surface-area-weighted electronic descriptors, and condensed-phase interaction properties. In contrast, many isolated functional group and single molecule electronic descriptors contribute negligibly once these effects are accounted for. These results demonstrate that incorporating physics-informed, multi-scale descriptors enables more accurate and physically interpretable MP predictions.
The glass transition temperature (Tg) of polyimides is a critical parameter determining their processability and application performance. Traditional experimental methods for measuring Tg are time‐consuming and costly, while existing machine learning prediction models predominantly rely on manually defined molecular descriptors, which often fail to fully capture detailed molecular structural information, limiting their prediction accuracy and generalization capability. To address this, this study proposes a hybrid feature engineering strategy combining Morgan fingerprints and molecular descriptors to comprehensively represent the chemical structure of polyimides. Based on a dataset of 1257 polyimide samples from a public database, we systematically compared six feature selection methods and employed multiple mainstream machine learning algorithms for modeling. The results show that the CATB model performed best, achieving a coefficient of determination (R2) of 0.882 and a mean absolute error (MAE) of 17.34 °C on an independent test set, with fivefold cross‐validation further confirming the model's robustness. SHAP interpretability analysis revealed the significant influence of key features such as the number of rotatable bonds, ether bonds, and ether‐linked oxyethylene units on Tg, providing clear guidance for molecular design. External validation demonstrated the model's strong generalization ability. This study not only achieves high‐precision and robust Tg prediction but also highlights the importance of hybrid feature strategies in polymer property modeling, offering a data‐driven foundation for the rational design of polyimides.
Peishuai Xing, Xiaodong Guo, Yang Wang et al.· Molecular Informatics· 0 citations
Accurate prediction of melting points for pure molecules remains a significant challenge in predictive chemistry, with implications across various scientific fields, including materials science, drug discovery, and separations chemistry. Traditional methods, such as group contribution (GC) techniques, have shown limited success due to the complex relationship between molecular structure and melting point. In this study, we present a data-driven machine learning (ML) approach to predict the melting points of organic compounds, leveraging both 2D and 3D molecular descriptors. Our results indicate that ML models can significantly improve melting-point predictions, providing a robust tool for the scientific community. Scientific contributionOur detailed analysis on melting point prediction, along with SHAP explainability, reveals the top influencing features for the prediction. The P2MAT application we developed as part of this study can predict both melting and boiling points from a SMILES string. P2MAT is available as an easy-to-install, user-friendly GUI for maximum outreach to the scientific community. Our benchmark analysis demonstrates the excellence of our method for predicting melting points.
Md Kamruzzaman, Alexander Landera, N. Menon et al.· Journal of Cheminformatics· 1 citation
Metal nitrides exhibit broad prospects as high-energy-density materials (HEDMs) due to their exceptional energy release characteristics and environmentally friendly decomposition products. However, a fundamental challenge for this class of HEDMs lies in the trade-off between high energy density and high stability. Here, we employed a crystal structure search method to obtain numerous metal nitride configurations with different elements and stoichiometries and characterized their properties using first-principles calculations and
ab initio
molecular dynamics (AIMD) simulations. Based on the dataset, we developed descriptors related to the elemental and structural properties and trained regression models to predict the energy density of the metal nitrides. Additionally, multiple classifier models were trained to assess their stability. Through these machine learning models, we analyzed the features affecting the energy density and stability of metal nitrides and identified three key descriptors: the average distance between each nitrogen atom and its nearest neighbor (DNN), the stoichiometric ratio (N:M), and the average number of N-N bonds per nitrogen atom (aveNNBonds). Based on these insights, we propose a design principle for advanced metal nitride HEDMs: prioritizing high nitrogen-to-metal ratios, light metal elements, and structures wherein nitrogen atoms are spatially separated by the metal matrix, which minimizes N-N bonds and favors dominant M-N bonding. Moreover, our analysis revealed a class of metal nitrides with high nitrogen content that combines considerable stability and high energy density, demonstrating the power of machine learning in material HEDM design.
Yaozhong Liu, Huifang Du, Caimu Wang et al.· Chinese Physics B· 0 citations
Aggregation‐induced emission (AIE) has revolutionized the design of photoluminescent materials by enabling strong solid‐state emission from molecularly nonemissive compounds. However, rational prediction of AIE properties remains challenging because photophysical behavior depends not only on molecular structure but also on aggregate‐state packing and measurement conditions. This study develops a quantitative and interpretable machine learning (ML) framework for predicting experimentally reported emission energies of AIE‐active molecules using continuous physicochemical descriptors derived from molecular structures. A dataset of 590 AIE luminogens—including conjugated organics, donor–acceptor (D–A) systems, silicon‐containing luminogens, and transition‐metal complexes—was analyzed using Gaussian process regression (GPR) combined with SHapley Additive exPlanations (SHAP). The optimized descriptor‐based model achieved moderate predictive performance (test
R
2
= 0.58) and provided chemically interpretable structure–property trends. Feature attribution indicated that nitrogen‐ and sulfur‐containing motifs, electrotopological‐state descriptors, Burden–CAS–University of Texas (BCUT) descriptors, and stereodefined vinylene units are statistically associated with lower emission energies within the present dataset. Morgan fingerprint baseline models showed higher random‐split accuracy, whereas leave‐one‐cluster‐out validation revealed cluster‐dependent degradation for structurally separated regions. This work therefore provides an interpretable initial screening strategy for AIE luminogens while clarifying the need for future models incorporating measurement conditions, solid‐state structural descriptors, and electronic‐structure‐informed features.