Aug 2026· Journal of Molecular Graphics and Modelling· Vol 148, pp.
109542
· 0 citations· 63 references
Medicine
Abstract
Experimental determination of flash points (FPs) for liquid mixtures is laborious and costly, highlighting the need for reliable predictive approaches for safety assessment and engineering applications. Although numerous models have been reported for binary miscible mixtures, most rely on fixed model parameters or empirical correlations, which limits their ability to capture the nonlinear relationship between molecular structure and FP. In this study, a quantitative structure-property relationship (QSPR) framework that tightly integrates differential evolution (DE) with support vector regression (SVR) was developed to predict the FP values of binary miscible mixtures, where DE was employed to globally optimize key SVR hyperparameters and enhance model generalization capability. A dataset consisting of 332 compositions from 33 binary mixtures formed by pairwise combinations of 20 pure components was employed, and multiple molecular descriptor representation strategies were deliberately adopted to construct distinct DE-SVR models, enabling a systematic investigation of the combined effects of descriptor representation and model optimization on predictive performance. Three DE-SVR models were established based on different descriptor sets, and their predictive accuracy, robustness, and stability were comprehensively evaluated. Among them, the model constructed using physicochemical parameters exhibited the best overall performance. Comparative analyses with existing FP prediction methods reported in the literature further confirmed the effectiveness and superiority of the proposed DE-SVR-based models. The results of this study provide a practical tool for FP estimation of binary mixtures and offer valuable insights into the joint roles of model optimization and molecular representation in mixture property prediction.
An additive fragment-based QSPR model for molar refractivity (MR) was developed using Monte Carlo optimization within the CORAL framework. Molecular structure was encoded using fragment-level attributes, and their contributions were statistically optimized across multiple independent training-validation splits. The resulting linear models retain the additive character of classical fragment-constant approaches while introducing data-driven weight optimization under explicit statistical control. Across three independent splits, the developed QSPR models exhibit consistently high coefficients of determination in both the training and external validation sets, with validation R2 values ranging from approximately 0.86 to 0.89. Root-mean-square errors remained consistent across splits, indicating reproducible behavior of the optimized descriptors. An applicability domain (AD) was defined using statistical defect metrics derived from descriptor distributions. Across independent splits, the applicability domain consistently identified a minority of compounds with rare or unevenly represented SMILES attributes; these statistically under-supported compounds were associated with a modest increase in prediction error, indicating that the defect-based domain reflects descriptor representativeness rather than acting as a strict error filter. The study presents Monte Carlo-optimized additive modeling as a statistically audited extension of fragment-constant schemes, integrating robustness analysis and applicability-domain assessment into a transparent QSPR workflow.
Aleksandar M. Veselinović, Jelena Zivkovic, S. Sunarić et al.· Journal of Molecular Graphic...· 0 citations
Melting point (MP) is an important thermophysical property for the chemical process industry, yet accurate prediction of MP for organic compounds in the absence of experimental data remains challenging due to the complex interplay between molecular packing, intermolecular interactions, and electronic structure. Traditional group contribution and quantitative structure-property relationship models, which rely primarily on static molecular descriptors, often fail to capture these critical condensed-phase effects. In this study, we present a hybrid machine learning framework that integrates cheminformatics descriptors with quantum chemical features and dynamic condensed-phase descriptors derived from molecular dynamics (MD) simulations. Using a curated subset of the DIPPR 801 database, multiple machine learning architectures, including light gradient boosting machine (LightGBM) and graph convolutional networks, were evaluated with feature sets of increasing physical fidelity. The best-performing model, based on LightGBM trained on Dragon descriptors augmented with MD and quantum chemical features, achieves a mean absolute error of 22.5 K, outperforming descriptor-only models and structure-based deep learning baselines. Shapley additive explanations interpretability analysis reveals that melting behavior is governed primarily by molecular topology, surface-area-weighted electronic descriptors, and condensed-phase interaction properties. In contrast, many isolated functional group and single molecule electronic descriptors contribute negligibly once these effects are accounted for. These results demonstrate that incorporating physics-informed, multi-scale descriptors enables more accurate and physically interpretable MP predictions.
Frank T. Mtetwa, N. Giles, W. Wilding et al.· Journal of Chemical Physics· 0 citations
This study provides a structured and reproducible assessment of the conditions under which descriptor-based ML models can be expected to succeed or fail in DES systems and highlights the importance of rigorous, leakage-aware evaluation in data-driven chemical modeling.
Hakim Faraji, Julio Brito Santana, R. Rodríguez-Ramos et al.· ACS Omega· 0 citations
Deep eutectic solvents (DESs) have considerable potential for NH
3
capture, but traditional solvent screening methods are unable to identify appropriate DES efficiently. One thousand nine hundred fifty‐nine experimental solubility data points for 72 DESs were used to construct and compare multiple machine learning models based on σ‐profile descriptors to predict NH
3
solubility in DESs. CatBoost achieved the best performance (
R
2
= 0.993, RMSE = 0.079). Nested cross‐validation and independent test sets confirmed that the model has good physical consistency and cross‐system generalization. SHAP analysis further quantified the contributions of key features. The final model was then employed to predict the NH
3
solubilities of 1140 DESs, and the highest‐ranked systems were selected for subsequent characterization and absorption experiments. The results of the gas absorption performance experiment are in excellent agreement with the predictions of the model. Finally, quantum chemical calculations were used to clarify the microscopic mechanisms underlying DES formation and NH
3
interaction.
Lu Gao, Ruixin Li, Lili Wang et al.· AIChE Journal· 0 citations
Accurate prediction of physicochemical properties such as the octanol-water partition coefficient’s logarithm (LogP) is critical in early-stage drug development. This study presents a novel and interpretable computational framework that integrates symbolic graph-theoretical descriptors, specifically topological indices derived from M-polynomials, with machine learning (ML) techniques for LogP prediction in oncology drug molecules. Molecular graphs were constructed from SMILES strings, and eight M-polynomial-based topological indices were computed as predictive features. A comprehensive suite of regression models was applied, ranging from univariate and multivariate linear regressions to regularized methods (least absolute shrinkage and selection operator (LASSO), ridge, ElasticNet) and nonlinear ensemble learners (random forest, XGBoost). Dimensionality reduction using principal component analysis (PCA) and rigorous validation via cross-validation and bootstrapping were conducted. Among all approaches, ensemble models combining XGBoost and random forest with bootstrapping yielded the most robust performance, achieving an R2 score greater than 0.6 and a low prediction error. These results demonstrate the effectiveness of M-polynomial-derived indices as interpretable molecular descriptors and affirm the predictive utility of advanced ML models in quantitative structure-property relationship (QSPR) modeling. To our knowledge, this is among the first comprehensive studies integrating ensemble-based ML methods and M-polynomial-derived topological indices for LogP prediction in oncology drugs.
Shabbir Ahmad, Sana Javed, S. Khalid et al.· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.