Skip to content
Open access

Glass Transition Prediction of Binary Copolymers Across Large Chemical Spaces Using Machine Learning and Physics-Based Modeling

Jul 2026 · Polymers · Vol 18 · 0 citations · 36 references
Medicine

TL;DR

A machine learning framework for the high-throughput prediction of Tg in binary copolymers, trained on experimental datasets encompassing both homopolymers and copolymers, and validated using physics-based molecular dynamics simulations.

Abstract

The glass transition temperature (Tg) is a pivotal design parameter for polymer performance across diverse applications, yet its rapid prediction within expansive chemical spaces remains a challenge. We present a machine learning (ML) framework for the high-throughput prediction of Tg in binary copolymers, trained on experimental datasets encompassing both homopolymers and copolymers. We evaluate various ML architectures, including graph-based algorithms, to effectively capture non-linear composition–property relationships. The optimized model achieves high predictive accuracy with an RMSE of ~14K and an R2 of ~0.98. Crucially, the framework accounts for the chemical diversity of monomeric units by integrating structural descriptors with molar composition ratios, enabling the model to capture complex dependencies of thermal stability on chemical structure and composition. We validate the model’s robustness using physics-based molecular dynamics (MD) simulations. To showcase the platform’s scalability, we generated a library of approximately 148,000 binary copolymer compositions and predicted their Tg, facilitating the rapid mapping of vast design spaces. This extensive virtual library enables the identification of optimal monomer pairings that would be experimentally inaccessible through traditional trial-and-error methods. Through these large-scale exploration studies, we demonstrate the ability to design copolymers for targeted applications, including a specific case study on elastomeric systems. This integrated approach, combining experimental data, ML modeling, and physics-based validation, offers a transformative path for the accelerated discovery and multi-property optimization of functional copolymers.

Read PDF

Similar papers

Open access Aug 2026

Boosting the Prediction Accuracy of Glass Transition Temperature in Polyimides: A Hybrid Machine Learning Approach Integrating Morgan Fingerprints and Molecular Descriptors

The glass transition temperature (Tg) of polyimides is a critical parameter determining their processability and application performance. Traditional experimental methods for measuring Tg are time‐consuming and costly, while existing machine learning prediction models predominantly rely on manually defined molecular descriptors, which often fail to fully capture detailed molecular structural information, limiting their prediction accuracy and generalization capability. To address this, this study proposes a hybrid feature engineering strategy combining Morgan fingerprints and molecular descriptors to comprehensively represent the chemical structure of polyimides. Based on a dataset of 1257 polyimide samples from a public database, we systematically compared six feature selection methods and employed multiple mainstream machine learning algorithms for modeling. The results show that the CATB model performed best, achieving a coefficient of determination (R2) of 0.882 and a mean absolute error (MAE) of 17.34 °C on an independent test set, with fivefold cross‐validation further confirming the model's robustness. SHAP interpretability analysis revealed the significant influence of key features such as the number of rotatable bonds, ether bonds, and ether‐linked oxyethylene units on Tg, providing clear guidance for molecular design. External validation demonstrated the model's strong generalization ability. This study not only achieves high‐precision and robust Tg prediction but also highlights the importance of hybrid feature strategies in polymer property modeling, offering a data‐driven foundation for the rational design of polyimides.

Peishuai Xing, Xiaodong Guo, Yang Wang et al. · 0 citations
Preprint Aug 2026

3D Molecular Representation Learning for Organic Mixtures: Viscosity and Density Prediction

The viscosity and density of organic mixtures are essential properties for designing lubricants, solvents, and heat transfer fluids. In engineering practice, formulating a functional fluid requires understanding how these properties change with composition and temperature. However, exhaustive experimental characterization across the full parameter space is impractical due to the vast number of possible species and combinations. Here we introduce a mixture-aware 3D molecular representation learning strategy, built upon a pre-trained molecular encoder, that jointly encodes component structures, mole fractions, and temperature to achieve accurate predictions for organic mixtures. Fine-tuning on publicly available datasets covering a wide range of binary organic mixtures yields test-set R2 values of 0.973 for dynamic viscosity and 0.996 for density, significantly outperforming traditional machine learning baselines. Beyond this overall accuracy, the model captures non-monotonic viscosity changes upon mixing, surpassing simple linear or logarithmic mixing rules. The architecture is extendable to ternary and multicomponent mixtures, as verified via preliminary experiments. Using this model, we quantitatively analyze how molecular structure-branching, cycloalkane, and aromatic rings-affects viscosity-temperature behavior, which benefits the design of lubricants with superior viscosity-temperature performance. Altogether, this work provides a practical, data-driven tool for mixture property prediction, accelerating the rational formulation of functional fluids in chemical engineering.

H. Qu, Yanyi Su, Ning Wang et al. · 0 citations
Aug 2026

Prediction of Glass Transition Temperatures for Organic Electronic Materials

The glass transition temperature (Tg) is a critical descriptor governing the morphological stability, emitter orientation, and interfacial integrity of amorphous thin films in organic electronics. However, experimental Tg measurements suffer from high resource costs and interlaboratory variability, while machine learning models are bottlenecked by scarce, noisy data sets. Here, we establish a physics-based atomistic molecular dynamics (MD) protocol to predict the Tg of 160 diverse organic electronic materials. To study computational throughput and predictive accuracy, we systematically benchmarked nine configurations spanning system sizes (5,000, 10,000, and 15,000 atoms) and cooling step relaxation times (5, 10, and 15 ns). Extracted via an automated, bias-free hyperbolic fitting scheme, our preferred standalone workflow (15,000 atoms, 15 ns) yields a correlation of R2 = 0.89 and a mean absolute error (MAE) of 10.9 K relative to experiment. Structural descriptor analysis confirms that accuracy remains uniform regardless of molecular weight or heteroatom density, establishing this transferable workflow as a digital sieve to accelerate the discovery of next-generation organic electronics.

Hadi Abroshan, Paul Winget, H. Kwak et al. · 0 citations
Aug 2026

Machine learning prediction of organic compound melting points informed by condensed-phase and electronic descriptors.

Melting point (MP) is an important thermophysical property for the chemical process industry, yet accurate prediction of MP for organic compounds in the absence of experimental data remains challenging due to the complex interplay between molecular packing, intermolecular interactions, and electronic structure. Traditional group contribution and quantitative structure-property relationship models, which rely primarily on static molecular descriptors, often fail to capture these critical condensed-phase effects. In this study, we present a hybrid machine learning framework that integrates cheminformatics descriptors with quantum chemical features and dynamic condensed-phase descriptors derived from molecular dynamics (MD) simulations. Using a curated subset of the DIPPR 801 database, multiple machine learning architectures, including light gradient boosting machine (LightGBM) and graph convolutional networks, were evaluated with feature sets of increasing physical fidelity. The best-performing model, based on LightGBM trained on Dragon descriptors augmented with MD and quantum chemical features, achieves a mean absolute error of 22.5 K, outperforming descriptor-only models and structure-based deep learning baselines. Shapley additive explanations interpretability analysis reveals that melting behavior is governed primarily by molecular topology, surface-area-weighted electronic descriptors, and condensed-phase interaction properties. In contrast, many isolated functional group and single molecule electronic descriptors contribute negligibly once these effects are accounted for. These results demonstrate that incorporating physics-informed, multi-scale descriptors enables more accurate and physically interpretable MP predictions.

Frank T. Mtetwa, N. Giles, W. Wilding et al. · 0 citations
Preprint Jul 2026

Informatics Modeling of High Tg Polymers: Assessing the Role of Processing versus Chemistry

Despite the advances in structure-based modeling of polymer properties, accurately predicting glass transition temperature (Tg) is still challenging for polymers whose behavior is strongly influenced by intermolecular interactions and processing conditions. We previously developed a machine-learning model based on polymer topological descriptors to predict Tg. The model performed well and was based solely on the chemistry and structure of the polymer without any inclusion of processing parameters. In this work, we have extended that work by first applying that same model to a larger range of polymers and second by integrating processing parameters into the feature set. The chemistry-based model still demonstrates consistent predictive performance for most polymers, indicating that Tg is indeed primarily chemistry and structure driven and not strongly impacted by processing. However, several polymers exhibited deviations between predicted and experimental Tg values. Detailed analysis reveals that these differences are related to strong intermolecular interactions and processing-dependent factors, particularly for polymers prepared by solution casting and high temperature annealing. These results demonstrate that molecular topology provides a strong foundation for Tg prediction; however, this approach also screens out those classes of polymers for with processing conditions play an important role.

Qinrui Liu, Scott R. Broderick · 0 citations
Open access Aug 2026

A Leakage-Aware Benchmark Study of Machine Learning Models for Deep Eutectic Solvent Property Prediction

Predicting the physicochemical properties of deep eutectic solvents (DESs) remains challenging due to the large combinatorial design space and the complex, composition- and temperature-dependent interactions governing their behavior. While machine learning (ML) has been widely applied to DES property prediction, reported performance is often sensitive to data set structure, feature representation, and validation design, raising questions about the reliability and transferability of existing models. In this work, we present a systematic evaluation of descriptor-based ML models for DES property prediction using a curated and high-confidence data set spanning 2003–2026. A unified feature representation combining molecular descriptors, molar composition, and temperature is employed, and model performance is assessed under a hierarchy of validation protocols designed to control for data leakage and progressively increase extrapolation difficulty. The results show that predictive performance is strongly dependent on the validation design. Density and refractive index exhibit relatively stable behavior, while surface tension shows moderate predictability. In contrast, electrical conductivity and viscosity display a strong dependence on temperature and limited contribution from descriptor-based features. Under extrapolative validation, performance for these properties deteriorates substantially, indicating limited transferability. Additional analyses based on feature-space distance and similarity-based baselines indicate that prediction errors are strongly associated with training-domain proximity and that local similarity in the current feature space is insufficient for reliable prediction outside observed data regions. These findings suggest that the primary limitation arises from the representational capacity of static descriptor-based features rather than model choice alone. Overall, this study provides a structured and reproducible assessment of the conditions under which descriptor-based ML models can be expected to succeed or fail in DES systems and highlights the importance of rigorous, leakage-aware evaluation in data-driven chemical modeling.

Hakim Faraji, Julio Brito Santana, R. Rodríguez-Ramos et al. · 0 citations