Machine learning for predicting the glass transition and melting temperatures of polymers: molecular representations, model architectures, and interpretability
Abstract
Glass transition (Tg) and melting (Tm) temperatures set the processing window, service range, and end-use performance of polymers, making their rapid prediction central to accelerated materials design. Experiments and simulations are accurate but costly, while classical structure–property models depend on hand-crafted descriptors; machine learning instead maps chemical structure to thermal transitions end to end. This review organizes machine-learning prediction of Tg and Tm around three pillars—molecular representation, model architecture, and interpretability—and foregrounds the features that separate polymer informatics from generic molecular machine learning: repeat-unit periodicity, chain-length-invariant encoding, copolymer sequence, and stereoregularity. We compare descriptors, fingerprints, graph neural networks, and coarse-grained schemes; traditional, deep, transfer, and multi-task models under data scarcity; and methods for quantifying predictive uncertainty and delimiting applicability domains. We treat Tg and Tm as physically distinct targets, since Tm additionally reflects crystal packing, hydrogen bonding, and chain symmetry. A recurring theme is label quality: calorimetric, dynamic-mechanical, and thermomechanical measurements define these transitions differently, so pooled datasets embed instrumental as well as chemical variance. We critically assess explainable artificial intelligence methods and the way accuracy is reported, arguing that headline metrics are not comparable across studies, and we examine how the first community-scale prediction challenge, chemistry-aware data splitting, and calibrated uncertainty can place reporting on a common footing. Finally, we connect representations, architectures, and interpretability to high-performance and sustainable polymer design, synthesizability-aware screening, and closed-loop discovery, and outline open challenges in data scarcity, domain transfer, and chemical-space extrapolation.