Graph-based models provide the lowest errors in property prediction, retain their advantage across the evaluated training-set sizes, and remain robust to increasing repeat-unit complexity, according to the PolyBench26 benchmark.
Abstract
Polymer property prediction lacks open, standardized benchmarks that enable rigorous comparison of machine-learning methods, with existing resources covering only a narrow fraction of polymer architectures, such as homopolymers. We introduce Polymer Benchmark 2026 (PolyBench26), an open dataset comprising nearly 250,000 polymer-property datapoints across eight physical properties, including data from experimental measurements, density functional theory, and molecular dynamics. The benchmark supports four evaluation tasks across homopolymers and alternating, random, and block copolymers: in-distribution property prediction, dataset-size scaling, repeat-unit complexity, and transfer to held-out polymer architectures. We compare language model, graph-based, and descriptor-based approaches and find graph-based models provide the lowest errors in property prediction, retain their advantage across the evaluated training-set sizes, and remain robust to increasing repeat-unit complexity. PolyBench26 provides a reproducible foundation for developing models for the increasingly complex polymer design space. The PolyBench26 benchmark is available open-source at https://github.com/rlearsch/PolymerBenchmark2026.
Polymer informatics currently lacks shared benchmarks that enable reproducible comparison of property prediction methods across datasets, representations, and models. Here, we present PolyBench, a public benchmark platform comprising 39 literature-derived prediction tasks with 42,169 curated structure–property record...
Polymers are fundamental to modern materials science because their backbone chemistry, monomer composition, and chain architecture can be systematically tuned to achieve a virtually unlimited range of mechanical, thermal, electronic, and optical properties. To accelerate the discovery and design of polymeric material...
Amberbir Alemayoh, Zhi-Xin Pan, Ning Wang· Polymer Science & Techno...· 0 citations
Polymers are highly versatile and cost-effective materials with a wide range of tunable physicochemical properties, making them indispensable across numerous technological and biomedical applications. However, their intrinsic chemical heterogeneity and complex structural organization pose significant challenges for t...
Lorand Gabriel Parajdi, I. Tóth, Alex-Adrian Farcas et al.· Scientific Reports· 0 citations
Polymer properties emerge from interactions across scales, yet existing polymer models typically preserve either detailed monomer chemistry without an explicit polymer graph or polymer connectivity with simplified monomer representations, due to computational constraints, as polymers typically contain tens of thousands...
This study explores data-driven approaches for predicting the thermal decomposition temperature of polymers using both classical machine learning (CML) models and a small language model (SLM), suggesting that small language models can serve as a valuable alternative modeling strategy for predicting polymer thermal prop...
Nguyen T. T. Duyen, Ngo T. Que, Hanh Bich Vu· Journal of Physics, Conferen...· 0 citations
An interpretable machine learning workflow is developed that predicts both the total dielectric constant and the HSE bandgap of polymer repeat units directly from a monomer structure, using an open density-functional-theory dataset of 284 four-block polymers.
F. Gül· Polymers· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 29, 2026
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.