Skip to content

Author

Caroline M. Coxwell

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Predicting Copolymerization Reactivity Ratios: Do We Need Better Models or Better Data?

Reactivity ratios are a key metric for understanding copolymer microstructure, yet they are challenging to predict a priori. Data-driven methods have recently been employed for the prediction of free radical copolymerization reactivity ratios, but model extrapolation remains modest with respect to accuracy. This has elicited discussions on variability within published reactivity ratio datasets and a call for standardization. Here, we seek to identify the source of accuracy limitation in data-driven reactivity ratio prediction through the addition of new model features, evaluation of the model strategy, and estimate of the quality of literature reactivity ratio data. Systematic studies using a literature-mined dataset of >450 reactivity ratios demonstrated that the inclusion of more relevant transition state features in multivariate linear regression models and the use of more complex machine learning models for reactivity ratios led to marginal improvements in model accuracy, evaluated through mean absolute error. This motivated the evaluation of the accuracy of literature-extracted reactivity ratios through the analysis of a representative library of 100 reactivity ratio values from 47 distinct studies. Independent measurements of the same comonomer pair disagree by an average standard deviation of 0.26 kcal/mol in ΔΔG‡, an estimate of the irreducible noise in the training labels. The error of every model evaluated here, and most literature models, is comparable to this dispersion. Therefore, the study concludes that the accuracy floor in data-driven methods for reactivity ratio prediction is set by the data rather than by the descriptors or architecture and improving predictive models for reactivity ratios will require a large, standardized dataset of reactivity ratios measured using modern synthesis, characterization, and statistical analysis. More generally, the concept presented here of estimating the measurement–noise floor of a literature-mined dataset provides an approach to determine whether a prediction task is limited by the model or by the data, and we propose that it can be generalized beyond reactivity ratios to polymer properties that are compiled from heterogeneous literature.

Caroline M. Coxwell, Dylan M Anstine, O. Isayev et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.