PREreview of "Benchmark Label Definition, Not Model Capability, Explains a Machine-Learning Advantage over TD-DFT for λmax Prediction"
Abstract
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/22983291. General assessment This manuscript addresses an important and underappreciated problem in computational chemistry and molecular machine learning: apparent differences in predictive performance can arise from benchmark construction and target definition rather than from intrinsic superiority of one modelling approach over another. The distinction between literature-reported absorption maxima and calculated vertical excitation energies is scientifically important, and the emphasis on genuinely external validation is a clear strength of the work. However, I find the central conclusion stronger than the evidence presently described appears to support. The results convincingly show that model performance changes substantially across datasets and label conventions. They do not yet demonstrate that label definition, rather than model capability, is the principal cause of the apparent machine-learning advantage over TD-DFT. Several factors change simultaneously across the comparisons presented in the manuscript, including chemical domain, dataset provenance, transition assignment, computational protocol, sample size, and the nature of the machine-learning representation. These factors make it difficult to isolate label definition as the causal explanation. The manuscript raises a valuable methodological concern, but the title and several conclusions should either be moderated or supported by additional controlled analyses. Major comments 1. The title makes a stronger causal claim than the study design appears to establish The title states that benchmark label definition, "not model capability," explains the machine-learning advantage over TD-DFT. This is a strong exclusionary claim. The results show that the apparent ranking between a Gaussian-process model and quantum-chemical calculations changes between evaluation datasets. However, label definition is not the only property that changes between these datasets. Their chemical composition, source literature, distribution of chromophores, spectral range, experimental conditions, and possibly computational protocols also differ. Therefore, demonstrating that a GP performs differently on a dataset with transition-specific labels does not by itself establish that label definition is the cause of the difference. A substantially stronger test would use the same set of molecules with two alternative target definitions—for example, a literature-reported overall λmax and an independently assigned transition-specific λmax—and evaluate both methods against both labels. A paired analysis of this type would isolate label definition while holding molecular distribution constant. Unless such a controlled analysis is provided, I suggest moderating the central conclusion to state that label definition and dataset provenance are important contributors to the reported ML advantage rather than demonstrating that they fully explain it. 2. One Gaussian-process model cannot represent "machine-learning capability" generally The machine-learning side of the comparison is represented by a Gaussian process using 256-bit Morgan fingerprints. This is a relatively specific and comparatively simple model representation. The conclusions therefore apply directly to this GP/fingerprint model, but not necessarily to molecular machine learning more broadly. Modern λmax prediction methods may use larger fingerprints, graph neural networks, message-passing architectures, pretrained molecular representations, three-dimensional information, explicit solvent descriptors, or quantum-chemical features. Such models may respond differently to both chemical and label distribution shifts. This issue is particularly important because the title contrasts "model capability" with benchmark effects in general terms. I recommend evaluating several substantially different ML approaches or restricting the conclusions explicitly to the model class studied here. The use of only 256 fingerprint bits also deserves justification. Such a short fingerprint may produce substantial bit collisions for chemically diverse datasets. Performance should ideally be compared with conventional 1024- or 2048-bit Morgan fingerprints to establish that the observed transfer behaviour is not partly a consequence of the selected molecular representation. 3. The TD-DFT comparison may not constitute a uniform computational benchmark The quantum-chemical reference data appear to originate from multiple published sources and include TD-DFT, sTDA, ωB97X-D3, and CAM-B3LYP/6-31G** calculations. These are not interchangeable computational protocols. Differences in functional, basis set, molecular geometry, conformer selection, solvent model, treatment of protonation or tautomerism, and transition assignment can all substantially affect calculated excitation energies. Consequently, describing the comparison simply as "machine learning versus TD-DFT" risks combining methodological heterogeneity with model-class differences. A stronger comparison would apply a single, clearly defined quantum-chemical workflow to the same molecules used for ML evaluation. At minimum, performance should be stratified by computational protocol rather than interpreted as representing a single quantum-chemical method. 4. The very small TD-DFT comparison set limits the strength of the in-corpus conclusion The reported in-corpus TD-DFT comparison contains only 36 compounds. An RMSE difference of 87.3 versus 108.7 nm may appear substantial, but with n = 36 the uncertainty around both values could also be substantial, particularly if a small number of molecules have very large errors. Confidence intervals or bootstrap distributions for the difference in RMSE should therefore be reported. More generally, model comparisons should be paired whenever predictions are available for the same molecules. Reporting distributions of paired absolute-error differences would be more informative than comparing headline RMSE values alone. The larger sTDA comparison is more convincing statistically, but it answers a somewhat different computational question and should not simply be treated as confirming the TD-DFT result. 5. Evaluation in wavelength units may distort the comparison A particularly important issue is the use of RMSE in nanometres. Excitation energy and wavelength are related non-linearly: E ∝ 1/λ. Therefore, a 50 nm error at 250 nm does not represent the same energetic error as a 50 nm error at 650 nm. RMSE measured in nm gives disproportionate influence to errors occurring at longer wavelengths and complicates comparisons across datasets with different spectral ranges. This is especially relevant here because different external datasets may contain substantially different wavelength distributions. The main comparisons should therefore also be reported in energy units, preferably eV or cm⁻¹. Conclusions that depend strongly on whether errors are expressed in nm or eV would require careful interpretation. The same consideration applies to the reported 57–65 nm systematic offsets. A constant wavelength offset does not correspond to a constant error in excitation energy across the spectrum. 6. Removing the systematic offset requires careful validation One of the key claims is that quantum-chemical calculations become substantially more accurate than the GP after each method's "removable offset" is discounted. The procedure used to estimate and remove this offset is therefore crucial. If the offset is estimated from the same test dataset on which residual error is subsequently reported, the corrected error is optimistically biased. This effectively gives the method access to information from the evaluation set. A fair calibration procedure would estimate the correction on an independent calibration or training dataset and then apply that fixed correction to an untouched external test set. The manuscript should explicitly state how offsets were estimated and whether test labels influenced the correction. Both uncorrected and prospectively calibrated results should be reported. Otherwise, comparing a calibrated quantum-chemical method against an uncalibrated GP may not constitute a balanced comparison. 7. Reproduction of a similar quantum-chemical offset does not establish its origin The observation that an approximately 57–65 nm underestimation in the in-corpus data reappears as approximately −61.9 nm in an external dataset is interesting. However, the conclusion that this supports a "definitional rather than functional origin" seems too strong. A reproducible systematic bias could also arise from features shared across quantum-chemical calculations, including the use of vertical excitations, similar density functionals, incomplete treatment of solvent effects, neglect of vibronic structure, conformational differences, or geometry approximations. The recurrence of the offset establishes reproducibility of the bias, but not necessarily its mechanistic origin. To distinguish label-definition effects from computational bias more convincingly, the manuscript would need controlled comparisons involving multiple electronic-structure methods and experimentally assigned transitions on the same molecules. 8. Label convention and chemical-domain shift are incompletely separated The external photoswitch dataset is an important test because it uses a specifically defined E-isomer π→π* transition. However, photoswitches are also a specialised chemical domain. Therefore, deterioration in GP performance may simultaneously reflect changes in molecular structure, photochemical behaviour, chromophore families, wavelength distribution, and target definition. The statement that "chemical novelty alone does not break it, reassigning the label does"