Skip to content

PREreview of "Benchmark Label Definition, Not Model Capability, Explains a Machine-Learning Advantage over TD-DFT for λmax Prediction"

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/22983291. General assessment This manuscript addresses an important and underappreciated problem in computational chemistry and molecular machine learning: apparent differences in predictive performance can arise from benchmark construction and target definition rather than from intrinsic superiority of one modelling approach over another. The distinction between literature-reported absorption maxima and calculated vertical excitation energies is scientifically important, and the emphasis on genuinely external validation is a clear strength of the work. However, I find the central conclusion stronger than the evidence presently described appears to support. The results convincingly show that model performance changes substantially across datasets and label conventions. They do not yet demonstrate that label definition, rather than model capability, is the principal cause of the apparent machine-learning advantage over TD-DFT. Several factors change simultaneously across the comparisons presented in the manuscript, including chemical domain, dataset provenance, transition assignment, computational protocol, sample size, and the nature of the machine-learning representation. These factors make it difficult to isolate label definition as the causal explanation. The manuscript raises a valuable methodological concern, but the title and several conclusions should either be moderated or supported by additional controlled analyses. Major comments 1. The title makes a stronger causal claim than the study design appears to establish The title states that benchmark label definition, "not model capability," explains the machine-learning advantage over TD-DFT. This is a strong exclusionary claim. The results show that the apparent ranking between a Gaussian-process model and quantum-chemical calculations changes between evaluation datasets. However, label definition is not the only property that changes between these datasets. Their chemical composition, source literature, distribution of chromophores, spectral range, experimental conditions, and possibly computational protocols also differ. Therefore, demonstrating that a GP performs differently on a dataset with transition-specific labels does not by itself establish that label definition is the cause of the difference. A substantially stronger test would use the same set of molecules with two alternative target definitions—for example, a literature-reported overall λmax and an independently assigned transition-specific λmax—and evaluate both methods against both labels. A paired analysis of this type would isolate label definition while holding molecular distribution constant. Unless such a controlled analysis is provided, I suggest moderating the central conclusion to state that label definition and dataset provenance are important contributors to the reported ML advantage rather than demonstrating that they fully explain it. 2. One Gaussian-process model cannot represent "machine-learning capability" generally The machine-learning side of the comparison is represented by a Gaussian process using 256-bit Morgan fingerprints. This is a relatively specific and comparatively simple model representation. The conclusions therefore apply directly to this GP/fingerprint model, but not necessarily to molecular machine learning more broadly. Modern λmax prediction methods may use larger fingerprints, graph neural networks, message-passing architectures, pretrained molecular representations, three-dimensional information, explicit solvent descriptors, or quantum-chemical features. Such models may respond differently to both chemical and label distribution shifts. This issue is particularly important because the title contrasts "model capability" with benchmark effects in general terms. I recommend evaluating several substantially different ML approaches or restricting the conclusions explicitly to the model class studied here. The use of only 256 fingerprint bits also deserves justification. Such a short fingerprint may produce substantial bit collisions for chemically diverse datasets. Performance should ideally be compared with conventional 1024- or 2048-bit Morgan fingerprints to establish that the observed transfer behaviour is not partly a consequence of the selected molecular representation. 3. The TD-DFT comparison may not constitute a uniform computational benchmark The quantum-chemical reference data appear to originate from multiple published sources and include TD-DFT, sTDA, ωB97X-D3, and CAM-B3LYP/6-31G** calculations. These are not interchangeable computational protocols. Differences in functional, basis set, molecular geometry, conformer selection, solvent model, treatment of protonation or tautomerism, and transition assignment can all substantially affect calculated excitation energies. Consequently, describing the comparison simply as "machine learning versus TD-DFT" risks combining methodological heterogeneity with model-class differences. A stronger comparison would apply a single, clearly defined quantum-chemical workflow to the same molecules used for ML evaluation. At minimum, performance should be stratified by computational protocol rather than interpreted as representing a single quantum-chemical method. 4. The very small TD-DFT comparison set limits the strength of the in-corpus conclusion The reported in-corpus TD-DFT comparison contains only 36 compounds. An RMSE difference of 87.3 versus 108.7 nm may appear substantial, but with n = 36 the uncertainty around both values could also be substantial, particularly if a small number of molecules have very large errors. Confidence intervals or bootstrap distributions for the difference in RMSE should therefore be reported. More generally, model comparisons should be paired whenever predictions are available for the same molecules. Reporting distributions of paired absolute-error differences would be more informative than comparing headline RMSE values alone. The larger sTDA comparison is more convincing statistically, but it answers a somewhat different computational question and should not simply be treated as confirming the TD-DFT result. 5. Evaluation in wavelength units may distort the comparison A particularly important issue is the use of RMSE in nanometres. Excitation energy and wavelength are related non-linearly: E ∝ 1/λ. Therefore, a 50 nm error at 250 nm does not represent the same energetic error as a 50 nm error at 650 nm. RMSE measured in nm gives disproportionate influence to errors occurring at longer wavelengths and complicates comparisons across datasets with different spectral ranges. This is especially relevant here because different external datasets may contain substantially different wavelength distributions. The main comparisons should therefore also be reported in energy units, preferably eV or cm⁻¹. Conclusions that depend strongly on whether errors are expressed in nm or eV would require careful interpretation. The same consideration applies to the reported 57–65 nm systematic offsets. A constant wavelength offset does not correspond to a constant error in excitation energy across the spectrum. 6. Removing the systematic offset requires careful validation One of the key claims is that quantum-chemical calculations become substantially more accurate than the GP after each method's "removable offset" is discounted. The procedure used to estimate and remove this offset is therefore crucial. If the offset is estimated from the same test dataset on which residual error is subsequently reported, the corrected error is optimistically biased. This effectively gives the method access to information from the evaluation set. A fair calibration procedure would estimate the correction on an independent calibration or training dataset and then apply that fixed correction to an untouched external test set. The manuscript should explicitly state how offsets were estimated and whether test labels influenced the correction. Both uncorrected and prospectively calibrated results should be reported. Otherwise, comparing a calibrated quantum-chemical method against an uncalibrated GP may not constitute a balanced comparison. 7. Reproduction of a similar quantum-chemical offset does not establish its origin The observation that an approximately 57–65 nm underestimation in the in-corpus data reappears as approximately −61.9 nm in an external dataset is interesting. However, the conclusion that this supports a "definitional rather than functional origin" seems too strong. A reproducible systematic bias could also arise from features shared across quantum-chemical calculations, including the use of vertical excitations, similar density functionals, incomplete treatment of solvent effects, neglect of vibronic structure, conformational differences, or geometry approximations. The recurrence of the offset establishes reproducibility of the bias, but not necessarily its mechanistic origin. To distinguish label-definition effects from computational bias more convincingly, the manuscript would need controlled comparisons involving multiple electronic-structure methods and experimentally assigned transitions on the same molecules. 8. Label convention and chemical-domain shift are incompletely separated The external photoswitch dataset is an important test because it uses a specifically defined E-isomer π→π* transition. However, photoswitches are also a specialised chemical domain. Therefore, deterioration in GP performance may simultaneously reflect changes in molecular structure, photochemical behaviour, chromophore families, wavelength distribution, and target definition. The statement that "chemical novelty alone does not break it, reassigning the label does"

View source

Similar papers

#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11
#computer vision Open access Sep 2012

Making the leap to a software platform strategy: Issues and challenges

A comprehensive taxonomy of the challenges faced when a medium-scale organization decided to adopt software platforms is provided, namely: business challenges, organizational challenges, technical challenges, and people challenges.

Yaser Ghanam, F. Maurer, P. Abrahamsson · 41 citations · ⚡3
#machine learning Open access Mar 2024

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction

MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently, offers an effective and efficient solution for PPI overall property predictions.

Yang Yue, Shu Li, Yihua Cheng et al. · 15 citations

PepPCBench is a Comprehensive Benchmarking Framework for Protein-Peptide Complex Structure Prediction

PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction, and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy.

Si-Long Zhai, Huifeng Zhao, Ji-Ke Wang et al. · 13 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

OmniMol is presented, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability, and achieves state-of-the-art performance in properties prediction.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 11 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.