Skip to content
Review

Quantitative structure-retention relationships (QSRR): Effect of experimental variables on QSRR model performance in HPLC - A critical review.

Aug 2026 · Journal of Chromatography A · Vol 1786, pp. 467335 · 1 citation · 55 references
Medicine

Abstract

The QSRR modeling framework enables researchers to use molecular descriptors for predicting chromatographic retention, including HPLC retention based on their physicochemical characteristics, as it can predict chromatographic retentions and not only HPLC retention. QSRR models show low transferability between different laboratories and instruments and experimental protocols because their performance depends on experimental conditions which remain poorly documented. The present study investigates the way experimental conditions affect both model architecture and descriptor selection through their experimental design which has not been studied in previous reviews. The review demonstrates how organic modifier type and concentration, mobile phase pH and buffer identity, stationary phase chemistry, and column temperature functions as the fundamental drivers of analyte retention which leads to QSRR model transferability problems. The research examines how dataset variation and laboratory differences and variable relationships affect study results. The study evaluates traditional modeling methods which include Linear Solvation Energy Relationships and Multiple Linear Regression and Partial Least Squares together with modern machine learning techniques which use Random Forests and Gradient Boosting and Support Vector Regression and Graph Neural Networks. The research establishes model validation standards which include applicability domain assessment and descriptor selection and standardized reporting.

View source

Similar papers

Open access Aug 2026

Condition-Specific LC Retention Time Prediction: Feature-Selected QSRR Versus Pretrained Graph Isomorphism Network Transfer Learning

Condition-specific liquid chromatographic retention time prediction remains challenging because retention depends on both molecular structure and experimental conditions. This study compared feature-selected quantitative structure–retention relationship (FS-QSRR) models with pretrained graph isomorphism network (GIN) transfer learning for three reversed-phase LC datasets measured under acidic, neutral, and basic conditions. Mordred descriptor-based QSRR models were developed using leakage-safe preprocessing, nested cross-validation, and feature selection. The selected FS-QSRR workflow for each dataset was then compared with pretrained GIN transfer learning using identical 100 repeated random 80/20 train–test splits. Feature selection substantially reduced descriptor dimensionality but did not consistently improve predictive accuracy over the best baseline descriptor models. Under matched validation, GIN transfer learning gave lower RMSE for all three datasets, decreasing error from 0.816 to 0.613 min under acidic conditions, from 0.868 to 0.665 min under neutral conditions, and from 1.019 to 0.876 min under basic conditions. The corresponding RMSE reductions were 24.8%, 23.4%, and 14.0%, respectively. Matched prediction error analysis showed that GIN particularly reduced the frequency and magnitude of large errors, although the improvement was more modest under basic conditions. Descriptor frequency analysis revealed condition-dependent contributions from lipophilicity, ionization-related, electronic, and topological descriptor groups. These findings support pretrained GIN transfer learning as the stronger predictive approach, while FS-QSRR remains valuable for model simplification and chemical interpretation.

R. Szucs, Emília Sýkorová, I. Boháčová et al. · 0 citations
Jul 2026

Monte Carlo optimization-based QSPR modeling of molar refractivity: Descriptor stability and applicability domain analysis.

An additive fragment-based QSPR model for molar refractivity (MR) was developed using Monte Carlo optimization within the CORAL framework. Molecular structure was encoded using fragment-level attributes, and their contributions were statistically optimized across multiple independent training-validation splits. The resulting linear models retain the additive character of classical fragment-constant approaches while introducing data-driven weight optimization under explicit statistical control. Across three independent splits, the developed QSPR models exhibit consistently high coefficients of determination in both the training and external validation sets, with validation R2 values ranging from approximately 0.86 to 0.89. Root-mean-square errors remained consistent across splits, indicating reproducible behavior of the optimized descriptors. An applicability domain (AD) was defined using statistical defect metrics derived from descriptor distributions. Across independent splits, the applicability domain consistently identified a minority of compounds with rare or unevenly represented SMILES attributes; these statistically under-supported compounds were associated with a modest increase in prediction error, indicating that the defect-based domain reflects descriptor representativeness rather than acting as a strict error filter. The study presents Monte Carlo-optimized additive modeling as a statistically audited extension of fragment-constant schemes, integrating robustness analysis and applicability-domain assessment into a transparent QSPR workflow.

Aleksandar M. Veselinović, Jelena Zivkovic, S. Sunarić et al. · 0 citations
Open access Aug 2026

Revisiting the LSER Approach in the Era of Machine Learning: Insights from IAM Chromatography

The present study demonstrates the integration of the linear solvation energy relationship (LSER) concept with machine learning (ML) methodologies to improve the predictive and interpretative capabilities of chromatographic retention modeling. Immobilized artificial membrane (IAM) chromatography was employed as a model biochromatographic system, and an in-house library of 993 structurally diverse compounds, with experimentally determined chromatographic hydrophobicity index of IAM (CHIIAM), was used to train LSER-ML models. LSER descriptors were calculated using Absolv and extended with ionization-state descriptors to evaluate the applicability of the LSER framework for realistic in silico virtual screening scenarios. Several regression algorithms were tested, including linear, neighborhood-based, kernel-based, and ensemble tree-based models. Among them, the support vector regression with the radial basis function kernel (SVR-RBF) demonstrated the most balanced performance across all validation metrics of R2train = 0.884, R2test = 0.853, and Q2cv = 0.811, achieving predictive errors (RMSEtrain = 5.546, RMSEtest = 4.609, and RMSEcv = 6.989) close to the analytical uncertainty. Model interpretability was achieved using SHapley Additive exPlanations (SHAP), which confirmed the mechanistic relevance of the Abraham descriptors and the dominant contribution of hydrophobic volume and hydrogen-bonding properties to IAM retention. Applicability domain was verified with a Williams plot (±3 standardized residuals and leverage threshold h*). The results indicate that the proposed LSER-ML approach provides an interpretable, robust, and generalizable tool for modeling membrane-mimetic chromatographic systems and can be effectively applied in virtual screening and property-based molecular design.

W. Nisterenko, K. Greber, Magdalena Kierkowicz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.