Skip to content

A reference-guided large language model workflow for mobile phase selection in thin-layer chromatography enabled by a polarity-space tetrahedron strategy.

Oct 2026 · Analytica Chimica Acta · Vol 1418, pp. 345859 · 0 citations · 11 references
Medicine

TL;DR

This work demonstrates that combining a locally constrained reference construction strategy with an LLM offers a practical and interpretable tool for TLC mobile phase optimization, without requiring large labeled datasets or task-specific model training.

Abstract

Mobile phase selection in thin-layer chromatography (TLC) still relies heavily on empirical trial-and-error. Existing machine learning methods often exhibit limited predictive performance, as they depend on manual descriptor engineering and are sensitive to data quality. Here, a reference-guided large language model (LLM) workflow for TLC mobile phase recommendation is proposed. The LLM is used as a flexible inference interface that integrates structural similarity, polarity descriptors, and example-based analogical inference. The key methodological contribution is a polarity-space tetrahedron strategy for selecting reference compounds. Three polarity descriptors, including molecular refractivity, topological polar surface area, and n-octanol-water partition coefficient, are used to construct a three-dimensional space, and a target compound is constrained within a tetrahedron formed by four reference compounds to enable interpolation-based inference. A similarity-weighted version of this tetrahedron strategy is also developed. Using a publicly available high-throughput automated TLC dataset and DeepSeek-Reasoner as the inference engine, four reference selection strategies are compared. The similarity-weighted tetrahedron method achieves the best performance, with an availability of 81.48% and a final score of 0.8436 across 45 compound pairs evaluated in triplicate. The success rate of 92.16% confirmed by independent experimental validation supports the practical relevance of the recommendations. The framework also provides interpretable outputs, including structured rationales and experimental suggestions. This work demonstrates that combining a locally constrained reference construction strategy with an LLM offers a practical and interpretable tool for TLC mobile phase optimization, without requiring large labeled datasets or task-specific model training.

View source

Similar papers

Open access Aug 2026

FastRet: Fast and Simple Retention Time Prediction in Liquid Chromatography

Feature annotation in liquid chromatography–mass spectrometry (LC–MS)-based untargeted metabolomics remains challenging. Retention time (RT) prediction can support candidate prioritization and improve annotation confidence. Here, we present FastRet, an R package predicting RTs using Least Absolute Shrinkage and Selection Operator (LASSO) and Boosted Regression Trees (BRT) on molecular descriptors. FastRet provides a flexible framework combining from-scratch model training, selective measuring to prioritize metabolites for remeasurement, and model adjustment to adapt existing models to changed chromatographic conditions. Model training and prediction are completed within seconds on a single CPU core, and FastRet is accessible both from the R console and through a web interface. We validated FastRet on three in-house data sets covering reversed-phase chromatography (RP; N = 458), RP–anion-exchange mixed-mode chromatography (RP-AXMM; N = 436), and hydrophilic interaction chromatography (HILIC; N = 388), plus one external HILIC data set from the Retip package (N = 970). Using a 2:1 training/test split, BRT models trained from scratch achieved a test-set coefficient of determination (R 2) of 0.86, 0.66, and 0.81 for the three in-house data sets. FastRet can also adjust a model to new chromatographic conditions from a few remeasured metabolites: using 25 RP metabolites measured under six modified conditions, adjustment reached R 2 of 0.74 to 0.84 on unseen metabolites, a mean 0.22 gain over from-scratch models. Compared with published methods on identical splits, FastRet showed competitive performance for de novo prediction and superior performance in low-data transfer scenarios, while generalizing to 14 external data sets (median held-out R 2 0.59). FastRet is available on CRAN with the web interface hosted at https://fastret.spang-lab.de.

F. Fadil, T. Schmidt, Christian Amesoeder et al. · 0 citations
Sep 2026

Consensus descriptor recurrence across complementary ANN workflows in chiral HPLC: An interaction-domain analysis.

Interpreting descriptor relevance in chiral high-performance liquid chromatography (HPLC) remains challenging because enantioseparation depends on multiple correlated structural and chromatographic factors. Here, descriptor-selection recurrence was compared across two consensus artificial neural network (ANN) workflows applied to the same Lux Cellulose-1/aqueous-acetonitrile system. The analysis used 76 learning-stage compounds for model optimization and descriptor-frequency analysis, with two external-test compounds reserved for a limited independent decision-level check. The workflows comprised a previously reported efficient enantioseparation (EES)-mobile-phase profile predictor and a new decision-oriented model for optimal mobile-phase recommendation. To reduce overinterpretation at the single-descriptor level, recurrent selection patterns were examined at both descriptor and interaction-domain levels, acknowledging that multiple descriptors may encode overlapping physicochemical information. Frequency-based analysis revealed distinct interaction-domain trends: the decision-oriented workflow was dominated by hydrogen-bond-related descriptors, whereas EES-profile prediction across the nine tested mobile phases favored topological/structural descriptors. π-π descriptors remained recurrent contributors in both workflows, with greater quantitative prominence in the predictive workflow. These results indicate that descriptor recurrence can serve as an operational proxy to explore and compare interaction-domain trends within the present dataset, but not as direct mechanistic proof. Beyond prediction, the EES-driven ANN framework enabled the extraction of chemically interpretable interaction-domain fingerprints, offering practical guidance for chiral stationary phase/mobile phase selection and cautious data-driven rationalization of chiral recognition.

Carlos Pardo-Cortina, S. Sagrado, L. Escuder-Gilabert et al. · 0 citations

Journal of Chromatography A

Jim Boelrijk, S. A. Molenaar, Tijmen S Bos et al. · 0 citations
Review Open access Aug 2026

How to Build Machine-Learning Models for Molecular Science: A Step-by-Step, Annotated Tutorial

This tutorial provides a comprehensive, end-to-end workflow from raw data to deployed models,icitly designed for environmental chemists with limited prior experience in ML modeling while also providing practical guidance for other users seeking to strengthen their modeling workflows.

Kai Zhang, Yushu Cheng, Hai-Ping Ai et al. · 0 citations
Jun 2025

READ: A Retrieval-Alignment Diffusion Framework for Structure-based Drug Design.

Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared binding patterns and intrinsic similarities among protein-ligand complexes. This fragmented perspective constrains their ability to capture the fundamental principles governing molecular recognition and binding specificity. Moreover, the limited availability of high-quality experimental data further hampers model generalization and real-world applicability. To address these challenges, we present READ, a retrieval-alignment molecular generation framework that conditions the generative process on small molecules targeting homologous proteins. Retrieved ligands are aligned with a diffusion model across multiple representational spaces and integrated as conditional guidance throughout successive stages of generation. Under a standardized docking-based evaluation protocol, READ achieves consistently strong performance against state-of-the-art SBDD methods. More importantly, it introduces a retrieval-alignment paradigm for structure-based molecular generation, offering a practical framework for early-stage computational hit generation while leaving prospective experimental validation as future work.

Dong Xu, Zhangfan Yang, Junchuang Cai et al. · 1 citation
Open access Aug 2026

Condition-Specific LC Retention Time Prediction: Feature-Selected QSRR Versus Pretrained Graph Isomorphism Network Transfer Learning

Condition-specific liquid chromatographic retention time prediction remains challenging because retention depends on both molecular structure and experimental conditions. This study compared feature-selected quantitative structure–retention relationship (FS-QSRR) models with pretrained graph isomorphism network (GIN) transfer learning for three reversed-phase LC datasets measured under acidic, neutral, and basic conditions. Mordred descriptor-based QSRR models were developed using leakage-safe preprocessing, nested cross-validation, and feature selection. The selected FS-QSRR workflow for each dataset was then compared with pretrained GIN transfer learning using identical 100 repeated random 80/20 train–test splits. Feature selection substantially reduced descriptor dimensionality but did not consistently improve predictive accuracy over the best baseline descriptor models. Under matched validation, GIN transfer learning gave lower RMSE for all three datasets, decreasing error from 0.816 to 0.613 min under acidic conditions, from 0.868 to 0.665 min under neutral conditions, and from 1.019 to 0.876 min under basic conditions. The corresponding RMSE reductions were 24.8%, 23.4%, and 14.0%, respectively. Matched prediction error analysis showed that GIN particularly reduced the frequency and magnitude of large errors, although the improvement was more modest under basic conditions. Descriptor frequency analysis revealed condition-dependent contributions from lipophilicity, ionization-related, electronic, and topological descriptor groups. These findings support pretrained GIN transfer learning as the stronger predictive approach, while FS-QSRR remains valuable for model simplification and chemical interpretation.

R. Szucs, Emília Sýkorová, I. Boháčová et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 14, 2026

New method enables AI for safety-critical situations

The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it.

GPT-Lab Sep 10, 2026

Responsible AI Must Consider Its Afterlife

AI may appear weightless, but every model depends on physical infrastructure. To understand responsible AI, we need to look beyond algorithms and consider the entire lifecycle of the hardware behind them. The post Responsible AI Must Consider Its Afterlife appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.