Skip to content
Preprint

DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery

Aug 2026 · 1 citation
Computer Science

TL;DR

An LLM-guided SR framework is introduced, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery, demonstrating that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.

Abstract

Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models. Large language models (LLMs) offer a promising avenue for injecting domain expertise into this search. Here, we introduce an LLM-guided SR framework, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery. The LLM performs two roles at each iteration: (1) a qualitative physicochemical critique of the best SR candidates, and (2) the proposal of new candidate rate expressions guided by the SR-generated models and embedded chemical knowledge. Our framework is evaluated on four in silico case studies of increasing complexity, spanning heterogeneous catalysis and bioprocess systems. Results show the LLM-guided framework reduces iterations to identify the ground-truth model by $41.7-79.3\%$ versus a state-of-the-art SR framework, with the LLM directly proposing the correct model structure in over half of the guided runs. In practical settings, where each iteration typically requires a new wet-lab experiment, this translates into a substantial reduction in experimental effort. Predictive performance on an independent validation set is equivalent between both approaches, with $R^2>0.98$ in all case studies. Ablation studies indicate that both the SR component and the LLM scale contribute to this performance, with a reduced-size LLM largely retaining discovery efficiency. These findings demonstrate that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.

View source

Similar papers

Open access Aug 2026

Automated Data-Efficient Symbolic Regression for Interpretable Bioprocess Model Development.

Bioprocessing is central to the sustainable manufacture of pharmaceuticals, food products, and renewable chemicals. Consequently, developing high-fidelity kinetic models to facilitate accurate process prediction, optimisation, and scale-up is a top research priority. However, bioprocess model construction remains hindered in practice by incomplete mechanistic understanding and limited data availability. Therefore, to accelerate the development of accurate bioprocess models, this work presents a data-efficient symbolic regression (SR)-based framework to simultaneously indentify interpretable model structures and aid knowledge discovery. A generic macroscopic kinetic model backbone was used to capture overall process behaviour, while SR was applied to strategically uncover the structures of critical kinetic terms within the backbone. Two implementation strategies were evaluated using an in-silico yeast fermentation case study. The first strategy, embedded SR directly into the kinetic model backbone while the second identified time-varying parameter profiles prior to SR. The results demonstrated that independently identifying individual kinetic terms was crucial for recovering the ground-truth model, while refining SR-generated candidates through a novel local iterative structural correction strategy significantly improved convergence to the true kinetic expressions, surpassing model-based design of experiments in data efficiency. This study therefore enables automated yet interpretable model construction for small-data bioprocess applications, paving the way towards augmented intelligence driven bioprocess modelling and accelerating digital twin development for process optimisation and control.

Luca Riezzo, Alexander W. Rogers, Harry Kay et al. · 0 citations
Preprint Jul 2026

Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System

The MEDA system is introduced, an LLM- and SR-powered agentic framework for discovering ordinary-differential-equation models of biological and biologically inspired dynamical systems and shows that knowledge-guided formalization and mechanistic constraints are load-bearing components, whereas numerical fitting alone can preserve trajectory-compatible but biologically incorrect equations.

D. Krongauz, A. Zulti, Eran Segal et al. · 0 citations
Jul 2026

MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models

Multi-Objective Tool-augmented Symbolic Regression (MOT-SR), a unified framework that integrates external analytical tools to extract structural priors and guide equation generation, while jointly optimizing for accuracy, complexity, and generalization via a multi-objective evaluation module that maintains a dynamic Pareto front is proposed.

Boxiao Wang, Runxian Wang, Kai Li et al. · 0 citations
Preprint Aug 2026

Multi-Granular Rationale-Guided Molecular LLM for Property Prediction

This is the first method to expose GNN-derived attributions to an LLM as evidence for property prediction, and achieves the best overall results among generalist models and narrows the gap to specialist models tuned for each task.

Junwoo Park, Minyoung Shin, C. Lee et al. · 0 citations
Jul 2026

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

SDABench is introduced, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics).

Chuhan Shi, Xiaoquan Ren, Sicheng Song et al. · 1 citation
Review Open access Jul 2026

Leveraging Large Language Models for Understanding Fundamental Principles of Catalysis

This perspective focuses on three opportunities where LLMs can significantly contribute to catalysis: (1) text to properties; (2) text to structure; and (3) text to mechanistic models.

Shane S Michtavy, Sinhara M. H. D. Perera, Marc D. Porosoff · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.