Aug 2026· Innovations in Systems and Software Engineering· Vol 22· 0 citations· 46 references
Computer Science
TL;DR
This work studies model management tools for transformation and editing in Model-Based Engineering and proposes an approach based on complementary mechanisms that select more accurately modeling tools to perform user instructions than off-the-shelf agents.
Abstract
In Model-Based Engineering (MBE), practitioners frequently have to choose appropriate tools from many different options. LLM-based agents are software components that depend on Large Language Models (LLMs) to autonomously select and apply software tools to perform specific tasks. Although LLMs have already been used in the MBE context, LLM-based agents to assist users of MBE tools remain underexplored. This is particularly true in industrial environments where only medium-sized on-premise LLMs can be considered due to policies related to security or data privacy. To investigate the potential of LLM-based agents for MBE, we study model management tools for transformation and editing. Currently, off-the-shelf agents such as Microsoft Copilot can invoke model management tools when the task is explicitly described. However, these agents struggle to select the correct transformation or operation when they only have limited contextual information, especially when coupled with medium-sized LLMs. To overcome this, we propose an approach based on complementary mechanisms. First, we provide a server and associated LLM-based agent with dedicated tools for each transformation available on this server. We also provide a similar server and agent for model editing operations. Then, to enable the two agents to efficiently select transformations and operations (respectively), we rely on a tool retrieval technique based on a tool relevance score computed by an LLM. We evaluate these agents using generated model management datasets that we contribute to the community. The obtained results show that our LLM-based agents select more accurately modeling tools to perform user instructions than off-the-shelf agents.
Results show that DocsChisel improves the task success rate of LLM agents by 95.89% over the original tool documentation and by 75.15%, on average, over existing baselines, while incurring limited optimization time and token overhead.
You Lu, Kun Zhang, Bihuan Chen et al.· 0 citations
A systematic literature review of technical approaches, including agent architecture, perception, memory, reasoning and planning, action space, orchestration, and self-improvement, reveals a field that has built agents able to act but not yet agents whose authority is bounded or whose behavior is auditable.
Jing-Jing Nie, Jiawei Guo, Krishna Meda et al.· 0 citations
This paper constructs a large-scale dataset of agent applications, tools, and tests, and manually label 2,572 test methods from 240 modules, and derives a taxonomy of 23 testing patterns across test fixtures, data, objectives, and assertions, and characterize tests by level.
Rangeet Pan, Tyler Stennett, D. Sankar et al.· 1 citation
The use of large language models (LLMs) to support non-modeling experts of multi-perspective EM are investigated and LLMs can be seen as assistive technology for certain tasks in EM.
Peter-Alexander Kolev, H. Pruss, J. Wilken et al.· Journal of Software and Syst...· 0 citations
A comprehensive overview of the existing tools and frameworks for implementing MAS in software engineering and a set of lessons learned and challenges that can help researchers and practitioners to select a suitable MAS framework according to their needs are provided.
Maria Sâmyla Serafim de Oliveira, M. Ibiyo, Marco Gianrusso et al.· 0 citations
A paired benchmark that tests whether LLM agents exhibit conscious allocation behavior under a fixed budget in two contexts: an abstract text-based formulation and a code-construction task finds that every frontier model testedacts near-optimally in the abstract framing but fails to transfer this ability to script-writing.
Daniel Wang, Andrew Xu· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.