Skip to content
Book

Improving Usability and Productivity of PETSc with Agent-Based Workflows

Jul 2026 · Practice and Experience in Advanced Research Computing · pp. 1-5 · 1 citation · 11 references
Computer Science

TL;DR

This work position PETSc as a domain-aware component in multi-step AI workflows that span question answering, code development, execution, and verification, and describes the infrastructure and prototype services that support these capabilities and outline their potential to enable more effective, reliable, and scalable AI-assisted workflows in scientific computing.

Abstract

Scientific computing software such as PETSc embodies deep expertise in numerical methods, solver configuration, and scalable implementation, yet this knowledge remains difficult for both users and large language models (LLMs) to access and apply effectively. As a result, even experienced researchers spend significant time selecting solvers, debugging configurations, and validating results. In this work, we present initial experiences in developing an AI-assisted, agent-based ecosystem to improve PETSc usability and productivity for scientific applications. Our approach integrates retrieval-augmented generation (RAG), grounded in PETSc manual pages and curated documentation, with modular services for code generation, compilation, execution, and validation, all exposed through lightweight agent interfaces. We position PETSc as a domain-aware component in multi-step AI workflows that span question answering, code development, execution, and verification. We describe the infrastructure and prototype services that support these capabilities and outline their potential to enable more effective, reliable, and scalable AI-assisted workflows in scientific computing.

View source

Similar papers

Preprint Aug 2026

CURATE: Leveraging LLM Agents to Compose, Catalog, and Deploy Reproducible Workflows

This work proposes CURATE - Composition, User-in-the-loop, Reuse, and Automated Task Execution - a novel human-in-the-loop multi-agent system that uses LLM agents to manage and develop composable workflows across their entire lifecycle.

Nolan Cutler, Chia-Chen Kuo, Nanda Velugoti et al. · 0 citations
Jul 2026

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents

SyntheticAgentTraceQA is proposed, an execution- first framework for generating scalable supervision data for tool- augmented agents and shows that execution-grounded supervision improves tool execution behavior, reference-trace agreement, and answer-generation performance on the evaluated tasks.

Hafsa Ouajdi, Francesco Giannuzzo, Alaa Boukhary et al. · 1 citation · ⚡1
Jul 2026

Knowledge-Centric Agents for Workflow Generation in ComfyUI

This work argues that successful workflow generation requires modeling knowledge itself, including its structure, hierarchy, and reasoning dynamics, and proposes a knowledge-centric framework that learns to invert, inject, and infer with knowledge across multiple abstraction levels.

Zhen-Dong Li, Lei Sun, Ruibo Ming et al. · 2 citations
Conference Jul 2026

KG-Augmented LLM for Efficient and Correct Dockerfile Generation

Container images are fundamental to cloud deployment, with their build instructions (e.g., Dockerfiles) critically impacting the efficiency and stability of cloud service. Manually authoring these instructions is error-prone, while Large Language Models (LLMs) lack the domain knowledge to generate both correct and optimized Dockerfiles reliably. This problem may cause runtime failures, prolonged deployment times and increased storage overhead. This paper introduces a novel knowledgeenhanced approach to automate Dockerfile generation. First, we construct a Dockerfile Instructions Knowledge Graph (DIKG) by analyzing a large corpus, capturing complex dependencies among images, packages, and commands. Leveraging DIKG, we design DKRAG, a retrieval-augmented generation system that guides an LLM to interpret user requirements and produce semantically accurate instructions. The output is further optimized via log-based repair and static dependency-aware refactoring for correctness, layer sharing, and minimal image size. Comprehensive experiments show our approach significantly improves the generation accuracy while also reducing build time and storage overhead compared to state-of-the-art methods.

Kun Wang, Yao Wu, Hao Fan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

UI-Venus-2 Technical Report

Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.

Venus Team, Zhuo-Hang Cai, Haoxin Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environments, or scoring protocols, limiting their comparability, interpretability, and reliability for deployment decisions. We introduce DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents), a benchmark designed to capture workload variation and support reliable agent evaluation. Built on a shared OpenClaw execution environment, DAREBench organizes 233 tasks selected and adapted from 22 source benchmarks into a $2\times3$ workload matrix defined by input modality and execution form, and evaluates them under a unified contract-based protocol with evidence-based score auditing. We evaluate 23 commercial API models and 12 locally deployed open-weight models over 7,587 model--task runs, reporting accuracy and token consumption alongside reference costs for API models. Results show that no single model dominates all workload groups, text and multimodal tasks exhibit distinct accuracy--cost trade-offs, and local open-weight models are competitive in several groups but still trail frontier commercial models overall. These findings suggest that agent deployment and model selection should consider workload profiles, deployment mode, and accuracy--cost trade-offs rather than rely on a single aggregate score.

Yu Liu, Zhi-Lin Liu, Zhi-Wei Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.