Skip to content

Autonomous Agent Learning in Production

Unknown authors
· 0 citations · 8 references

TL;DR

ABL is presented, a fully autonomous optimizer that closes the loop between production signals and agent improvements, and takes a data-centric view: each session constructs its own dataset from recent production requests sent to the target agent, anchors evaluation in a project-defined LLM-judge rubric, and then runs an agent-guided tree search over candidate edits to the agent’s source.

View source

Similar papers

Aug 2026

DBAgent: An RL-Based Agent for Autonomous Database Operations and Maintenance

DBAgent is presented, an autonomous agent for Huawei Cloud Data Warehouse Service (DWS) integrated with Autopilot (DWS's production monitoring, alerting, and auto-remediation service) that consumes DWS telemetry and Autopilot alerts and produces evidence-grounded reports with low hallucination on DWS benchmark.

Xu Chen, Jun-Ming Chen, Shuncheng Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

AgentBrew is proposed, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts, and demonstrates that fine-grained offline learning can recover useful supervision from raw trajectories that filterin...

Zhiyi Lyu, Ye-Wen Li, Longtao Zheng et al. · 2 citations
Jul 2026

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

AgentHPOBench, a sequential benchmark comprising 30 executable machine learning tasks across seven research categories, shows that current agents exhibit measurable experimental optimization ability across domains, but still face clear limitations in sustained iterative refinement, complex log diagnosis, and consistent...

Tianyu Huai, Tingshuo Fan, Xinchi Chen et al. · 0 citations
Preprint Aug 2026

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

The same harness runs across five backend LLMs from three model families, indicating the harness generalizes across backends without tuning, even as different models induce distinct execution styles under the same workflow.

Jingsheng Zheng, Xinyuan Fang, Jintian Zhang et al. · 0 citations
Jul 2026

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

MARS is proposed, a search-based framework that formulates MAS repair as a Monte Carlo Tree Search (MCTS) process and navigates the vast space of potential repairs via diagnosis-guided expansion with taxonomy-augmented evaluation.

Hanxiao Lu, Tian-Yi Zhang · 0 citations
Conference Open access Sep 2026

Toward Reliable Agents

An agenda for a shared MDP vocabulary, uncertainty-aware oversight, and scaling human-AI collaboration across volume, complexity, and expertise is opened.

Hua Wei · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.