The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model inten...
Ge-Yi Yang, Zi-Kun Qu, Xiang Li et al.· 0 citations
Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Un...
Zi-Hao Chen, Fan-Xiang Xiong, Hong-Ran Ren et al.· 0 citations
The EXPonential-weight algorithm for prompt Optimization} (EXPO) is proposed to automatically optimize the task description and meta-instruction in the meta-prompt for LLM-based agents and is extended to additionally optimize the exemplars (i.e., history of interactions) in the meta-prompt to further enhance the perfor...
Ming-Ze Kong, Zhiyong Wang, Yao Shu et al.· arXiv.org· 7 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.