This work proposes that LLM web agents can learn simple environment observations at test time, and introduces trial steps for agents to decompose a complex environment observation into sub-modules, and implements a label-free learning method, Test-Time Environment Decomposition (TTED), to adapt agent behaviors with experience during inference.
Jun-Xuan Li, Zijun Liu, Zi-Yi Huang et al.· 0 citations
State2State is proposed, an environment-derived mid-training method that converts explored environment states into training objectives, challenging agents to reach a specified target state by deriving tasks from environment exploration and verifying success through rule-based state matching.
Xuanyu Lei, Yiqi Zhu, Chenliang Li et al.· 1 citation
GMA is presented, a benchmark for evaluating general mobile assistants in challenging real-world scenarios, and shows that appropriate harness design can meaningfully improve performance, particularly on demanding workflows, while the effectiveness of specific designs can vary across foundation models.
Yi-Qi Zhu, Feiyu Gao, Jiakang Fan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.