DeBERTa-based and T5-based models that rank HTML elements by their relevance to the task and a zero-shot ColBERT-based retriever that is able to retrieve the ground-truth element with a recall of 0.52 on Mind2Web and 0.47 on WebArena are developed.
Abstract
Autonomous web agents, powered by Large Language Models (LLMs), have garnered significant attention for automating various web-based tasks with multi-step reasoning and decision-making capabilities. An open research question in the development of these agents lies in the format of the webpage input. Raw HTML source code, with its extensive and often irrelevant details, poses difficulties for LLMs with limited context windows. To address this challenge, we first reproduce baseline models such as GPT-3.5 and LLaMA-2-70B on the WebArena (Zhou et al., 2023) benchmark, identifying common failure modes. We then propose two retrieval strategies to filter out irrelevant context for LLM agents. We develop DeBERTa-based and T5-based models that rank HTML elements by their relevance to the task. We fine-tune them on Mind2Web trajectory data and transfer them to WebArena. Experiments show that our DeBERTa-based model improves the success rate of the LLaMA-2-70B LLM agent on WebArena from 1.97% to 2.96%. Moreover, we develop a zero-shot ColBERT-based retriever that is able to retrieve the ground-truth element with a recall of 0.52 on Mind2Web and 0.47 on WebArena.
It is found that correctly inferring user preferences is necessary but not sufficient for task success, as execution failures in downstream web interactions remain a significant bottleneck even when agents align with the target preference.
Dong-Chan Shin, Xing Han Lù, Jiaqi Deng et al.· 0 citations
A novel RAIE taxonomy along four scaling dimensions is proposed, which optimizes the entire thought process through search algorithms and self-verification, and introduces a task-oriented guideline for choosing the best TTS strategy.
Jia-Yu An, Zheng Chen, Yong-Cheng Jing et al.· Proceedings of the Thirty-Fi...· 0 citations
AnyAct is a universal action layer that unifies available capabilities into a self-evolving action space, enabling agents to operate efficiently and reliably in large-scale, dynamic tool ecosystems and optimizes for a balance between task success rate and execution cost.
Ling-Rui Xu, Ya Jiang, Jia-Chang Zhang et al.· 0 citations
We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the sit...
Collaboration between a small language model (SLM) and a large language model (LLM) offers an opportunity to combine the efficiency of smaller models with the strong reasoning capabilities of larger ones. Existing approaches primarily frame such collaboration as a computation allocation problem, determining which model...
This work proposes that LLM web agents can learn simple environment observations at test time, and introduces trial steps for agents to decompose a complex environment observation into sub-modules, and implements a label-free learning method, Test-Time Environment Decomposition (TTED), to adapt agent behaviors with exp...
Jun-Xuan Li, Zijun Liu, Zi-Yi Huang et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.