Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent prefere...
Dong-Chan Shin, Xing Han Lù, Jiaqi Deng et al.· 0 citations
A variant of FocusAgent significantly reduces the success rate of prompt-injection attacks, including banner and pop-up attacks, while maintaining task success performance in attack-free settings, highlighting that targeted LLM-based retrieval is a practical and robust strategy for building web agents that are efficien...
Imene Kerboua, S. Shayegan, Megh Thakkar et al.· arXiv.org· 15 citations
These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden sta...