Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed returns also depend on s...
X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al.· 0 citations
Nontrivial dynamics can emerge in large language model (LLM)-based multi-agent systems, and preliminary evidence exists that formalisms from statistical mechanics can be effective at modeling and predicting such behaviors. In parallel, designing multi-agent communication topology for optimal task-solving is an active r...
Wen-Wen Zheng, Yuzhe Yang, Helen Qu et al.· 0 citations
Neuro-symbolic computer use is introduced, in which a recurring workflow is executed by a learned policy rather than re-derived by an agent on each run, and a pre-action verifier guards each state-mutating step at deployment.
Hyewon Suh, Than Minh Nguyen, Chih-Lun Lee et al.· 0 citations
The results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system.
X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al.· 1 citation
Results show that TRACE converts high model potential into stable, consistent performance gain, and bridge the gap between potential and reliable performance to just 4.0 points.
Wen-Hao Wu, Meng-Hao Zhang, X. Wang et al.· 2 citations
This work introduces Multi- Agent Contextual Exploration (MACE), a lightweight framework that explicitly promotes exploration through structured peer selection that substantially improves exploration behavior and downstream task performance and shows theoretically that the value of exploration increases with agent dive...
Hyeong Kyu Choi, Jiatong Li, Wendi Li et al.· arXiv.org· 1 citation
These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden sta...