Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult. Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachm...
Zhi-Yuan Qi, Jie-Rui Li, Yi-Fan Shen et al.· 0 citations
An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Err...
Kun-Lun Zhu, Xu-Yan Ye, Yi-Bo Li et al.· 0 citations
This survey synthesizes agentic reasoning methods into a unified roadmap bridging thought and action, and outlines open challenges and future directions, including personalization, long-horizon interaction, world modeling, scalable multi-agent training, and governance for real-world deployment.
Tian-Xin Wei, Ting-Wei Li, Zhining Liu et al.· 39 citations· ⚡6
The results reveal a pattern distinct from prior findings on long-tail vulnerability during acquisition and retention: among facts that models already answer correctly, those associated with highly connected entities are more likely to be corrupted by neighboring updates, and updates to such facts propagate errors more...
Yu-Ji Zhang, Wei-Bing Wang, Cheng Qian et al.· 0 citations
PRISK is proposed, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses.
Yumeng Wang, Yu-Chen Wu, Cheng Qian et al.· 0 citations
This survey focuses on co-evolution in agentic systems, a multi-component form of self-evolution in which multiple agents and their environment impose adaptive pressure on one another.
Qing Zong, Jiayu Liu, Junhao Shen et al.· 0 citations
MetaEvolve is presented, a framework designed to develop meta-skills that can transfer broadly to open-ended problems where such rich training signals are scarce, and aims to inspire generalizable domain-agnostic meta-skills that can transfer broadly to open-ended problems where such rich training signals are scarce.