Results show a unified System-2 agent enables adaptive humanoid service across simulation and reality.
Abstract
We propose HODAgent, a System-2 embodied agent for humanoid robots in service settings, addressing situated intent, responsive execution, task revision, and outcome verification. Its semi-duplex architecture integrates an Env-Interactor, Planner, Executor, and hierarchical Memory to maintain coherent interaction, planning, and task state during service episodes. This allows handling new requests during motion, retaining progress, revising actions, and grounding closure in execution outcomes. A shared interface connects simulation and physical robots (Unitree G1), isolating platform-specific control. In an interactive simulation with 164 cases, HODAgent achieves 84.8% and 91.5% Joint Success under two VLM backbones, outperforming baselines by 9.8 and 18.9 points. On physical robots, pass rates are 92% (atomic), 72% (composite), and 63.3% (complete tasks). On multiple embodied benchmarks, it improves over baselines by 0.7-9.0 points. Results show a unified System-2 agent enables adaptive humanoid service across simulation and reality.
The World-Cognition Model is presented, a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime and introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks.
The Embodied Task Agent is introduced, a new paradigm for extending digital agents into the physical world, and OpenETA is released as its open-source implementation, which provides replaceable Planners, composable Tools and Skills, auditable memory, replayable trajectories, and common interfaces for simulation and real robots.
Yitong Chen, Zezheng Huai, Sixian Li et al.· 1 citation
CoMuRoS enables runtime, event-driven replanning on physical robots and supports flexible multi-robot and human-robot collaboration across diverse scenarios.
Suraj S. Borate, Bhavish Rai B, Vipul Pardeshi et al.· Frontiers in Robotics and AI· 0 citations
RoboBRIDGE is presented, a modular framework that provides an orchestration layer over five coordinated modules, namely Monitor, Perceptor, Planner, Controller, and Robot Interface, to compose robust robotic agents from off-the-shelf components, including pretrained VLAs.
Agent-based modeling and simulation (ABMS) has been widely employed to study emergent processes in collective robotic construction (CRC), where global architectural structures arise from local agent interactions. While these approaches reveal how complex assemblies can emerge without centralized control, they remain limited when an architectural goal is known but the behaviors required to achieve it are not. Most CRC workflows still depend on handcrafted heuristics. This paper presents a hybrid CRC workflow that integrates large language models (LLM) into the ABMS behavior design process. The system enables human–AI co-creation of robot behaviors, allowing an LLM agent to generate and negotiate behavioral strategies toward user-defined construction goals under partial observability. The approach is evaluated in simulation, comparing an LLM–heuristic hybrid against a heuristic-only baseline behavior. For well-documented swarm patterns, the LLM matches heuristic performance; for geometrically novel tasks, handcrafted heuristics retain an advantage. By embedding language-based reasoning within ABMS, this work expands participation in CRC behavior design and demonstrates a pathway for translating high-level design intent into adaptive, goal-oriented multiagent construction processes.
Samuel Slezák, Lasath Siriwardena, S. Leder et al.· Construction Robotics· 0 citations
We consider a task planning scenario in which robots sharing a persistent environment are assigned tasks one at a time from a held-out sequence. Standard task planners, lacking foresight of future tasks and inconsiderate of others'constraints, solve each task in isolation, leaving terminal states that increase future cost for all, side effects that compound over lengthy task sequences. To reduce cost over the sequence, a robot must anticipate how its actions now may impact performance on future tasks for all robots sharing the environment. Therefore, we present courteous anticipatory planning, wherein a model-based planner proposes candidate plans and selects the one that jointly minimizes immediate cost and aggregated expected future cost across all robots, estimated via independent per-robot learned estimators. This factored formulation avoids combinatorial joint rollouts and supports modular deployment: adding a robot requires only training its own estimator. We evaluate in two persistent PDDL domains, a home environment with robots that have similar capabilities but different responsibilities, and a restaurant environment where robots'distinct capabilities create states that other robots lack the capability to resolve. During lengthy task sequences, our planner reduces total cost by 10.43% versus myopic and 4.03% versus selfish anticipatory planning in a two-robot home environment and by 17.41% and 13.24%, respectively, in a three-robot restaurant.
Md Ridwan Hossain Talukder, Roshan Dhakal, Elizabeth Phillips et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.