Human-like memory empowers embodied robots for long-term object navigation
Abstract
Effortless object finding by humans, even in cluttered or unseen environments, relies on the seamless integration of perception, memory, and contextual inference. In contrast, embodied robots operating under egocentric perception and partial observability frequently struggle with dynamic spatial relations and long-term consistency, leading to inefficient, repetitive search behaviors. Here we present Human-like Memory Navigation (HM-Nav), a brain-inspired architecture that bridges this cognitive gap by integrating transient sensory inputs with an evolving internal model to enable long-term object navigation. HM-Nav employs three synergistic pillars: (i) Perception: a multi-view fusion module that reconciles viewpoint inconsistencies into unified representations; (ii) Memory: an adaptive dynamic knowledge graph that accumulates semantic-spatial associations over extended timescales; and (iii) Inference: an experience-driven trajectory optimization mechanism that learns from past failures to suppress cyclic and suboptimal search patterns. Validated in simulations and real-world trials, HM-Nav demonstrates superior navigation performance and robust sim-to-real transfer, significantly outperforming existing benchmarks. Our findings suggest that emulating human-like memory structures is essential for achieving resilient, long-term autonomy in complex, open-ended environments.