Skip to content

Author

Wenhan Cao

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

2026

CoIN: Interactive Navigation With Counterfactual Reasoning via Vision–Language Models

Interactive navigation requires robots to actively modify cluttered environments to create traversable paths, going beyond passive obstacle avoidance. However, existing methods either depend on global maps and lack the reasoning capabilities to make interaction decisions from local observations, or are restricted to interactions with simple geometric objects, limiting their applicability in partially observable, unstructured environments. To address these challenges, we propose counterfactual interactive navigation, named CoIN, a vision-language model (VLM)-based hierarchical framework that integrates high-level interaction reasoning with low-level loco-manipulation policies for diverse objects. Specifically, we propose CoIN-VLM, a VLM that internalizes counterfactual reasoning to evaluate the effect of object removal on goal reachability, thereby deciding when interaction is necessary and which object to interact with. To further align such reasoning with the robot’s physical capabilities, we inject robot skill descriptions into the VLM context and ground them into a metric-scale environmental representation, ensuring that the generated plans remain physically feasible. To execute the generated high-level plans, we develop a comprehensive skill library through reinforcement learning, specifically introducing traversability-oriented strategies to manipulate diverse objects for path clearance. Furthermore, a systematic benchmark in Isaac Sim is proposed to evaluate both the reasoning and execution aspects of interactive navigation. Extensive simulations and real-world experiments demonstrate that CoIN significantly outperforms representative baselines, achieving a 17% higher overall success rate and over 80% improvement in complex long-horizon scenarios compared to the best-performing baseline, while exhibiting robust generalization across diverse object categories. Our project page is available at https://coins-internav.github.io/ Note to Practitioners—This work addresses the practical challenge of enabling autonomous robots to reach goals in cluttered indoor environments where the path is blocked by movable objects, without relying on a global map. The primary application is warehouse automation, facility inspection, and disaster-response robots that must decide when to interact, which object should be moved, and how to interact with diverse objects. The fine-tuned vision-language reasoning module determines the timing of interaction and selects the object whose removal is most likely to open a useful path. The learned skill library then executes efficient physical interactions, such as pushing obstacles or opening doors, to create traversable space for navigation. This improves navigation efficiency and reliability by reducing unnecessary detours and allowing the robot to complete tasks that are infeasible for passive obstacle avoidance.

Kangjie Zhou, Zhe-Jia Wen, Zhiyong Zhuo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.