M2P-DRL: Navigation and Active Perception for Partially Observable Vehicular Networks
Abstract
Autonomous vehicular navigation in unknown environments faces many challenges due to partial observability and limited sensing ranges. Although Vehicle-to-Infrastructure communication provides global topological information to overcome local deadlocks, the associated network overhead and latency constraints often hinder real-time performance. In this paper, we investigate the collaborative navigation and active perception problem to address these constraints by balancing physical routing efficiency with on-demand perception. We propose the Multi-branch Proximal Policy Optimization-based Move-to-Perceive Deep Reinforcement Learning (M2P-DRL) framework, featuring a dual-branch architecture that jointly optimizes movement control and communication requests. To drive efficient exploration in unknown topologies, we design an information-theoretic reward mechanism that balances information gain against dynamic communication penalties. Simulation results demonstrate that the proposed scheme achieves a high navigation success rate and maintains low communication costs.