AdvNav is proposed, a behavior-guided black-box adversarial attack framework that disturbs an agent's first-person views during navigation, which demonstrates the effectiveness and generality of AdvNav, reveals critical perception vulnerabilities and offers insights for the design of future resilient VLN models.
Abstract
Despite progress in Embodied AI, Vision-and-Language Navigation systems remain vulnerable to adversarial visual disturbances. Most existing methods rely on white-box access to target model gradients, which is often unrealistic for real-world deployed systems and computationally exhaustive due to recursive backpropagation for optimization, limiting their applicability. While previous black-box methods predominantly target single-step, instantaneous decision tasks, they struggle to handle the task complexities and temporal dependencies. This highlights the need for a gradient-free attack method that can effectively disrupt the multistep sequential perception-action loop using only observable inputs and outputs. Therefore, we propose AdvNav, a behavior-guided black-box adversarial attack framework that disturbs an agent's first-person views during navigation. To construct an informative surrogate objective for effective optimization guidance in gradient-free search under the black-box setting, we design a dual-granularity behavior-based feedback, aggregating a trajectory-level performance score representing overall navigation degradation, an action-level reward score considering the potential decision risk, and a deviation indicator, all of which are extracted from the agent's self-output behaviors. This feedback guides a hybrid optimization strategy that heuristically tunes perturbation strength via adaptive updates and evolves noise spatial structure genetically, to iteratively discover the most disruptive noise configuration. Evaluated against Transformer-based HAMT and LLM-based MapGPT with two types of backbones on R2R dataset, AdvNav achieves 49.70/65.96/87.30% Attack Success Rate. The result demonstrates the effectiveness and generality of AdvNav, reveals critical perception vulnerabilities and offers insights for the design of future resilient VLN models.
Results support a bounded conclusion: flow-rollout coupling provides additional optimization signal for the evaluated $\tau$-indexed ordinary differential equation (ODE) rollouts, but broader architectural generalization and defense effectiveness require future extensive studies.
Mengxiang Liu, Ruilong Deng, Rong Guo et al.· Security and Safety· 0 citations
Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision. To address this, we propose an anticipatory risk-guided reinforcement learning framework. Leveraging privileged simulator states, we construct a directionally aligned future collision risk map based on the Closest Point of Approach (CPA). Through an asymmetric actor-critic architecture, the network is trained to self-predict this structured risk, which explicitly guides the visual policy during deployment. A lightweight spatio-temporal encoder extracts motion cues directly from onboard depth sequences, bypassing explicit object tracking or optical flow estimation. Extensive simulated and real-world experiments demonstrate that our method effectively improves safety margins and flight efficiency in dense dynamic clutters compared to existing baselines. Furthermore, the learned policy achieves robust zero-shot Sim-to-Real transfer on a physical quadrotor, relying purely on abstracted spatio-temporal depth sequences and its self-predicted risk priors, validating the effectiveness of our approach and its robust generalization from simulation to reality.
Yuchao Mei, Guohao Zhang, Luxia Ai et al.· arXiv.org· 0 citations
UniTexture is introduced, a cross-task universal adversarial texture attack that uses a single textured 3D object to induce targeted deviations in VLA action predictions across multiple tasks and reveals shared cross-task vulnerabilities in multitask VLAs that can be systematically exploited through a single adversarial surface texture.
Yu Dai, Mingzhe Dai, Tianshi Wang et al.· 0 citations
Reactive local planners such as the Dynamic Window Approach (DWA) often fail in non-convex dead-end structures because their greedy objective drives the robot into local minima. This paper presents a ROS 2-native mapless navigation framework that trains a Proximal Policy Optimization (PPO) agent using a Procedural Adversarial Trap Generator (ATG) in Gazebo. The generator systematically produces U-shaped traps, corners, and narrow passages so that the agent learns proactive avoidance and recovery behavior rather than merely reacting to nearby obstacles. In simulation, the proposed DRL policy achieves an 88% success rate in complex maze scenarios, while the DWA baseline drops to 8%. A zero-shot deployment on a physical Unitree Go2 further confirms that the learned behavior transfers to real hardware despite LiDAR noise and odometry uncertainty.
Ardiansyah Al Farouq, Yuya Hosoda, Jooho Lee· 2026 23rd International Conf...· 0 citations
FLARE is proposed, an optimized physical spotlight attack framework that exploits vulnerabilities to minor environmental perturbations via targeted illuminations, dropping baseline task success rates to zero without any access to model internals.
Reachability analysis for visuomotor policies is difficult because large visual encoders make end-to-end set propagation computationally expensive and excessively conservative. We therefore freeze the visual encoder and confine set propagation to a low-dimensional interface between it and the downstream policy, with the interface set calibrated from held-out camera-pose perturbations. Propagating this set through the policy with zonotopes yields a terminal output-enclosure width that set-based training optimizes directly. During evaluation, camera-pose perturbations are sampled from the prescribed distribution, and rollout-level split conformal calibration converts the resulting action-deviation scores into a probabilistic reachable-action radius with finite-sample coverage. In controlled manipulation experiments, set-based training reduces this radius while preserving closed-loop task capability, and matched behavior-only, observational-consistency, and pointwise-adversarial controls all leave a larger radius.
Yanliang Huang, Zhuocheng Zhang, Peng Xie et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.