Results support a bounded conclusion: flow-rollout coupling provides additional optimization signal for the evaluated $\tau$-indexed ordinary differential equation (ODE) rollouts, but broader architectural generalization and defense effectiveness require future extensive studies.
Abstract
Flow-matching Vision-Language-Action (VLA) policies generate actions through $\tau$-indexed ordinary differential equation (ODE) rollouts, exposing intermediate velocity-field structure that endpoint-only adversarial objectives do not directly target. We study this mechanism on the evaluated $\pi_{0.5}$ policy in LIBERO and propose the Tau-Path Coupled Attack (TPCA), a white-box visual attack that couples executed-window endpoint displacement with velocity-field divergence along the flow rollout. The study is deliberately scoped as a single-architecture mechanism analysis rather than a general claim about all flow-matching policies. Under matched compute against Visual-PGD, TPCA and endpoint-only optimization both produce near-saturated task failure at $3/255$, while TPCA yields larger action-space endpoint displacement. As the perturbation budget tightens, the task-level gap becomes visible: at $1/255$, TPCA induces a Drop of $0.38$ whereas endpoint-only optimization induces $0.17$, with over $+129\%$ larger action-space endpoint displacement. Cross-task checks on two additional LIBERO spatial tasks show consistent action-space endpoint-displacement advantages for the evaluated setting, while task-failure advantages are budget- and task-dependent. These results support a bounded conclusion: flow-rollout coupling provides additional optimization signal for the evaluated $\pi_{0.5}$ policy, but broader architectural generalization and defense effectiveness require future extensive studies.
This work introduces DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy.
Reachability analysis for visuomotor policies is difficult because large visual encoders make end-to-end set propagation computationally expensive and excessively conservative. We therefore freeze the visual encoder and confine set propagation to a low-dimensional interface between it and the downstream policy, with the interface set calibrated from held-out camera-pose perturbations. Propagating this set through the policy with zonotopes yields a terminal output-enclosure width that set-based training optimizes directly. During evaluation, camera-pose perturbations are sampled from the prescribed distribution, and rollout-level split conformal calibration converts the resulting action-deviation scores into a probabilistic reachable-action radius with finite-sample coverage. In controlled manipulation experiments, set-based training reduces this radius while preserving closed-loop task capability, and matched behavior-only, observational-consistency, and pointwise-adversarial controls all leave a larger radius.
Yanliang Huang, Zhuocheng Zhang, Peng Xie et al.· 0 citations
AdvNav is proposed, a behavior-guided black-box adversarial attack framework that disturbs an agent's first-person views during navigation, which demonstrates the effectiveness and generality of AdvNav, reveals critical perception vulnerabilities and offers insights for the design of future resilient VLN models.
Chenyang Li, Kaige Li, Zeyu Jiang et al.· 0 citations
Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stochasticity is largely passive: repeated sampling may yield diverse behaviors, but users cannot directly choose among valid continuations from the same state. We propose Source-Lifted Flow Matching (SL-FM), a source-intervenable flow-matching policy that exposes such a handle while keeping the velocity field shared and latent-free. The handle selects only the source endpoint of the conditional flow, not a mode-specific field, preserving the standard formulation while avoiding decomposition into separate mode-conditioned dynamics. The core mechanism is \textbf{Orthogonal Source Lifting}, designed to prevent path-crossing ambiguity. Instead of partitioning target actions by mode, SL-FM lifts handle-specific sources into auxiliary orthogonal coordinates and keeps targets in the original action subspace. This preserves the demonstrated action distribution while allowing one shared field to carry different branches without merging at crossings. To keep handles usable across states, we learn a state-dependent source mixture end to end and use a responsibility floor, giving each handle weak supervision and mitigating dead modes. Experiments on crossing-flow diagnostics and robot-control benchmarks show that SL-FM converts passive source randomness into an actionable intervention variable. It removes crossing-induced composite trajectories, changes future routes in 91.1\% of matched-prefix interventions, and achieves strong free-deployment performance, with improvements in several benchmark settings. Overall, source geometry provides actionable multimodal control without conditioning the velocity field on the selected mode.
He Zhang, Ying Sun, Pengteng Li et al.· 0 citations
Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state to derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps.
Kaiyang Ye, Yuan Ge, Junxia Zhang et al.· arXiv.org· 0 citations
Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused by distribution shifts during deployment. While offline reinforcement learning (RL) provides a practical way to improve deployed policies using rollout data, existing methods either ignore failure data or exploit it only at the trajectory level, resulting in low learning efficiency and persistent errors. We propose **RedFlow**, a fine-grained offline RL framework that redirects failure experiences into action-level corrective supervision for flow-matching VLA policies. RedFlow consists of two key components: (1) a **Context-Aware Corrective Matching** mechanism that identifies failure-inducing actions and retrieves successful alternatives from similar contexts as corrective targets, and (2) an **Adaptive Redirection Objective** that jointly reinforces successful actions, suppresses undesirable ones, and redirects recoverable failures toward corrective targets. By converting both successful and failed experiences into dense supervision, RedFlow enables robust recovery learning from mixed-quality data. Experiments on the LIBERO benchmark and three real-world manipulation tasks show that RedFlow consistently outperforms state-of-the-art offline RL baselines, improving the real-world success rate from 56.7% to 74.7%. It also matches strong on-policy methods (PPO, GRPO, and DDPO) while requiring roughly an order of magnitude fewer training samples.
Zhengyang Yan, Junhao Li, Fangqi Zhu et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.