Skip to content
Conference

Poster: Task-Aware Dynamic Visual Token Pruning for Efficient Vision-Language-Action Inference

Aug 2026 · IEEE International Conference on Embedded and Real-Time Computing Systems and Applications · pp. 228-229 · 0 citations · 3 references

Abstract

Vision-language-action (VLA) models map language instructions and multi-view observations to robot actions, but dense visual-token processing imposes substantial inference overhead. This paper presents TADP, a training-free dynamic visualtoken pruning framework for efficient VLA inference. TADP estimates task-conditioned visual-token importance after early vision-language interaction, decouples main- and wrist-view budgets, reuses keyframe scores under temporal consistency, and refreshes view-specific caches independently. On LIBERO, TADP achieves a 96.25% average success rate while using only 39% of OpenVLA-OFT FLOPs, corresponding to a 61% FLOP reduction.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.