Preprint
Aug 2026
Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models
AtVLA, a framework that inserts learnable register tokens into the visual encoder and improves the average LIBERO success rate, is introduced, a framework that inserts learnable register tokens into the visual encoder and improves the average LIBERO success rate.
Jin Cui, Yanbin Hu, Xinyue Long et al.
· 0 citations