Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference
A system level acceleration strategy that reduces computation in both perception and action generation and compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision is proposed.