Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models
Artificial Foveated Perception is proposed, a lightweight, policy-agnostic module that takes the same vision and language inputs as Vision-Language-Action and World Action Model pipelines and predicts task-conditioned masks over relevant objects, the robot, and other action-critical regions and reduces fine-tuning time, suppresses overfitting, and improves generalization under environmental perturbations.