Scouting the Dynamics Gap: Test-Time Policy Adaptation via Action-Outcome Feedback
While pretrained robotic policies exhibit impressive capabilities in controlled environments, unobserved physical properties and dynamics require these policies to rapidly adapt during deployment. Existing test-time adaptation methods typically rely on sparse scalar rewards, failing to exploit the rich geometric and dy...