TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models
TTRSD separates update direction from update position, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher, demonstrating cross-dataset generalization while preserving inherent reasoning integri...