Let VLMs Grade Their Own Thoughts: A Self-Quantification Approach to Reasoning-Aware Reward Modeling
This work proposes leveraging the model’s intrinsic self-evaluation to guide its optimization, and designs two novel reward functions: Sequential Confidence Rigorous Evaluation (SCRE) for challenging problems that demand strict logical reasoning, and intra-group Score Re-ranking (IGSR) for general-purpose, open-ended scenarios.