Skip to content

Let VLMs Grade Their Own Thoughts: A Self-Quantification Approach to Reasoning-Aware Reward Modeling

· 0 citations · 38 references

TL;DR

This work proposes leveraging the model’s intrinsic self-evaluation to guide its optimization, and designs two novel reward functions: Sequential Confidence Rigorous Evaluation (SCRE) for challenging problems that demand strict logical reasoning, and intra-group Score Re-ranking (IGSR) for general-purpose, open-ended scenarios.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

This work introduces Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers, and shows that unsupervised reasoning can emerge through cooperative multi-agent training.

Yunhao Yang, Yuexin Bian, Yunjie Tian et al. · 0 citations
Jul 2026

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

This work presents two converging lines of evidence that linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations.

Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit et al. · 0 citations
Jul 2026

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents that achieves improved reasoning performance using similar or fewer inference tokens compared to heavily scaled homogeneous baselines is introduced.

Anmol Kankariya, Sercan Ö. Arik · 0 citations
Preprint Aug 2026

StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning

This work introduces StructReward, a compute-efficient framework that provides dense reinforcement signals through structured step-level reward alignment and substantially reduces the computational overhead of multimodal reinforcement learning.

Yifan Li, Ruxi Sun, Tong-Zhou Zhao · 0 citations
Jul 2026

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

LatentRM is a reward modeling framework that learns intermediate reasoning traces as discrete latent variables to explicitly maximize the likelihood of downstream scalar rewards through on-policy optimization of the latent reasoning space end-to-end.

Sanwoo Lee, Clive Bai, Hsiu-Yuan Huang et al. · 0 citations
Preprint Aug 2026

Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models

CDS is introduced, a training-free meta-reasoning framework equipped with residual demand assessment: at each step, an LLM-based progress evaluator characterizes the residual reasoning required to arrive at a solution rather than merely evaluating the previous step.

John Scoville, Shengzhuang Chen, Yejin Bang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.