Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggre...
Bo-Yuan Wang, Cheng-Yao Yu, Jia-Xi Ren et al.· 0 citations
Large reasoning models (LRMs) often suffer from overconfidence when expressing their uncertainty. Confidence-aware reinforcement learning (RL) offers a promising way to optimize calibration. However, it relies on on-policy rollouts and is thus constrained by the model's pre-RL confidence distribution, which we term con...
Shuo-Yuan Wang, Bei-Er Luo, Hao Zeng et al.· 0 citations
Occupancy-based Quantile Risk Control (OQRC), a novel method that provides tight risk control bounds with finite-sample validity and establishes a finite-sample guarantee showing that OQRC yields tight risk control bounds that converge to the optimal bounds at a provable rate of $\mathcal{O}_ p(n^{-1/2})$.
Zi-Hao Shi, Hua-Jun Xi, Bing-Yi Jing et al.· 0 citations
Tail-Aware Top-$k$ OPD is proposed, a novel distillation method that restores the missing tail probability signal and better aligns the student's next-token distribution with the teacher's, preventing the increase in tail probability and entropy caused by top-$k$ normalization.
Huipeng Huang, Hong-Xin Wei· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.