How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that...
Bang Yang, Jing-Yuan Li, Jia-Jun Fan et al.· 0 citations
This work is the first to unveil the critical role of the null space and harness it for model optimization, and uses a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.
Hong-Bo Ma, Sansheng Cao, Jia-Jun Fan et al.· 0 citations
This work introduces Constraint-First Reasoning (CFR), a training-free two-stage prompting protocol that improves direct CoT on multiple backbones and positions CFR as a targeted test-time intervention whose benefit depends on recoverable constraints and reliable Stage 1 extraction.
Hongbo Ma, Bang Yang, Y. Cheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.