Jul 2026
EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
EvoThink is proposed, a framework that reduces redundant verification and encourages the exploration of new reasoning paths that not only substantially reduces inference-time token usage but also improves the reasoning capability of LRMs.
Xinbang Dai, Zheyu Xin, Hui-Kang Hu et al.
· arXiv.org · 0 citations