Preprint
Aug 2026
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
Results show that test-time compute can be more effective when used to refine sampled trajectories rather than only to sample more candidates or rely on verifier-guided selection.
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer et al.
· 0 citations