Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive reasoning, which limits systematic search, refinement, and backtracking. While recursive models such as Hierarchical Reasoning Model (HRM) and Tiny Recursive Model (TRM) address this limitation through iterative latent-state refinement, they are typically task-specific and do not leverage pretrained language priors. We propose R-Qwen, a recursive reasoning framework built upon a pretrained Qwen backbone. R-Qwen repeatedly refines a candidate solution through programmatic self-recursion and deep supervision, combining the structured iterative computation of recursive models with the linguistic and reasoning priors of pretrained LLMs. We further adapt Hierarchical Supervision Weighting (HSW) to autoregressive models by exponentially weighting losses across recursive steps. HSW reduces gradient variance by at least 50\%, improves the signal-to-noise ratio of stochastic gradients, and accelerates convergence. Across eight challenging benchmarks, R-Qwen consistently outperforms prior recursive reasoning models and substantially larger LLMs while using a comparable number of trainable parameters. Notably, on ARC-AGI dataset, our model achieves a 27.6\% improvement over the baseline, highlighting the effectiveness of recursive refinement for general symbolic reasoning. These results suggest that recursive reasoning mechanisms and pretrained language model priors are complementary approaches for improving symbolic puzzle-solving. Code and models will be released after acceptance.
Omid Nejati Manzari, Guillaume Lajoie, H. Rivaz· 0 citations
This work interprets chain-of-thought reasoning as a latent variable modeling problem and demonstrates that this distribution-matching paradigm of LLM fine-tuning can serve as an effective alternative to maximum-likelihood training and reward-maximizing policy optimization.
Edward J. Hu, Moksh Jain, Eric Elmoznino et al.· International Conference on...· 110 citations· ⚡19
A new algorithm for amortized inference in sparse probabilistic graphical models (PGMs) is presented that enables off-policy training but avoids the need to instantiate all the random variables for each parameter update, thus speeding up training considerably.
J. Falet, Haebeom Lee, Esmeralda S. Whitammer et al.· International Conference on...· 9 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.